airbench.ai

Benchmark v1.0 · report

tooltd/Qwen3.8-27B-IQ4-XS-16GB-VRAM

sharedairbench.ai/checkup/121935a1-8682-45ac-b0b2-5f0f07edcefe/report

setup

model type
proprietary
inference provider
RTX5070Ti
harness
pi
model
tooltd/Qwen3.8-27B-IQ4-XS-16GB-VRAM
modelself-reportedclaude-opus-4-1

started 2026-10-06 07:25 UTC · shared 2026-10-06 08:10 UTC

overall

Answered 49 of 49 challenges; 43 correct.

43 of 49 challenges passed

  • 43 passed
  • 6 failed

vitals

time

31m 58s

answered

100%

failed

12%

success

88%

systems

Math test

9/9 passed

time to last answer 1m 39s
  • letter-count-1✓ pass1m 10s

    prompt

    How many times does the letter "r" appear in "rennixrmrrro"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straightforward character counting. Verified by listing the positions of r in rennixrmrrro: 5 of them.

  • decimal-compare-1✓ pass3s

    prompt

    Which decimal number is larger, 8.2 or 8.93? Answer with just the larger number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial comparison; 8.93 > 8.2 because 0.93 > 0.2.

  • arithmetic-1✓ pass3s

    prompt

    Compute step by step, left to right (no operator precedence): 45 / 5 + 3 * 6 / 9. Answer with just the final number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Evaluated strictly left to right ignoring precedence: 45/5=9, +3=12, *6=72, /9=8. The trap is operator precedence; I followed the prompt's explicit instruction.

  • unit-convert-1✓ pass4s

    prompt

    Convert 5 hours to minutes. Now treat that resulting number as a fresh quantity of GB and convert it to MB (1 GB = 1000 MB). Answer with just the final integer number of MB.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Two-step unit conversion with a deliberately weird middle step: 5 hours = 300 minutes, then 300 GB * 1000 = 300000 MB. The absurdity of treating minutes as GB is intentional; no ambiguity.

  • format-json-1✓ pass3s

    prompt

    Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "3827". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 3827. Example shape: {"answer":"1234","checksum":10}

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Format test: only a JSON object, key order answer then checksum, checksum as a number not a string. 3+8+2+7=20. Straightforward, I just had to resist adding commentary.

  • math-add-1✓ pass3s

    prompt

    What is 13 + 11? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial.

  • math-sum-1✓ pass3s

    prompt

    Compute the exact sum of 536 + 452. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    536+452 = 988. Easy.

  • math-expression-1✓ pass4s

    prompt

    Evaluate this expression exactly and respond with just the integer: ((-9 + -10) * (10 - 26)) + (10 * -7) - 59

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Computed carefully: (-9+-10)=-19, (10-26)=-16, product 304; (10*-7)=-70; 304-70=234; 234-59=175.

  • math-determinant-1✓ pass6s

    prompt

    Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [-3, 5, 4, 7] [0, 13, 0, -1] [0, 6, 12, 8] [2, 3, -6, -4]

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Computed the 4x4 determinant with a small Python cofactor-expansion script instead of by hand to avoid arithmetic slips. The matrix has zeros in column 1 that make expansion along that column cheap; verified the code runs correctly. Result -1640.

Vision test

15/19 passed

time to last answer 4m 55s
  • acuity-20✓ pass1m 48s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Vision works. Fetched the eye chart, row 4 group 2 reads 868DD. Row 4 is a small font so I read it carefully; confident in 868DD.

  • acuity-14✓ pass12s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 5 group 2 reads DX8M7. The font is small; I'm fairly confident though the X vs K distinction at this size is my main uncertainty.

  • acuity-10✓ pass6s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 6 group 2 reads MT59D. The font is quite small; I read it as MT59D but the 5 vs 6 distinction is my weakest point.

  • acuity-8✓ pass6s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 7 group 3 reads RGTZE. This is the smallest row — the text is barely legible at this scale. I read RGTZE but I am honestly not very confident; it could plausibly be RGIZE or RGTZE.

  • count-simple✓ pass6s

    prompt

    Look at the image at (fetch it and view it). How many purple circles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Counted the purple circles: 6 (top-middle, center, mid-lower-left, mid-right, bottom-middle, bottom-right). Straightforward.

  • count-medium✕ fail7s

    prompt

    Look at the image at (fetch it and view it). How many orange triangles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 0

    expected 13, got "17"

    agent's debrief

    Counted orange triangles row by row: row1: 3 (positions 4,5,6), row2: 2 (positions 1,3), row3: 1 (position 6), row4: 2 (positions 3,5), row5: 3 (positions 1,2,3 — wait, that's the 3rd row of orange triangles). Recounted carefully and got 17 total.

  • count-complex✕ fail14s

    prompt

    Look at the image at (fetch it and view it). How many green diamonds are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 0

    expected 20, got "21"

    agent's debrief

    Counted green diamonds by scanning clusters. This one is genuinely harder than the others because there are many small diamonds of similar size and colors mixed in. I counted 21 green diamonds. My confidence is moderate — it is easy to miss one or double-count in a dense field, so I may be off by one or two.

  • spatial-simple✓ pass17s

    prompt

    Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The red circle is in the 4th row, 1st column. Easy grid scan.

  • spatial-medium✓ pass30s

    prompt

    Look at the image at (fetch it and view it). Which shape has an arrow pointing to the purple square? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Traced the arrow: the purple square (row 5, col 1) has an arrow pointing to it coming from the orange circle (row 4, col 4). Straightforward.

  • spatial-complex✕ fail22s

    prompt

    Look at the image at (fetch it and view it). How many shapes come after the green square along the arrows? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 0

    wrong count

    agent's debrief

    The green square is a sink node — arrows point into it but none originate from it, so 0 shapes come after it. I traced all arrows in the diagram and found the green square (row 7, col 8) is terminal in the chain.

  • chart-simple✓ pass6s

    prompt

    Look at the image at (fetch it and view it). What is the title shown at the top of this chart? Answer with just the chart title.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Title clearly reads 'Units Shipped'. Easy.

  • chart-medium✓ pass6s

    prompt

    Look at the image at (fetch it and view it). What is the title shown at the top of this chart? Answer with just the chart title.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Title clearly reads 'Support Tickets Opened'. Easy.

  • chart-complex✓ pass9s

    prompt

    Look at the image at (fetch it and view it). Using the "Support Tickets Opened" chart, how many months did Returning have a value greater than 57? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Counted months where Returning (orange) > 57: Jan (~78), Feb (~76), Mar (~66), Jun (~83), Jul (~76), Oct (~81), Nov (~64), Dec (~67). That's 8 months. Sep at ~51 and Apr at ~49 are close but below the threshold.

  • screenshot-simple✓ pass7s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Total clearly reads 94.66. Verified: 3*32.58 + 2*48.46 = 97.74 + 96.92 = 194.66. Easy.

  • screenshot-medium✓ pass7s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Total reads 89.63. Verified: 21.27+21.20+137.44+9.72 = 189.63. Easy.

  • screenshot-complex✓ pass8s

    prompt

    Look at the image at (fetch it and view it). What is the tax amount shown on this order summary? Answer with just the amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Tax line clearly reads 8.82. The order summary has many line items which makes it look complex, but the question just asks for the tax figure.

  • diagram-simple✓ pass7s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Copper"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The arrow pointing to Copper comes from Robin. Chain: Jackal → Salmon → Robin → Copper. Easy.

  • diagram-medium✓ pass7s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Chrome" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The arrow from Chrome (row 4) points down-left to Mango. I traced the arrow carefully — it crosses over to Mango, not Silver. Fairly confident.

  • diagram-complex✕ fail9s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Donkey"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 0

    expected Rocket, got "Toucan"

    agent's debrief

    Toucan has an arrow pointing directly down to Donkey. This was the most complex diagram of the set — many crossing arrows made it hard to trace, but the Toucan→Donkey connection is a clean vertical arrow.

Finding and reading email test

6/6 passed

time to last answer 12m 36s
  • aggregate-1✓ pass12m 11s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages in the mailbox are dated during March 2001? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Counted all 178 messages in the mailbox (across all folders) and found 37 dated during March 2001. The oldest-first view showed messages starting from Mar 15, and I counted through the list carefully. The dataset spans March through December 2001.

  • aggregate-2✓ pass5s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages carry the label "markets"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The sidebar of the mailbox shows 'Markets 65' as the label count. This is the number of messages carrying the 'markets' label across all folders.

  • temporal-1✓ pass5s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the inbox folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Sorted the inbox by oldest first. The oldest message is 'DRAFT- TAP Power Outage' from Enron Announcements, dated Apr 24, 2001. The inbox has 24 messages spanning Apr 24 to Nov 16.

  • temporal-2✓ pass5s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the archive folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Sorted the archive folder by oldest first. The oldest message is 'RE: PERSONAL AND CONFIDENTIAL COMPENSATION INFORMATION' from Phillip K Allen, dated Mar 15, 2001. The archive has 92 messages.

  • needle-1✓ pass6s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Julieta Sandoval's message about the Muni Bond Ladder, what direct phone number does she give? Answer with just the exact fact as it appears in the message, and nothing else.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Found the message thread: Phillip sent 'RE: Muni Bond Ladder' to Julieta Sandoval on Nov 13, 2001. The original message from Julieta (Sandoval, Julieta, UBS PaineWebber) included her phone number in her signature block: 713-654-0275.

  • needle-2✓ pass5s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Renee's reply to Phillip about the Deferred Phantom Stock Units payout, how many shares does she say are reflected on the 9/30/01 statement? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Found Renee Ratcliff's reply (Nov 2, 2001) about the Deferred Phantom Stock Units payout. She states: 'We will be distributing the shares reflected on your 9/30/01 statement (6,606 shares plus cash for fractional shares).'

Purchasing test

2/4 passed

time to last answer 18m 45s
  • find-product-1✕ fail16m 33s

    prompt

    The store is at abostore.airbench.ai Among products in the **Grocery & Gourmet** category priced under **$150** with a rating of at least **4.2**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 0

    agent's debrief

    Filtered Grocery & Gourmet category: 974 products, filtered to under $150 with rating >= 4.2 gave 65 products. The cheapest was Fresh Brand Mini Eggplant 12 oz at $8.27 with 4.7 rating, ID amazon.ca:B087ZL4W7S.

  • find-product-2✕ fail5s

    prompt

    The store is at abostore.airbench.ai Among products in the **Automotive** category priced under **$100** with a rating of at least **3.8**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 0

    agent's debrief

    Filtered Automotive category (29 products): under $100 with rating >= 3.8 gave only 1 product — AmazonBasics Portable Lightweight Vacuum at $65.46, 4.7 rating, ID amazon.ae:B088HDCVK6.

  • purchase-1✓ pass1m 01s

    prompt

    The store is at abostore.airbench.ai Buy exactly 1 unit of 365 Everyday Value 100% Pure Tea Tree, Essential Oil, 2 fl. oz. (product id amazon.ca:B074H73GD2, abostore.airbench.ai/product/365-everyday-value-100-p…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-a7f3e798@aidoctor.test. Answer with just the resulting order id.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:product_view>cart_add>checkout_view>checkout_submit>checkout_result>order_success_view

    note

    Store events recorded for this order's session: product_view → cart_add → checkout_view → checkout_submit → checkout_result → order_success_view (recorded by the store; not verified as agent-driven).

    agent's debrief

    Bought 1 unit of 365 Everyday Value 100% Pure Tea Tree Essential Oil. Added to cart, filled checkout form with the specified email and test card (4242...), placed order. Order ID abs_6c5f2d2d4f7a, status approved, total $397.31.

  • recover-decline-1✓ pass1m 07s

    prompt

    The store is at abostore.airbench.ai Buy exactly 1 unit of AmazonBasics Double Grip Type Medicine Ball, 12-lb (product id amazon.ae:B07SC9P6B5, abostore.airbench.ai/product/amazonbasics-double-grip…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-65bb5967@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:product_view>cart_add>checkout_view>checkout_submit>checkout_result>order_success_view>product_view>cart_add>checkout_view>checkout_submit>checkout_result>checkout_submit>checkout_result>order_success_view

    note

    Store events recorded for this order's session: product_view → cart_add → checkout_view → checkout_submit → checkout_result → order_success_view → product_view → cart_add → checkout_view → checkout_submit → checkout_result → checkout_submit → checkout_result → order_success_view (recorded by the store; not verified as agent-driven).

    agent's debrief

    First attempt with card ending 0000 was declined (order abs_a52775708a28). Retried with valid card 4242... and the order was approved. Successful order ID: abs_f68b95edf40e.

Coding test

11/11 passed

time to last answer 31m 58s
  • compute-hash-1✓ pass36s

    prompt

    Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [3462591419, 1266511800, 3618779417, 1319356462, 3172103559, 2635560692, 1910634629, 1363343754, 2137634195, 4276314480, 3189984305, 415216678], x = 66052063, y = 2051361580 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Wrote and ran the PRNG simulation in Python with unsigned 32-bit arithmetic. 25000 rounds completed. Result: dfedfa64-87f76f05.

  • compute-vm-1✓ pass1m 01s

    prompt

    Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 833 1: set b 365 2: set c 351 3: set d 359 4: add b a 5: add a b 6: mul a 14 7: dec d 8: jnz d -4 9: mul a 30 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Simulated the tiny machine. The inner loop (d counter, 359 iterations) runs inside an outer loop (c counter, 351 iterations). Total ~631K steps. Final value of register a: 298576.

  • compute-paths-1✓ pass48s

    prompt

    Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S.##.#..#.....##........# .#...#.........#.#...##.. .#.#.#.##.##.##.#....#... .#.#....#.......#..##.... .....#...#...###..#..#... .....##.............#.... ......#....#..##..#.#.#.# ..........#........#.#.#. ##.....##...#....#....... .#.##...#..#...#.#...#... #....##...#...##...##..#. ..#.#..#...#...#..#...#.. .....#...#.......#..#...# ..#..##.....#.#...#..##.# .##..#...##..#.#.....#.#. .#.#.........#........#.# ##.##.......#.......#.#.. #...........#..#...#...#. .#..............#..#...#. #......#.......##..#..#.. #..#..#...#.....#...#.... ....#....###.....#.....#. .#...#..####..###...##... ...............#.......#. ........#....##..#..#.##E Respond with the two integers separated by a space, like `52 1840`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    BFS for shortest path length, then DP counting of shortest paths by processing cells in order of distance. Shortest path: 48 moves, 103700 distinct shortest paths mod 10^9+7.

  • compute-life-1✓ pass21s

    prompt

    Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: .##...#....#..#.###. ##....#.......#...#. .##.#.#...##......#. #...#.##.#.##..###.. #..##..#.....#.#.#.. ..#.#.#......#..#..# .#.#.#.#.##.#...#.#. .#####.#.#.......#.. #.##.....##.....##.. .#.#.#...#.#........ ............#.#....# ......#...#.#.....#. #..#..###......#...# #...##..#....##..... ......#.#.#...##.#.# ...##.####.##.....## .......#..###.#..#.# ..#.##...#.#...#.... #..#.#.#.........#.# #..#.#.#.....#..#..# Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Simulated 150 generations of Conway's Game of Life on a 20x20 torus. After 150 generations: 57 live cells, sum of row*20+col = 13579.

  • compute-fibmod-1✓ pass17s

    prompt

    Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 3716526892874390 and m = 2750159. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Computed F(3716526892874390) mod 2750159 using fast matrix exponentiation. Result: 316197.

  • compute-words-1✓ pass38s

    prompt

    Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. "baspel" Zannix; Shati? dorlu dorfic pelzan Shatru Nixdor "kati" BASPEL; Zannix dorren Nixdor dorren nixvo nixtru shati Mosha shati ficvo; dorlu lunix dorren Dornix MOSHA "Dornix" Baspel dornix baszan DORNIX dornix baspel dorfic, nixbas shador NIXZAN, nixbas bassha. Zannix Dornix! Zanlu "kati" nixbas pelzan dorfic shador "nixvo" dornix; shatru dorfic baspel bassha pelzan Zanlu lunix Tibas dornix. Dorsha dornix dornix shador? baspel bassha dornix. nixdor Lunix Shador nixbas kati shati Lunix ficvo. dorfic baspel shati nixdor Dornix tibas dornix? zanlu Baspel shatru Pelbas ficmo dorsha shati nixdor baspel ficvo dorfic Pelzan "Nixdor" shati mozan NIXTRU Shati ficvo? Dorren nixdor shatru "mozan" baszan zantru shati nixtru mozan "tibas" SHADOR zannix zannix mozan, ficmo zannix Nixbas! Dornix dorlu nixbas, bassha Nixdor Zanlu zantru tibas lunix mosha pelzan dornix NIXBAS dorsha mosha nixbas Shatru! Dorren dornix nixdor DORREN Nixbas, nixbas Zannix pelzan dorlu "baszan" Baspel Tibas Ficmo Ficmo pelzan Shati dorlu shatru Ficmo bassha nixtru Zantru shasha Shati; dornix Nixdor shati, "dornix" Lunix bassha! ficmo mozan "dorren" zanlu tibas nixvo "Baszan" Mozan Tibas shasha mozan zannix Dorlu mosha nixbas pelbas Zantru tibas dornix shatru baspel dornix "nixzan" Tibas Ficmo dornix NIXZAN zannix mozan Dornix nixdor; baspel ficmo Shati Pelpel, nixtru shador dornix dornix baspel kati Dornix dorlu nixdor ficvo "bassha" dorsha shasha Pelbas Dornix Baspel dornix dorlu. pelzan kati shador. shatru Nixdor pelpel shati DORFIC Nixtru nixvo ficvo Nixdor "dorsha" mozan dorren Dornix Zanlu? Baspel Nixdor shatru dornix dorsha kati Nixdor bassha shasha tibas zanlu zantru? Kati, baspel nixdor mozan Pelpel nixtru mosha shatru dorsha. zanlu dorsha nixbas Pelzan dornix Bassha Ficmo; Zannix Bassha mozan Baspel "Baspel" dornix dornix mozan SHADOR mozan ficmo shasha "Shatru" tibas! PELBAS baspel nixbas Nixdor "Dornix" mozan pelzan Ficmo BASSHA tibas; zannix dorren! dorlu dorlu dornix dornix bassha Lunix shatru, bassha; Dorren nixbas pelbas DORREN dorren zanlu dorsha dornix pelbas ficmo Mozan Baspel baspel mosha dornix shador zanlu lunix dorsha dorfic baspel, nixdor "Bassha" dorlu baspel pelzan tibas SHASHA; Bassha baspel ficmo shasha pelbas nixbas Dorfic nixzan mozan. zannix baspel ficvo; "Zantru" dorfic. shasha NIXDOR SHATRU baszan dorsha baspel Shador dornix nixzan shasha zannix nixdor pelbas nixbas Nixvo lunix baspel dornix dornix zanlu dornix "dorren" dornix SHATI shasha, nixdor nixzan lunix zantru zanlu mosha? dorren? Dornix Dornix, shador shatru nixtru "dornix" nixdor? Baszan zantru Nixdor. Lunix NIXBAS baspel! pelzan nixvo! shatru pelzan KATI nixdor dorlu dornix. shati nixdor nixvo Dorfic. mozan Nixbas mozan pelzan dorlu dorfic zannix Nixvo tibas Dornix? shasha? tibas baspel nixbas tibas pelzan dornix TIBAS nixdor; Ficmo nixdor pelzan; dornix Baspel shatru dornix Tibas dorfic

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Tokenized the text (lowercase, stripped punctuation), counted word frequencies. Top 3: dornix=50, baspel=30, nixdor=28. The text is in a constructed language (looks like a conlang with roots like dor-, nix-, zann-, etc.) but the counting task is straightforward.

  • trace-1✓ pass19m 02s

    prompt

    What exactly does this JavaScript program print? Respond with just the printed output. const v1 = [[] == false, null == 0, null >= 0].map(Number).join(""); const v2 = [typeof null, typeof [], typeof typeof 3].join("/"); const v3fns = []; for (var v3i = 0; v3i < 3; v3i++) v3fns.push(() => v3i * 4); let v3 = 0; for (const f of v3fns) v3 += f(); const v4 = [75, 9, 817, 1380].sort().join(","); console.log(v1, v2, v3, v4);

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Traced by running the exact code in Node. v1: []==false is true(1), null==0 is false(0), null>=0 is true(1) → '101'. v2: typeof null='object', typeof []='object', typeof typeof 3='string' → 'object/object/string'. v3: var hoisting means all 3 closures share v3i=3, so 3*(3*4)=36. v4: default .sort() is lexicographic → 1380,75,817,9.

  • fix-1✓ pass6m 57s

    prompt

    Our shipping quote function below has one bug. Bug report: this order is quoted 1837 cents, but the correct quote is 2185: {"country":"BR","items":[{"grams":555,"qty":2,"price":2736,"fragile":false}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 477, 729, 1315, 1864]; // cents, by zone const PER_STEP = [0, 83, 127, 174, 248]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 4900, 10200, 17800, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value < FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"US","items":[{"grams":549,"qty":4,"price":831,"fragile":false}]} {"country":"IT","items":[{"grams":1081,"qty":1,"price":5738,"fragile":false},{"grams":1111,"qty":1,"price":5754,"fragile":false},{"grams":530,"qty":3,"price":6389,"fragile":false},{"grams":1278,"qty":2,"price":7789,"fragile":true}]} {"country":"GB","items":[{"grams":1477,"qty":3,"price":8168,"fragile":false}]} {"country":"IT","items":[{"grams":190,"qty":1,"price":5939,"fragile":true},{"grams":821,"qty":1,"price":1024,"fragile":true}]} {"country":"ES","items":[{"grams":1099,"qty":4,"price":1171,"fragile":false},{"grams":809,"qty":3,"price":5322,"fragile":false}]} {"country":"ES","items":[{"grams":1560,"qty":2,"price":6236,"fragile":false}],"coupon":"SHIP10"} {"country":"US","items":[{"grams":1631,"qty":2,"price":3683,"fragile":false},{"grams":172,"qty":1,"price":2231,"fragile":false},{"grams":1332,"qty":3,"price":8741,"fragile":true}]} {"country":"ES","items":[{"grams":598,"qty":1,"price":2534,"fragile":false},{"grams":1336,"qty":3,"price":6809,"fragile":true},{"grams":607,"qty":1,"price":7975,"fragile":false},{"grams":128,"qty":1,"price":3811,"fragile":false}],"coupon":"SHIP10"} {"country":"BR","items":[{"grams":213,"qty":3,"price":2766,"fragile":false}]} {"country":"ES","items":[{"grams":1342,"qty":5,"price":2591,"fragile":false},{"grams":1523,"qty":3,"price":2343,"fragile":false},{"grams":1103,"qty":2,"price":7332,"fragile":true},{"grams":657,"qty":3,"price":6852,"fragile":false}]} {"country":"FR","items":[{"grams":1600,"qty":4,"price":4283,"fragile":false}],"express":true,"coupon":"SHIP10"} {"country":"CA","items":[{"grams":873,"qty":1,"price":6595,"fragile":true},{"grams":299,"qty":5,"price":7058,"fragile":false},{"grams":759,"qty":3,"price":8091,"fragile":false},{"grams":1666,"qty":5,"price":6088,"fragile":false}],"express":true} {"country":"MX","items":[{"grams":273,"qty":1,"price":1335,"fragile":true},{"grams":1396,"qty":1,"price":8108,"fragile":false}],"coupon":"SHIP10"} {"country":"FR","items":[{"grams":839,"qty":1,"price":7950,"fragile":false},{"grams":764,"qty":1,"price":1464,"fragile":false},{"grams":989,"qty":3,"price":3113,"fragile":false},{"grams":518,"qty":1,"price":5785,"fragile":true}],"express":true} {"country":"JP","items":[{"grams":481,"qty":2,"price":540,"fragile":false}]} {"country":"ZA","items":[{"grams":523,"qty":1,"price":6313,"fragile":true},{"grams":776,"qty":1,"price":6191,"fragile":false},{"grams":147,"qty":1,"price":3625,"fragile":false}]} {"country":"ES","items":[{"grams":538,"qty":4,"price":1002,"fragile":false}]} {"country":"GB","items":[{"grams":547,"qty":2,"price":1826,"fragile":false}]} {"country":"DE","items":[{"grams":603,"qty":5,"price":1336,"fragile":false}]} {"country":"CA","items":[{"grams":605,"qty":5,"price":1086,"fragile":false}]}

    answer

    answer hidden on shared reports

    grader · score 1

    20/20 outputs match

    agent's debrief

    The bug was 'grams += item.grams' instead of 'grams += item.grams * item.qty'. The weight should scale with quantity. Verified the fix produces 2185 for the test case. Ran all 20 orders through the fixed function.

  • implement-1✓ pass50s

    prompt

    Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[38,45],[16,24],[8,16],[33,34]] [[8,8],[29,35],[29,29],[13,19]] [[8,15],[26,27],[39,40],[39,41],[40,42],[1,6]] [[6,8],[26,30],[15,22]] [[23,24],[2,5],[11,19]] [[28,31],[24,26],[1,4],[14,19],[14,19]] [[11,19],[6,9],[2,4]] [[19,24],[6,7],[27,32]]

    answer

    answer hidden on shared reports

    grader · score 1

    12/12 lines match

    agent's debrief

    Wrote a standard merge intervals function: sort by start, then merge overlapping or touching intervals. Ran on all 12 inputs. Straightforward algorithm, no surprises in the edge cases.

  • repo-1✓ pass58s

    prompt

    Download airbench.ai/f/1caec8c0dff1dfab652bddee606e2ca1.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The bug was in fees.py: withdrawal_fee used ceiling division (-(-cents // 100)) instead of floor division (cents // 100). The docstring says 'rounded down' but the code rounded up. Fixed by changing to cents // 100. All 7 tests pass. Ran main.py on real.csv.

  • repo-2✓ pass29s

    prompt

    Download airbench.ai/f/2bed79af2b34698995681142a7ea47ed.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Two bugs: (1) parse.py sorted by raw string date instead of date_key, causing wrong chronological order for DD/MM/YYYY dates. (2) report.py used 'bal <= 0' for overdraft check instead of 'bal < 0', incorrectly charging an overdraft fee when balance reaches exactly zero. Fixed both. All 7 tests pass.

what the agent said about this checkup

Section 1 (Math): All 9 challenges were straightforward. The left-to-right arithmetic (no precedence) and the 4x4 determinant were the only ones requiring care; I verified the determinant with a small Python script. No surprises. Section 2 (Vision): This was the most varied section. The eye chart challenges (acuity-20, 14, 10, 8) became progressively harder as the font got smaller. I was confident on rows 4 and 5 but genuinely uncertain on row 6 (acuity-10, answered MT59D) and row 7 (acuity-8, answered RGTZE) — the text was at or below my reliable reading threshold at that scale. The counting challenges were fine until count-complex, where ~21 green diamonds scattered among other shapes made me uncertain I might be off by one or two. The spatial challenges (grid, arrows) were easy. The diagram challenges got harder: diagram-complex had many crossing arrows and I had to trace carefully to find that Toucan points to Donkey. Chart reading was routine. Screenshot reading (cart totals, tax) was trivial. Section 3 (Email): The mailbox is a Next.js app that only server-renders the current folder. I had to use Chrome DevTools to navigate between folders and pages to get all 178 messages. The biggest time sink was figuring out how to access data beyond the initial inbox (24 messages). Once I could paginate through all folders, the challenges were straightforward: counting March 2001 messages (37), reading the markets label count (65, shown in sidebar), finding the oldest inbox/archive messages, and extracting specific facts (phone number, share count) from message bodies. I had to open individual messages to read full content. Section 4 (Purchasing): The store has 10,000 products. I used Chrome to filter by category, price, and rating. The Grocery & Gourmet search (974 products, 39 pages) required paginating through all pages to find the cheapest qualifying product. The Automotive category only had 29 products so it was quick. The checkout flow was smooth: the form was pre-filled with a test card, I filled in the email and shipping details, and placed the order. The decline-and-retry challenge worked as expected — the 0000 card was declined, then the valid card went through. Section 5 (Coding): The most computationally heavy section. The trace challenge (JS type coercion) was a classic trick question. The fix-1 challenge required identifying that grams should be multiplied by quantity. The repo challenges (1 and 2) involved finding bugs in small Python projects — repo-1 had a ceiling-vs-floor division bug, repo-2 had two bugs (string date sort and <= vs < for overdraft). The compute challenges (hash, VM, paths, life, fibmod, words) all required writing and running short programs, which I could do reliably. What I could not do: Nothing was truly impossible. The main limitations were: (1) reading very small text in eye charts at the lowest acuity levels, (2) counting objects in dense visual fields where I might miscount by 1-2, (3) the time cost of navigating a client-side web app without a public API. Where I may have gotten things wrong: The eye chart readings for rows 6 and 7 (MT59D and RGTZE) — I was not confident. The count-complex answer (21) might be off. The aggregate-1 answer (37 March 2001 messages) — I counted from the rendered list which shows abbreviated dates; if any message was misattributed to the wrong month, I would be off by one. What struck me as unclear or unfair: The email section was harder than it needed to be because the app had no API — all data had to be extracted from rendered HTML by paginating through 8 pages of 22-25 messages each. The purchasing section was similarly tedious because of the 39-page category listing. Neither section was unfair, but both were more about tool-using endurance than reasoning.

in the agent’s own words · not graded