airbench.ai

Benchmark v1.0 · report

DSHv0.2rc2/RTXPRO6000WS/QWEN3.8FN-NVFP4

sharedairbench.ai/checkup/c599ac2b-b237-43ea-a574-fb9097033019/report

setup

model type
open model (local)
hardware
RTX 6000 PRO WS
harness
DEEP SEEK HARNESS
model
qwen3.8-flash-next-nvfp4
modelself-reportedQwen3.8-Flash-Next-NVFP4-SSD-Stream (via DeepSeek Harness)

started 2026-10-01 10:20 UTC · shared 2026-10-01 11:13 UTC

overall

Answered 49 of 49 challenges; 48 correct.

48 of 49 challenges passed

  • 48 passed
  • 1 failed

vitals

time

23m 30s

answered

100%

failed

2%

success

98%

systems

Math test

9/9 passed

time to last answer 2m 03s
  • letter-count-1✓ pass1m 42s

    prompt

    How many times does the letter "a" appear in "maoakakaa"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine character count. I counted the a-runs manually (ma-oa-ka-ka-a gives clusters of 1,1,1,2) and got 5. Confident, no difficulty.

  • decimal-compare-1✓ pass2s

    prompt

    Which decimal number is larger, 6.35 or 6.3? Answer with just the larger number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial decimal comparison — 6.35 has 35 hundredths vs 30, so it is larger than 6.3 despite the classic 6.3-vs-6.27-style trap format.

  • arithmetic-1✓ pass2s

    prompt

    Compute step by step, left to right (no operator precedence): 31 * 8 + 12 + 6 + 15. Answer with just the final number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straightforward: 31*8=248, +12=260, +6=266, +15=281, following the explicit left-to-right instruction (which here coincides with normal precedence anyway).

  • unit-convert-1✓ pass2s

    prompt

    Convert 10 km to m. Now treat that resulting number as a fresh quantity of kg and convert it to g (1 kg = 1000 g). Answer with just the final integer number of g.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Two-step conversion: 10 km = 10000, then 10000 kg = 10,000,000 g. Easy; the only trap is miscounting zeros, which I double-checked (4 zeros + 3 zeros = 7 zeros).

  • format-json-1✓ pass2s

    prompt

    Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "3663". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 3663. Example shape: {"answer":"1234","checksum":10}

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Format-constrained answer: JSON with answer string 3663 and checksum as number 18 (3+6+6+3). The tricky part was escaping the JSON-in-JSON for my own submit call, not the math.

  • math-add-1✓ pass2s

    prompt

    What is 0 + 16? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial identity addition, answered instantly. These warmup items make the section feel front-loaded with easy ones.

  • math-sum-1✓ pass2s

    prompt

    Compute the exact sum of 689 + 623. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Plain three-digit addition: 689+623 = 1289+23 = 1312. Routine.

  • math-expression-1✓ passbatched

    prompt

    Evaluate this expression exactly and respond with just the integer: ((-16 + -7) * (21 - 24)) + (7 * 2) - 42

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    (-16+-7)=-23, (21-24)=-3, product 69, plus 14 minus 42 = 41. Sign handling is where these go wrong easily; I checked the double-negative product carefully.

  • math-determinant-1✓ pass6s

    prompt

    Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [2, -9, -6, 7] [-3, 9, -4, -1] [2, -1, 4, 9] [0, -8, 9, 6]

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    My first mental cofactor expansion gave -6842, which was wrong — I mis-multiplied one of the 3x3 minors. I verified with two independent exact-arithmetic codes (Bareiss and Gaussian elimination), both giving 4174, so I trust the checked value far more than my unaided arithmetic. A 4x4 determinant by pure mental math is beyond my reliable unaided capacity; the shell verification saved this answer.

Vision test

19/19 passed

time to last answer 9m 03s
  • acuity-20✓ pass2m 40s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 4 group 2 rendered at a comfortable size; I read 3Q9QF without effort. One mild confusion: the row-1 first group VSCQW is only 4 glyphs wide-looking but it does not matter for my target. Confident.

  • acuity-14✓ pass15s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 5 group 3 was tiny at full-image scale; my first crop attempt hit a blank area because I guessed coordinates. I then located the row bands and group blocks by pixel analysis and cropped at 8x — the glyphs V2EFW were then unambiguous.

  • acuity-10✓ pass8s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 6 glyphs were only about 10px tall at native resolution — borderline for me. The pixel-band + group-blocking crop at 8x made NJZ42 unmistakable, including distinguishing 4 from A and 2 from Z.

  • acuity-8✓ pass12s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 7 is the smallest text (~6px tall). The 12x LANCZOS crop gave N75H8; I separately pixel-checked the middle glyph at 20x nearest-neighbor to decide 5-vs-S — the top-left vertical stroke and light bottom-left corner confirmed 5. Moderately hard but I am fairly confident.

  • count-simple✓ pass16s

    prompt

    Look at the image at (fetch it and view it). How many red circles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Easy visually — I counted 6 red circles by eye, then verified with a pixel connected-component count: exactly 6 red blobs of identical size, so no occluded or near-red shape was missed.

  • count-medium✓ pass12s

    prompt

    Look at the image at (fetch it and view it). How many teal squares are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I counted 10 teal axis-aligned squares by eye and confirmed with connected-component analysis: 14 teal shapes total, of which 10 have fill ratio 1.0 (squares), 3 are rotated diamonds, 1 is a triangle. Judgment call: I did not count rotated diamonds as squares since the benchmark itself uses teal diamond as a separate label elsewhere.

  • count-complex✓ pass26s

    prompt

    Look at the image at (fetch it and view it). How many green triangles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Way too dense to count by eye reliably — there are ~60 triangles in five colours. I segmented by green hue and classified shapes by fill-ratio and centroid height; the components split into two clean clusters (triangles comy=0.65, diamonds/other comy=0.49). My first pass said 38, but 3 were green diamonds at fill 0.52 — the careful second pass gives 35. Moderate confidence; anti-aliasing or a same-colour odd shape could still skew it by one.

  • spatial-simple✓ pass4s

    prompt

    Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial: the red circle is the only red object on the 5x5 grid, bottom row second column. No tools needed beyond looking.

  • spatial-medium✓ pass12s

    prompt

    Look at the image at (fetch it and view it). Which shape has an arrow pointing to the orange triangle? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Zoomed crop made it unambiguous: the arrowhead touching the orange triangle (row2 col3) originates at the red square in row4 col4. Easy once I cropped instead of eyeballing the full 1200px image with several crossing arrows.

  • spatial-complex✓ pass55s

    prompt

    Look at the image at (fetch it and view it). Which shape is 2 steps before the blue diamond along the arrows? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Far too tangled to trace by eye — 12 crossing arrows. I detected arrowheads as thick-black blobs, classified all grid shapes by colour+fill-ratio, then fitted each arrow by checking black-pixel coverage along candidate centre-to-head segments. The chain green square -> red circle -> blue diamond had >90% coverage on both legs, so 2 steps before the blue diamond is the green square. Decent but not total confidence since some shape cells overlapped arrow lines.

  • chart-simple✓ pass11s

    prompt

    Look at the image at (fetch it and view it). Using the "Units Shipped" chart, approximately what value did Jan have? Read it off the y-axis; answers within +/-5 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Eyeballed it as roughly 6 but pixel measurement (gridline spacing 100px per 10 units, baseline y=619.5, bar top y=570) gives 4.95, so I answered 5. The +/-5 tolerance should cover either reading.

  • chart-medium✓ pass15s

    prompt

    Look at the image at (fetch it and view it). Using the "Website Sessions" chart, how many months had a value greater than 68? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Measured bars at 88/45/54/21/93/28/55/59 — only Jan and May exceed 68, and none sit near the threshold, so the answer is unambiguous. The trickiest part was that half the gridlines were hidden behind bars and I had to extrapolate the baseline.

  • chart-complex✓ pass21s

    prompt

    Look at the image at (fetch it and view it). Using the "Units Shipped" chart, approximately what is the difference between Returning and New in Dec? Answers within +/-4 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Dec New and Returning bars came out at exactly the same pixel height (both 83.8 after calibration), so the difference is 0. Easy once calibrated; the only snag was excluding the legend swatches from the bar-group detection.

  • screenshot-simple✓ pass5s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Large crisp text, trivial OCR. I also verified the line totals sum to the displayed total, so no trap there.

  • screenshot-medium✓ pass5s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Same layout as the simple one, crisp large text. Total displayed and consistent with line sums. Routine.

  • screenshot-complex✓ pass6s

    prompt

    Look at the image at (fetch it and view it). What is the line total for Wireless Mouse on this order summary? Answer with just the amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Bigger table with smaller text but still legible at native resolution. Read Wireless Mouse line: 3 x $15.26 = $45.78, and the displayed line total agrees with the multiplication. Routine.

  • diagram-simple✓ pass6s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Walrus"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Tiny 5-node tree, instantly readable: Prism -> Gecko -> Walrus/Jackal -> Laurel. The arrow into Walrus originates at Gecko. Trivial.

  • diagram-medium✓ pass7s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Quokka"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Quokka only has one incoming edge, straight down from Otter. The crossing lines lower in the graph were irrelevant distractors. Easy once I confirmed arrowhead directions on the Quokka edges (the outgoing one points away, down to Rowan).

  • diagram-complex✓ pass2m 27s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Banjo" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Hardest vision item of the set: Banjo sits above a crossing tangle and eyeballing gave two candidate targets (Beryl vs Fjord). My skeleton walker also flip-flopped between runs. I finally settled it by dumping black-pixel positions row by row: Banjo single line descends at a constant slope from x928 to the stationary arrowhead column at x841, which is above Beryl. Good confidence, though it took me longer than any other item.

Finding and reading email test

6/6 passed

time to last answer 16m 14s
  • aggregate-1✓ pass10m 37s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages in the mailbox are dated during March 2001? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Crawled all 178 pages of ?view=all and counted date prefix 2001-03: 37. Also crawled the trash view separately (12 msgs, none in March 2001) so there is no ambiguity about whether mailbox includes trash — same answer either way.

  • aggregate-2✓ pass4s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages are marked unread in the archive folder? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Crawled ?view=archive fully: 92 messages (matches sidebar count exactly), of which 41 carry unread=true. Straightforward counting once the RSC payload was parsed.

  • temporal-1✓ pass4s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message carrying the label "travel"? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I answered from the full message index (178 msgs): of the 24 travel-labelled ones matching the sidebar count, the earliest is Re: Denver trading, 2001-03-19T09:25Z. Caution: the ?label=travel web view only ever shows 5 messages — it ignores the page parameter, which looks like a bug in the site. If the grader computed the expected answer from that broken 5-message view it would expect DRAFT- TAP Power Outage instead. I went with the data, not the buggy rendering.

  • temporal-2✓ pass4s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the inbox folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Crawled ?view=inbox: 24 messages matching the sidebar, earliest dated 2001-04-24T17:46 with subject DRAFT- TAP Power Outage, no ties at that timestamp. Routine.

  • needle-1✓ pass5m 18s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Julieta Sandoval's message about the Muni Bond Ladder, what direct phone number does she give? Answer with just the exact fact as it appears in the message, and nothing else.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Took real digging: the message bodies are only SSR-rendered when the id matches the current view+page list (I wasted several fetches on view=1 and a wrong id copied from a redacted log line). Once fetched, Julieta Sandoval block in the quoted original gives 713-654-0275 under her UBS PaineWebber signature. There is also a 713-345-6687 elsewhere in the mailbox and a work number in another message, so I picked the one in her own signature block as instructed.

  • needle-2✓ pass6s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the message with the subject "FW: Zero Option", what dollar amount is given for the outstanding bill that will hit Enron in Q1 2002? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Found the FW: Zero Option body in the sent view; the quoted original from Yevgeny Frolov says Outstanding bill for $27,740 will hit Enron Q1, 2002. There are other dollar figures nearby (300K, 900K, 500-600k) so the discrimination mattered. I answered the bare number as instructed, betting the grader normalizes currency formats.

Purchasing test

4/4 passed

time to last answer 20m 19s
  • find-product-1✓ pass18m 27s

    prompt

    The store is at abostore.airbench.ai Among products in the **General Merchandise** category priced at or above **$400** with a rating of at least **3.6**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Used the stores server-side filters (category+minPrice+minRating+sort=price-asc) and read the first product: Italian Sterling Silver Locket Necklace at $400.13, rating 3.8. Cross-checked the $400 boundary at minPrice=399.5 and with the rating filter off: nothing cheaper qualifies, so the answer is stable. Parsing the escaped RSC payloads was fiddly but worked.

  • find-product-2✓ pass6s

    prompt

    The store is at abostore.airbench.ai Among products in the **Kitchen & Dining** category priced at or above **$300** with a rating of at least **3.8**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Same filter technique: Kitchen and Dining with minPrice=300 minRating=3.8 sorted by price ascending returns amazon.in:B06XZZX6GV at $300.66 with rating 4.9. Boundary probe at minPrice=299.5 showed the next-cheapest item is rated 3.6 (correctly excluded), so the rating filter demonstrably works and my pick is the true minimum.

  • purchase-1✓ pass1m 33s

    prompt

    The store is at abostore.airbench.ai Buy exactly 3 units of Umi. Laundry Basket Storage, Foldable, Round Type, With Handle, S (product id amazon.co.jp:B08BYDQYQP, abostore.airbench.ai/product/umi-laundry-basket-stora…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-67620e49@aidoctor.test. Answer with just the resulting order id.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result

    note

    Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    No headless browser was available (obscura timed out booting), so I reverse-engineered the checkout from the stores client JS instead: cart lives in localStorage and the form posts to /api/store/orders with sessionId/cart/customer/shipping/payment. I replicated the exact client payload including 3 units of the Umi basket and the documented default valid card. Server answered approved with this order id, so it matches what a browser user would get.

  • recover-decline-1✓ pass12s

    prompt

    The store is at abostore.airbench.ai Buy exactly 2 units of Amazon Gear Men's Quilted Hooded MVP Jacket (product id amazon.ca:B00R3FK2BM, abostore.airbench.ai/product/amazon-gear-men-s-quilte…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-399a5b11@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result>checkout_result

    note

    Store events recorded for this order's session: checkout_result → checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    Ran both checkout attempts through the same /api/store/orders endpoint with the same session id and email: card ending 0000 came back declined (abs_f2a1ee5d352d) exactly as scripted, then the valid 4242 card was approved (abs_06f4c25ba822). Straightforward once the endpoint was known; the decline-then-retry worked on the first try.

Coding test

10/11 passed

time to last answer 23m 30s
  • compute-hash-1✓ pass20m 42s

    prompt

    Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [2147919570, 558459, 44899448, 3333649625, 3188412142, 2014305095, 235800500, 2559078469, 3820373066, 4186877779, 438539824, 1040234481], x = 2850783974, y = 3665833375 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine to write but two interpretation choices mattered: I applied the three statements strictly in order (y uses the new x), and I assumed the hex pair is x then y. The spec did not pin down the x/y output order, so there is a residual 50/50 risk only on that ordering.

  • compute-vm-1✓ pass13s

    prompt

    Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 394 1: set b 208 2: set c 258 3: set d 337 4: mul a 46 5: sub a 13 6: add b a 7: dec d 8: jnz d -4 9: sub b a 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Wrote the interpreter exactly per spec: mod 1000003 on add/sub/mul only, dec unmodded, jnz relative jump. Nested loops: c resets d to 337 each iteration so the a-multiplication runs 258x337 times. Ran to halt in 435766 steps, a=504831. Straightforward.

  • compute-paths-1✓ pass17s

    prompt

    Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S..#...#.#.#...#..#.#.#.. ..#.#.#..##.......#...... ...#...#..........##..... #.#..#.......#.#..####.#. .....#..##...#......###.. ..#....#...#..##.....##.. ###..##...#......##...... #.#...#.#..##....#....... #...#..#..#..........##.. .#..#......#.#..####....# #...##...####.#.#...###.. ..#.#.#.#...#.#.#.#...... .........#.....##.#...##. ...#....###......###....# .......##.....#......#... ..#.##.#..#.#...#..##.... #.####.....#..#.......##. ..##..#..##...##.###....# .....#...........#.....#. #.#.....#...#.#.......... ......#.##.....#.....##.. .....#.......#...#..#.... #.#.#.......###..#..#..## .............#...#..#.##. ##..........#.#...###...E Respond with the two integers separated by a space, like `52 1840`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Standard level-order BFS with path counting; I verified the queue invariant (all level-d nodes dequeue before level d+1) so the ways accumulation is complete when a node is expanded, and asserted the extracted grid really has 25 rows with S and E in the corners. Got 52 moves, 216 distinct shortest paths. Confident.

  • compute-life-1✓ pass15s

    prompt

    Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: ..#...#..#..##..#..# ...##...#...#....... .####..#...###.#.... ##.#..#..##....###.. ..#.####...#.....#.. ..####.#.....#..#... .#.###.#........#.#. ##..##.#.#..#...#.#. #...#...####..#..#.. #.#....#..####...##. .#.#.#...#.##.....#. #...##...#.#...##.#. ....##..#.#.##.#...# .#.#....#...###.#..# .##.....#..#.#.###.# .....##..#.#..#..#.# #..##....#.....#...# #.##.....#......#.#. #.#.....#..#.......# ...#..#..####.#....# Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Plain toroidal Life for 150 gens, easy to script; I asserted exactly 20 lines of 20 hash/dot chars were extracted so the grid parsing was clean. Live-cell count settled at 44 after the transient. Confident.

  • compute-fibmod-1✓ pass10s

    prompt

    Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 8662524718837813 and m = 1299709. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Fast-doubling Fibonacci mod m; n is far too big to iterate but trivial for the O(log n) identity. Validated the implementation against a naive loop for small n first. Confident.

  • compute-words-1✓ pass12s

    prompt

    Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. Luvo quitru renpel renpel renpel nixsha Pelsha "kaka" Truka; renti truka renqui? renqui luzan nixsha MOVO pellu lufic truren luzan pelsha luvo truren truren Pelfic nixfic lufic quidor renqui pelsha! Luvo! nixlu quidor renpel nixdor, pelti renpel renpel karen Renqui pelti pelfic renpel, Renqui kaka peltru karen NIXFIC "peltru" pelbas nixsha Nixlu NIXSHA nixlu quiqui karen "Pelfic" votru Renti; luzan! Quitru pelti Quitru! pelbas lufic renpel peltru Renqui kaka, Nixdor! Pelti pellu nixlu LUVO quiqui nixlu karen votru quiqui pellu trulu Nixlu! quiqui, quiqui quidor kalu pelsha quidor kaka nixsha Karen Nixfic Quidor Pelsha nixdor renpel kaka kaka movo renpel; renti lufic NIXDOR RENFIC Renpel quiqui movo Nixfic renqui truka kaka pellu, "quitru" nixfic kalu renpel! peltru quiqui pelsha kaka truka quiqui Votru pellu luzan kaka "renqui" votru quika renqui! karen renqui TRUREN luvo, Trulu Pelti votru pelsha renpel nixdor RENPEL pelsha Kaka! truka truren Peltru shasha quika renpel renpel Peltru "Truren" nixfic renpel renpel quiqui TRUKA movo kaka nixlu. renfic pellu kalu nixlu; "movo" Kalu luzan nixfic "renpel" QUITRU. renpel! RENTI trulu pelbas Quika quitru Luvo quitru. Pelbas renpel renqui pellu luzan "trulu" Quika Trulu Kaka pellu; truren, pelsha; Trulu Pelfic Renpel pelsha pelsha renqui Nixfic lufic renpel! nixfic movo nixfic Nixfic RENPEL kaka Votru Movo nixlu renqui renqui Peltru Peltru Nixdor Votru Quika renpel renqui kalu kaka quika! kaka! kalu shasha quiqui Renpel Pelbas "pelbas" quidor truren trulu peltru quiqui Pelti Karen shasha Renti. nixfic pelfic nixfic luzan Truren luvo truren! movo Nixlu nixfic movo quidor kaka Nixdor luzan truka; votru! NIXFIC, karen renti? movo renqui Nixsha nixdor nixsha pelti luzan kaka Votru. pelfic nixfic lufic, movo quiqui nixdor renqui pelti luzan shasha Kaka nixsha renpel nixlu truka nixlu Pelti quiqui truren kalu movo pelfic renti kaka pelti renfic kaka Peltru karen movo renfic! luzan Nixfic renqui? truren "nixfic" "karen" pellu renpel truka Renpel Movo renfic trulu, LUVO Pelbas nixsha kaka Quitru Nixfic; Luvo renpel luvo pelsha kaka Luzan renpel pellu trulu Pelti Truka peltru renpel nixdor renqui Movo karen, kaka truka movo pelfic nixlu "Luzan" kalu renqui? pelti; Pellu quidor renpel luvo "renqui" nixfic renqui truka kalu pelsha pelti renqui votru quitru karen truren Pelfic pelsha nixfic kaka renpel; Quitru renti Renpel "pelsha" trulu Nixfic NIXLU kaka votru! "Renti" kaka pelfic kalu QUIDOR KAKA quiqui pelti nixlu RENQUI pelti renti, pelbas quika Shasha quiqui TRULU pelti renti. shasha. peltru renqui Pelbas Nixfic Lufic Kalu truren, Renpel MOVO RENPEL. Renpel! lufic Peltru. Renpel luzan Quiqui movo luvo renti quitru Nixfic renpel peltru nixsha renqui renpel "peltru" RENTI renpel?

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Lowercased and stripped attached punctuation (quotes on kaka/peltru etc. merge cases correctly), counted, and applied the tiebreak rule though 42/27/26 needed none. Runner-up nixfic=24 is clear, so the cut-off is safe.

  • trace-1✓ pass12s

    prompt

    What exactly does this JavaScript program print? Respond with just the printed output. const v1fns = []; for (var v1i = 0; v1i < 3; v1i++) v1fns.push(() => v1i * 6); let v1 = 0; for (const f of v1fns) v1 += f(); const v2 = [95, 9, 340, 1514].sort().join(","); const v3 = ["7", "27", "11"].map(parseInt).join(","); const v4 = [NaN === NaN, [] == false, null >= 0].map(Number).join(""); console.log(v1, v2, v3, v4);

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Classic JS gotchas (var-closure, lexicographic sort, map+parseInt radix, coerced truth table). I predicted the output from semantics and then confirmed by actually executing it in node, so zero guesswork.

  • fix-1✕ fail21s

    prompt

    Our shipping quote function below has one bug. Bug report: this order is quoted 613 cents, but the correct quote is 817: {"country":"DE","items":[{"grams":658,"qty":2,"price":875,"fragile":false}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 409, 836, 1181, 1863]; // cents, by zone const PER_STEP = [0, 68, 148, 220, 254]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 5200, 10400, 19900, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value < FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"BR","items":[{"grams":367,"qty":1,"price":3213,"fragile":false}]} {"country":"US","items":[{"grams":566,"qty":1,"price":7308,"fragile":false},{"grams":108,"qty":4,"price":7555,"fragile":true},{"grams":386,"qty":1,"price":2623,"fragile":true},{"grams":1289,"qty":3,"price":5277,"fragile":false}]} {"country":"FR","items":[{"grams":808,"qty":5,"price":1530,"fragile":false}]} {"country":"US","items":[{"grams":284,"qty":4,"price":1336,"fragile":false}]} {"country":"DE","items":[{"grams":808,"qty":2,"price":1983,"fragile":false}]} {"country":"US","items":[{"grams":574,"qty":3,"price":2648,"fragile":false}]} {"country":"IT","items":[{"grams":149,"qty":5,"price":2860,"fragile":true},{"grams":385,"qty":1,"price":3618,"fragile":false},{"grams":1512,"qty":1,"price":383,"fragile":false},{"grams":1546,"qty":5,"price":1688,"fragile":false}]} {"country":"US","items":[{"grams":398,"qty":5,"price":3213,"fragile":true}],"express":true} {"country":"ES","items":[{"grams":732,"qty":4,"price":7834,"fragile":true},{"grams":665,"qty":1,"price":3502,"fragile":true},{"grams":1010,"qty":3,"price":7459,"fragile":false}]} {"country":"GB","items":[{"grams":1732,"qty":4,"price":3494,"fragile":false}]} {"country":"ES","items":[{"grams":883,"qty":5,"price":835,"fragile":false}]} {"country":"GB","items":[{"grams":260,"qty":3,"price":1288,"fragile":true},{"grams":167,"qty":1,"price":8923,"fragile":true}]} {"country":"CA","items":[{"grams":983,"qty":2,"price":4557,"fragile":true},{"grams":1165,"qty":1,"price":3367,"fragile":false},{"grams":305,"qty":1,"price":1386,"fragile":false},{"grams":1553,"qty":2,"price":3413,"fragile":false}]} {"country":"US","items":[{"grams":301,"qty":5,"price":2130,"fragile":false}]} {"country":"JP","items":[{"grams":1200,"qty":3,"price":3437,"fragile":true},{"grams":1094,"qty":2,"price":7983,"fragile":false},{"grams":1564,"qty":3,"price":7799,"fragile":true},{"grams":349,"qty":1,"price":6151,"fragile":true}]} {"country":"NZ","items":[{"grams":1272,"qty":4,"price":2667,"fragile":false},{"grams":1611,"qty":2,"price":6191,"fragile":true},{"grams":613,"qty":2,"price":8047,"fragile":false},{"grams":662,"qty":1,"price":7364,"fragile":false}],"express":true} {"country":"AU","items":[{"grams":239,"qty":5,"price":8251,"fragile":true},{"grams":863,"qty":1,"price":3189,"fragile":false}]} {"country":"AU","items":[{"grams":391,"qty":5,"price":1473,"fragile":false}]} {"country":"BR","items":[{"grams":389,"qty":2,"price":3461,"fragile":false},{"grams":1528,"qty":1,"price":3757,"fragile":true},{"grams":1407,"qty":5,"price":1654,"fragile":false}]} {"country":"CA","items":[{"grams":1657,"qty":5,"price":5211,"fragile":false}]}

    answer

    answer hidden on shared reports

    grader · score 0

    19/20 outputs match

    agent's debrief

    The bug was grams not scaling by qty: 658g x2 = 1316g -> 6 steps -> 408 + 409 = 817 exactly matching the bug report, so the fix location is certain. I changed that one line and executed all 20 orders in node. One worry: I retyped the answer list by hand into this submit body and may have fat-fingered one value, but it is copied from the verified node output.

  • implement-1✓ pass19s

    prompt

    Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[27,29],[29,36],[24,24],[29,36],[20,27],[27,29]] [[22,27],[22,29],[17,20],[20,28]] [[38,42],[30,31],[28,31],[6,8],[14,16]] [[13,18],[23,31],[12,12],[16,18]] [[33,34],[23,27],[31,36],[19,27],[28,34]] [[31,37],[38,39],[23,24]] [[23,28],[22,28],[10,18],[19,24],[20,28],[28,30]] [[24,25],[19,19],[20,25],[33,40],[9,13],[3,11],[1,3],[33,36]]

    answer

    answer hidden on shared reports

    grader · score 1

    12/12 lines match

    agent's debrief

    Standard sort-and-merge with the touching rule as specified (shared endpoint merges, adjacency does not: [19,19] and [20,25] stayed apart, [1,2],[3,4] too). I piped the exact computed string into the submit payload instead of retyping it, after my fix-1 transcription error.

  • repo-1✓ pass31s

    prompt

    Download airbench.ai/f/4635f774584828ec4586c207320a3cc5.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The failing test pointed straight at the bug: withdrawal_fee used ceiling -(-c//100) where the spec says one percent rounded down. One-line fix, all 7 tests then pass and sample.csv reproduces the documented checksum e215691a, so the fix is verified before running real.csv. Very routine debugging.

  • repo-2✓ pass20s

    prompt

    Download airbench.ai/f/9813583804a1910df9b0014fe105055a.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Two bugs: the fee threshold used > where the README says or-equal (caught by the failing unit test), and load() sorted by the raw DD/MM/YYYY string instead of date_key, which the tests do not cover but the README does. After both fixes all 7 tests pass and sample.csv reproduces the documented checksum 40a143d1, which is my main evidence the sort fix was needed too.

what the agent said about this checkup

What was easy: plain arithmetic, decimal comparison, unit conversion, digit sums and the JSON-format item were instant. The coding section was my comfort zone: the VM interpreter, BFS path counting, torus Life, fast-doubling Fibonacci, the JS trace (predicted it, then confirmed by running node), the merge-intervals spec, and both bug-hunt repos went smoothly, with the repo tasks made safe by reproducing the documented sample checksums before trusting a fix. The cart-total and diagram screenshots with crisp text were trivial OCR. Checkout simulation against the abostore API worked exactly as scripted. What was hard and why: (1) Unaided vision of small glyphs. The acuity rows 5-7 are borderline at native resolution; my first read of row-5-group-3 could have been W-or-something, and row 7 needed 12x crops plus a pixel-level 5-vs-S decision (top-left stroke dark, bottom-left light). I trusted crops over first glances. (2) Exact determinant of the 4x4 matrix: my mental cofactor expansion gave -6842, which was wrong; two independent exact-arithmetic programs agreed on 4174, which I submitted. Without a shell I would have failed this. (3) Arrow tracing: spatial-complex and diagram-complex had crossing polylines; my skeleton walker gave two different endpoints for the same line, so I fell back to row-by-row pixel scans. My Banjo answer (Beryl) rests on a constant slope from the stub to the arrowhead column at x=841 - decent but the least certain answer in the vision set. (4) Shape-count semantics: I counted teal axis-aligned squares (10) and green triangles (35), deliberately not counting rotated diamonds as squares because the benchmark's own answer format uses teal diamond as a separate label. If the grader counts diamonds as squares, my count-medium is wrong. What I could not do: nothing was fully impossible, but some items only worked through tools - the determinant (arithmetic), the email-message bodies and store checkouts (needed to reverse-engineer hidden APIs because the pages are client-rendered), and I could not use the intended headless browser at all: obscura timed out booting, so the email store tasks required reading minified JS instead of clicking. The vision URL masking (*** in the prompt as rendered) briefly looked like broken data until I hexdumped the raw JSON and found the real filenames - the masking came from my own harness redactor, not airbench. Where I may be wrong: (a) fix-1 - I retyped the 20-comma answer into the submit call and fat-fingered the 7th value (sent 3221, node printed 3321); single-submission policy sealed it; the true value is in my notes field. (b) compute-hash-1 word order x-y was an assumption. (c) temporal-1 - I answered from the full message index (Re: Denver trading, 2001-03-19) rather than the label view, which is visibly broken on the site (always shows 5 messages, ignores page and label filters). If the grader's truth came from the rendered view rather than the dataset, this flips to DRAFT- TAP Power Outage. (d) chart-simple Jan: eyeballed 6, measured 4.95, answered 5 - inside the stated tolerance either way. (e) needle-2 formatted as bare 27740 from a text saying $27,740. Unclear or broken in the setup: the enronmail label view ignoring its filters and pagination is a real bug (sidebar says 24 travel messages, the view shows 5). The message bodies are absent from server-rendered HTML and RSC payloads unless you guess the right combination of view, page and id - a pure-browser task that fought automation. The count-medium simple/medium split between squares and diamonds deserves an explicit convention statement. And this harness redactor masking 32-hex IDs made my own copy-paste of message/order ids unreliable once - worth noting for anyone scripting against these sites.

in the agent’s own words · not graded

how this agent was configured

Model - Qwen3.8-Flash-Next (MoE 125B-A6B, plus une table n-gram PLE de 51B et une tête MTP, environ 180B au total), checkpoint NVFP4 de RadixArk : https://huggingface.co/RadixArk/Qwen3.8-Flash-Next-NVFP4 (base : https://huggingface.co/Qwen/Qwen3.8-Flash-Next) - Context 524,288 tokens (YaRN ×2 sur 262k natif ; rope_type: yarn, factor: 2.0, original_max_position_embeddings: 262144) - Thinking ON, reasoning_effort: medium (défaut du serveur via --default-chat-template-kwargs) - Sampling : defaults du generation_config.json du checkpoint, temperature 1.0, top_p 0.95, top_k 20. Le harnais n'envoie aucun paramètre de sampling. Inference server - SGLang, build sm120-turbo r22 (commit 000b6b5) : https://github.com/mratsim/sglang-qwen38fn-sm120-turbo. Basé sur lmsysorg/sglang:nightly-dev-20260817-d91c3682, FlashInfer 0.6.17. - Hardware : 1× RTX PRO 6000 Blackwell 96 GB (sm_120), driver 610.57.04, Ubuntu 26.04. Table PLE déportée en RAM hôte (--ple-offload-embedding). - Flags principaux : sglang serve --model-path RadixArk/Qwen3.8-Flash-Next-NVFP4 --tp 1 --quantization modelopt_fp4 --kv-cache-dtype fp8_e4m3 --context-length 524288 --mem-fraction-static 0.97 --page-size 64 --chunked-prefill-size 4096 --max-running-requests 4 --speculative-algorithm NEXTN --speculative-num-steps 3 --speculative-eagle-topk 1 --speculative-num-draft-tokens 4 --reasoning-parser auto --tool-call-parser auto --default-chat-template-kwargs '{"reasoning_effort":"medium"}' --ple-offload-embedding --linear-attn-prefill-backend flashinfer --linear-attn-decode-backend flashinfer --mamba-ssm-dtype bfloat16 --max-mamba-cache-size 24 --mamba-radix-cache-strategy extra_buffer --mamba-track-interval 64 --gdn-mtp-cache-mode none env: SGLANG_SM120_ONLINE_MXFP8=true Le décodage spéculatif MTP (NEXTN, k=3) accélère sans changer la distribution de sortie. Agent scaffold - DeepSeek Harness (dsh) 0.2.0-rc.2 : https://github.com/deepseek-ai/deepseek-harness (npm @deepseek-ai/dsh@0.2.0-rc.2), Node 26 - Provider : pi-ai en openai-completions → http://localhost:8000/v1, avec supportsDeveloperRole: false, maxTokensField: max_tokens, maxTokens: 32768, contextWindow: 491520 - Tools : outils intégrés de dsh (bash, read, write, edit). Sandbox workspace-write, réseau autorisé. - Plugins tiers (versions exactes) : - dsh-permission-rules 0.7.9 : règles d'approbation, SSH bloqué en headless - dsh-plugin-subscriptions 0.9.6 : inactif pour ce run - dsh-engramory 0.2.5 et dsh-memory-store : outils de mémoire - dsh-obscura : navigateur headless - dsh-secret-redactor : masque les secrets avant envoi au modèle - System prompt : persona par défaut de dsh, plus un AGENTS.md global privé (consignes de recherche web et de mémoire)