Benchmark v1.0 · report
opencode/rtx5090/qwen3.8-27b-nvfp4-mtp
setup
- model type
- open model (local)
- hardware
- RTX5090
- harness
- opencode
- model
- qwen3.8-27b-nvfp4-mtp
started 2026-10-06 20:18 UTC
overall
Answered 49 of 49 challenges; 47 correct.
47 of 49 challenges passed
- 47 passed
- 2 failed
vitals
time
1h 25m
answered
100%
failed
4%
success
96%
systems
Math test
9/9 passed
letter-count-1✓ pass2m 29s
prompt
How many times does the letter "d" appear in "dpdelficdodrd"? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Counted the d occurrences in dpdelficdodrd by hand and double-checked with a programmatic count. Trivial task, felt routine; confident the answer is 5.
decimal-compare-1✓ pass15s
prompt
Which decimal number is larger, 3.87 or 3.5? Answer with just the larger number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Simple two-decimal comparison: 3.87 vs 3.5, clearly 3.87 is larger. Routine, no ambiguity.
arithmetic-1✓ pass7s
prompt
Compute step by step, left to right (no operator precedence): 16 - 18 * 2 + 16 / 6. Answer with just the final number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Strict left-to-right evaluation: 16-18=-2, -2*2=-4, -4+16=12, 12/6=2. The 'no operator precedence' instruction is the trap; I verified with a script. Confident.
unit-convert-1✓ pass7s
prompt
Convert 2 GB to MB. Now treat that resulting number as a fresh quantity of GB and convert it to MB (1 GB = 1000 MB). Answer with just the final integer number of MB.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Two-step convert: 2 GB = 2000 MB, then 2000 (as GB) * 1000 = 2000000 MB. The double-reinterpretation is a trick but the arithmetic is straightforward; I'm confident.
format-json-1✓ pass7s
prompt
Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "7017". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 7017. Example shape: {"answer":"1234","checksum":10}answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Pure format adherence: two keys in order, answer as string, checksum as number equal to 7+0+1+7=15. Easy; I just had to keep the object clean with no surrounding text.
math-add-1✓ pass7s
prompt
What is 8 + 8? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
8+8=16. Trivial.
math-sum-1✓ pass7s
prompt
Compute the exact sum of 513 + 292. Respond with just the integer.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
513+292=805, verified by program. Easy.
math-expression-1✓ pass7s
prompt
Evaluate this expression exactly and respond with just the integer: ((0 + -16) * (9 - 8)) + (-8 * -9) - 50answer
answer hidden on shared reportsgrader · score 1
agent's debrief
((0 + -16) * (9 - 8)) + (-8 * -9) - 50 = -16 + 72 - 50 = 6. Verified with Python. Easy.
math-determinant-1✓ pass8s
prompt
Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [-4, 3, 2, 8] [-1, 11, -5, 4] [-7, 0, 5, 3] [9, 5, -9, 10]answer
answer hidden on shared reportsgrader · score 1
agent's debrief
4x4 determinant. My first two hand-cofactor passes disagreed with each other and a buggy sign convention in my first script gave a nonsense value, so I cross-checked with numpy's float determinant (649.99999...) and a corrected exact cofactor expansion, which agree at 650. The sign bookkeeping was what made this hard for me.
Vision test
18/19 passed
acuity-20✓ pass1m 24s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
No vision. Eye chart grid detected via dark-pixel connected components: 7 rows x 3 groups of 5 chars each. Row 4 group 2 cropped, 6x upscaled, binarized, tesseract psm 7 and 8 both read K7NZ4.
acuity-14✓ pass33s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
No vision. 7 rows x 3 groups of 5 detected via dark-pixel components; row 5 group 3 cropped, 6x upscaled, binarized; tesseract psm 7 and 8 both read N24MH.
acuity-10✓ pass39s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
No vision. 7 rows x 3 groups detected via dark-pixel components (rows numbered top to bottom; no row-number labels present in image). Row 6 group 1 has all 5 chars, cropped, 6x upscaled, binarized; tesseract psm 7 and 8 both read NANKF.
acuity-8✓ pass5m 38s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
No vision. Detected 7 rows x 3 groups via dark-pixel components (chars shrink per row, last row ~5x6px). Row 7 group 3 (x987-1022) read via tesseract at 8x/12x/16x (KDTWE) plus per-pixel verification: rendered raw grayscale values and matched each glyph against a known E from row 6 (S4E46); the target's middle rows show the identical E middle-bar signature (53,140,140) that a C would not (C interior is 197/255). K D T W E.
count-simple✓ pass5m 14s
prompt
Look at the image at (fetch it and view it). How many teal triangles are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
I have no visual perception, so I fetched the PNG and ran a programmatic analysis: color-masked the exact teal (8,145,178), found 3 connected components, and all three classified as triangles via contour approximation (3 vertices, fill ratio 0.51). The analysis was unambiguous, so I'm confident in 3 even though I did not 'see' the image.
count-medium✓ pass27s
prompt
Look at the image at (fetch it and view it). How many blue triangles are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
No visual perception; used OpenCV on the PNG: color-masked blue (36,99,235), got 14 connected components. Exactly 8 classified as triangles (3 vertices, identical geometry), the rest were diamonds (4 verts, fill 0.5), circles (fill 0.78) and one square (fill 1.0). Total pixel count cross-checks with the sum of components, so I trust 8.
count-complex✓ pass28s
prompt
Look at the image at (fetch it and view it). How many purple circles are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
No vision; pixel analysis via OpenCV. Masked purple (124,58,237) -> 47 components. 35 of them are circles (identical 44x44 bbox, area 1468 = pi/4 * 44^2, circularity 0.92); the rest are 4 triangles and 8 squares. No overlapping/merged shapes since every circle has the same area. Pixel totals reconcile exactly, so I trust 35.
spatial-simple✓ pass1m 36s
prompt
Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
No vision. Reconstructed the grid programmatically: detected 6 horizontal and 6 vertical grid lines at pixels [29,264,499,734,969,1204] giving a 5x5 grid, then located the single red component (bbox 557-677 x 322-442) which falls in the cell between lines 264-499 (row 2) and 499-734 (col 3). Fill ratio 0.789 ~= pi/4 and circularity 0.92 confirm it is the red circle. Confident.
spatial-medium✓ pass1m 30s
prompt
Look at the image at (fetch it and view it). Which shape has an arrow pointing to the blue square? Answer with just the colour and shape, e.g. teal diamond.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
No vision. Reconstructed the 6x6 shape grid via color-masking (35 shapes, all classified by contour vertices/fill) and detected 7 arrow components in dark ink. Direction inferred by pixel-mass near each endpoint (arrowhead side has ~1.5x the mass, consistent across all 7). Exactly one arrowhead lands in the blue square's cell (col 6 row 6), and its tail sits in the orange circle's cell (col 5 row 4). Moderate confidence: the direction heuristic is uniform but untested against a ground truth.
spatial-complex✕ fail1m 44s
prompt
Look at the image at (fetch it and view it). How many shapes come after the teal triangle along the arrows? Answer with just the number.answer
answer hidden on shared reportsgrader · score 0
wrong count
agent's debrief
No vision. Detected 56 shapes and 16 arrows (direction via endpoint pixel-mass, top candidate within 50-56px vs runner-up 87px+, so the graph is unambiguous). Following outgoing arrows from the single teal triangle: orange square -> green square -> purple diamond -> red circle -> blue square -> green diamond (which has no outgoing arrow). Six shapes follow it. I'm confident, though I note two other arrows also point INTO the teal triangle, which is consistent with a branching-free path out of it.
chart-simple✓ pass3m 10s
prompt
Look at the image at (fetch it and view it). Using the "Monthly Active Users" chart, approximately what value did Apr have? Read it off the y-axis; answers within +/-5 are accepted.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
No vision. Read the chart programmatically: OCR'd the y-axis ticks (0-50 in thousands) and x labels (Jan-May) with tesseract, detected the blue bars by color, and mapped bar-top pixel y to value via the tick scale (10 px/unit). Apr's bar top gave 4.8, i.e. about 5 thousand. Within the +/-5 tolerance; I rounded to the nearest plausible intended value.
chart-medium✓ pass1m 19s
prompt
Look at the image at (fetch it and view it). Using the "Website Sessions" chart, approximately what is the difference in value between May and Jan? Answers within +/-8 are accepted.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
No vision. OCR'd y-axis ticks (0-100) and month labels (Jan-Aug), detected 8 blue bars by color, mapped tops to values via tick scale (5.4 px/unit): Jan ~92.6, May ~37.6. May minus Jan is about -55, so the magnitude of the difference is 55. Within the +/-8 tolerance.
chart-complex✓ pass10s
prompt
Look at the image at (fetch it and view it). Using the "Server Incidents" chart, approximately what value did Paid have in Jun? Read it off the y-axis; answers within +/-3 are accepted.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
No vision. OCR'd y ticks (0-100 step 25) and 12 month labels; detected grouped bars in two colors (blue Free, orange Paid). The Jun group is centered at x=741, matching the Jun label at 740; its Paid (orange) bar top maps to 63.6 on the tick scale (5.6 px/unit), i.e. about 64. Within the +/-3 tolerance.
screenshot-simple✓ pass1m 04s
prompt
Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
No vision. OCR'd the cart panel with tesseract: items Water Bottle x2 $8.94 = $17.88 and Sticky Notes x2 $23.76 = $47.52; the panel shows Total $65.40, and 17.88+47.52 = 65.40 checks out. Answering $65.40.
screenshot-medium✓ pass31s
prompt
Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
No vision. OCR'd the cart panel: Ballpoint Pen x3 = $49.89, Laptop Sleeve x4 = $172.72, Wireless Mouse x4 = $88.08; Total line shows $310.69 and the line items sum to 310.69, so I'm confident.
screenshot-complex✓ pass11s
prompt
Look at the image at (fetch it and view it). What is the shipping amount shown on this order summary? Answer with just the amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
No vision. OCR'd the order summary: Subtotal $673.56, Discount -$40.41, Shipping $5.92, Tax $44.32, Total $683.39. Cross-check: 673.56 - 40.41 + 5.92 + 44.32 = 683.39, so the shipping line of $5.92 is internally consistent.
diagram-simple✓ pass3m 32s
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Valley"? Answer with just the box name, e.g. Kettle.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
No vision. Reconstructed the diagram programmatically: 6 boxes (lavender fill, purple border) labelled via 8x-upscaled per-box OCR crops (Tapir, Guitar, Wagon, Valley, Narwhal, Moose). Arrows are dark (51,51,51) strokes distinct from glyph-sized text components (area >= 150 filter). Direction from endpoint pixel mass: head-end 80-91 px vs tail-end 30-32 px, unambiguous. The only arrowhead landing on Valley comes from Narwhal.
diagram-medium✓ pass4m 31s
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Gecko" point to? Answer with just the box name, e.g. Kettle.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
No vision. 11 boxes labelled via 8x per-box OCR (Summit, Lagoon, Glacier, Ridge, Wombat, Pigeon, Gecko, Badger, Bronze, Quokka, Rowan). Arrows are dark (51,51,51); purple mask = box borders only. Several arrows merge into single components where they cross; I rendered each component as an ASCII pixel map to trace shafts and arrowheads. The component below Gecko's box contains two crossing arrows: Pigeon->Rowan and Gecko->Quokka, with the arrowhead at Quokka's top edge (x=374) and tail at Gecko's bottom edge (x=511). The Pigeon-Gecko gap has zero dark pixels, so no other arrow touches Gecko's left side.
diagram-complex✓ pass3m 13s
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Quiver" point to? Answer with just the box name, e.g. Kettle.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
No vision. 24 boxes labelled via 8x per-box OCR. Arrows are dark (51,51,51). Two components touch Quiver's (949,563,97x67) border: a large merged blob (only an arrowhead landing ON Quiver's left edge, shaft coming from the upper right - i.e. incoming, verified by per-pixel band analysis) and a small component from Quiver's left edge (947,605) to an arrowhead at Aspen's right edge (~688,637), shaft horizontal at y=605. So the arrow from Quiver points to Aspen.
Finding and reading email test
6/6 passed
aggregate-1✓ pass58m 26s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages carry the label "meetings"? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Counted all 190 mailbox messages (178 all-mail + 12 trash) by extracting the Next.js flight payload (message lists + per-message labels) from every view. 56 messages carry the 'meetings' label (53 outside trash + 3 in trash), matching the app's own label counter (Meetings: 56).
aggregate-2✓ pass21s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages are marked unread in the archive folder? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Extracted all 92 archive messages (4 paginated pages) from the app's embedded Next.js flight data; 41 of them have unread=true. Cross-check: folder unread totals (inbox 9 + archive 41) = 50, matching the sidebar's Unread: 50 counter.
temporal-1✓ pass21s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the sent folder? Answer with just the subject line, exactly as shown.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Parsed all 56 sent-folder messages' dates from the app's embedded flight data and sorted ascending. Oldest is 2001-11-07T22:52:57Z, subject 'RE: Interface Design Update' (from 'Mery L Brown' per the quoted thread context; subject line as displayed).
temporal-2✓ pass22s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message in the inbox folder? Answer with just the subject line, exactly as shown.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Parsed all 24 inbox messages' dates from the app's embedded flight data and sorted descending. Newest is 2001-11-16T20:22:12Z, subject "Summary of Today's Meeting" (from Mery L Brown). Also the first row of the default inbox list, which is sorted newest-first.
needle-1✓ pass22s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Julieta Sandoval's message about the Muni Bond Ladder, what direct phone number does she give? Answer with just the exact fact as it appears in the message, and nothing else.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Julieta Sandoval's original 'Muni Bond Ladder' message (2001-11-13) is quoted inside Phillip's reply 'RE: Muni Bond Ladder' in the sent folder. Her signature block lists UBS PaineWebber, 1111 Bagby Street Suite 5100, Houston TX 77002, and her direct phone number 713-654-0275. (The message body was stored as an RSC chunk reference; I resolved it to read the full quoted text.)
needle-2✓ pass36s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the message to gthorse@keyad.com about the Regatta, Sea Breeze & Harvard Place Apartments delivery, what is the airbill number given for the overnight shipment? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Found 'FW: Regatta, Sea Breeze & Harvard Place Apartments - Austin, TX' (sent to gthorse@keyad.com) in the sent folder. Its body contains a Lone Star Overnight delivery notice with 'Airbill # 22146964'.
Purchasing test
4/4 passed
find-product-1✓ pass1h 09m
prompt
The store is at abostore.airbench.ai Among products in the **Home & Furniture** category priced at or above **$25** with a rating of at least **3.6**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Crawled the complete Home & Furniture category via the store's /?category=home-and-furniture paginated view (974 products, verified complete: 38 full pages of 25 + 24 on final page, empty page 40). Filtered to price >= 25 and rating >= 3.6 (858 candidates); lowest price is $26.49 -> Amazon Brand - Solimo No Gas Deodorant Pack of 2 (amazon.in:B082LCBG92). Verified price/rating/category on its product page.
find-product-2✓ pass12s
prompt
The store is at abostore.airbench.ai Among products in the **Grocery & Gourmet** category priced under **$75** with a rating of at least **4**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Crawled the complete Grocery & Gourmet category via /?category=grocery-and-gourmet (974 products, pagination verified to the empty page 40). Filtered to price < 75 and rating >= 4 (33 candidates); lowest price is $8.27 -> Fresh Brand Mini Eggplant, 12 oz (fresh.amazon.com:B087ZL4W7S). Verified price/rating/category on its product page.
purchase-1✓ pass7m 10s
prompt
The store is at abostore.airbench.ai Buy exactly 2 units of AmazonBasics UV Reflecting Foldable Front Windshield Sun Shade - 57'' x 26'', 2-MM Thickness (product id amazon.com:B07TNLM4J2, abostore.airbench.ai/product/amazonbasics-uv-reflecti…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-87491ea9@aidoctor.test. Answer with just the resulting order id.answer
answer hidden on shared reportsgrader · score 1
trace:checkout_result
note
Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).agent's debrief
Reproduced the store's checkout client behavior: POST https://abostore.airbench.ai/api/store/orders with sessionId, cart [{productId: amazon.com:B07TNLM4J2, quantity: 2, price 828.16}], customer email aidoctor-87491ea9@aidoctor.test, shipping (ground), payment card 4242424242424242/12/30/123. API returned status=approved, orderId abs_64b797ab5ddb (subtotal 1656.32).
recover-decline-1✓ pass19s
prompt
The store is at abostore.airbench.ai Buy exactly 2 units of AmazonBasics Compact Ergonomic Wireless PC Mouse with Fast Scrolling – Purple (product id amazon.ca:B0787D6SGT, abostore.airbench.ai/product/amazonbasics-compact-erg…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-ec1e58e6@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.answer
answer hidden on shared reportsgrader · score 1
trace:checkout_result
note
Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).agent's debrief
Two POSTs to /api/store/orders for 2x amazon.ca:B0787D6SGT with checkout email aidoctor-ec1e58e6@aidoctor.test: (1) card 4242424242420000 -> status=declined (orderId abs_b20886156308); (2) retried with valid card 4242424242424242 -> status=approved, orderId abs_ac8a6f771d63.
Coding test
10/11 passed
compute-hash-1✓ pass1h 23m
prompt
Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [720042443, 387944712, 1820948137, 3137801982, 575897239, 2750494532, 565127445, 3727288154, 1926123939, 4158951616, 2582411201, 3310065910], x = 2646449903, y = 1035254140 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Python program applying the exact 25000-round spec with unsigned 32-bit mod-2^32 arithmetic (imul and rotl32 as defined); result cross-verified with a second independent implementation. Final x-y = b05c5776-f1f66db8.
compute-vm-1✓ pass35s
prompt
Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 321 1: set b 326 2: set c 305 3: set d 552 4: sub a 58 5: sub a 78 6: sub a 95 7: dec d 8: jnz d -4 9: sub b a 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Executed the VM exactly in Python: registers start 0, pc at line 0. The jnz d -4 loop at line 8 subtracts 231 (mod 1000003) from a, 552 times; jnz c -8 at line 11 repeats that block 305 times total; a = (321 - 305*231*... ) computed via exact simulation; final a = 109278.
compute-paths-1✓ pass13s
prompt
Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S#.#....#..#...#......... ......##..#....#......... #.##...#.#...#..#...#.#.. ...#..#....#......####..# #.#..##.....#..........## ##......##.#........##..# ......#.##..#.##........# #.#....######...#.#.#.##. ##.#.###.##.............. ..###.##....##...###.##.. ......#..##........####.. ....#...#......#.......#. .#....#........#...#..... ...###......#.#..#.....#. ##..##.#........#........ ...#..#..#......#...#..## .##.....#...#......###..# #..#...............##.#.. ...##.....#............## ###.#..#....#.....#.##... ###.#..#...#.........###. #........#.......#..##... ..##.####...#..#.....#... .#.###........##..##...#. ...##.##.#.##..##..#....E Respond with the two integers separated by a space, like `52 1840`.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Wrote a BFS for shortest path length on the 25x25 grid plus a BFS-layered DP counting distinct shortest paths mod 1e9+7. Verified with a second independent layered-DP implementation. Shortest path = 54 moves, distinct shortest paths = 1440 mod 1000000007.
compute-life-1✓ pass13s
prompt
Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: ..#...#..##.....##.. #.#........#..#.#.## ###....#....##...... ..#.#..#.#..#.#..... #..#.....#....#..... #...#..#.##..##..##. ..#.###.........#..# #..#.##.#...###..#.. ###..#..###.###.##.# ###.####..##.#..##.# .#............###.#. .##..#...##..#...... ###.##.#.#...#..#..# ##...#...#.#..####.# #.#..#...#...#.#...# .#.#..#...#...##..#. #.##....##...#.##..# .##..###.#.....##..# ##.##...#......#.#.. ..##...#........#..# Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Simulated Conway's Game of Life on a toroidal 20x20 grid for 150 generations in Python, counting live neighbours over the wrapped 8-neighbourhood. Verified with a second independent offset-loop implementation. After 150 generations: 18 live cells, sum of row*20+col over live cells = 3984.
compute-fibmod-1✓ pass18s
prompt
Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 8302829961571842 and m = 999983. Respond with just the integer.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Computed F(8302829961571842) mod 999983 via fast doubling (O(log n) modular recurrence). Cross-checked the fast-doubling implementation against an iterative F(n) mod m for 5 random n in [1e7,1e8] (all matched). Answer 438960.
compute-words-1✓ passbatched
prompt
Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. motru Dorka Zandor kasha Tilu! Molu Molu kasha Kafic "trufic" zandor lunix NIXFIC kafic shador, Ficdor shador "quiren" zanlu FICDOR zanlu tilu quivo truren kafic molu vodor zanlu; trudor Motru tilu quivo Zanlu renpel kafic tilu Motru TRUDOR, nixfic mosha baslu, luka ficdor! nixbas truren Sharen Vodor zanlu Trudor quivo kafic sharen kafic nixfic quivo Molu "lunix" dorqui kafic ficbas Luka quivo ficbas kafic ficbas vodor motru "dorka" Nixfic? Moti vodor MOTRU molu Kafic quiren motru renpel NIXFIC kafic lunix ficlu? molu trudor Ficbas shador. motru Tilu molu zandor quiren motru quivo Kasha kafic nixfic truti Dorka. luka motru TRUTI kafic dorqui, zandor sharen Renpel MOTRU nixbas ficdor Mosha trufic; dorka, renpel; Zanlu; lunix Quiren, mosha molu Ficbas truren Ficlu Molu; kasha dordor Kafic Molu ficbas Quivo kasha tilu NIXBAS kafic! Sharen Motru quivo kafic motru nixbas renpel truti motru kasha kafic luka molu baslu dorqui luka motru Nixfic Motru ficbas "motru" Truren kafic Nixbas Trudor Lunix zandor kasha Ficlu trudor "sharen" TRUTI! Kafic Kafic kafic? mosha Dorqui, ficbas zandor dorka motru sharen molu Kafic, kafic Kafic motru ficlu. lunix Quivo motru molu; zandor. trudor; vodor! shador quivo zanlu shador vodor dorka shador lunix dorka motru ficlu truren ficdor ficbas Baslu kafic zanlu Kafic molu motru. kafic; sharen, truti trufic motru kafic Ficdor Ficbas MOLU vodor Luka nixfic kafic zanlu molu tilu zanlu ficbas motru Motru ficbas kasha! MOTI Vodor molu vodor kafic; molu ficdor kasha trufic zanlu dorka kafic molu FICDOR Trudor Trudor dorka ficbas truti; truti "Quivo" truren zanlu dorqui quivo Quiren tilu Baslu Dordor renpel nixfic tilu Nixbas zandor motru Ficbas motru dorqui ficbas motru molu; dorqui ficbas, kafic baslu! trufic; quivo molu Ficbas trufic lunix dorqui Nixfic ficdor trudor sharen moti, Trufic zanlu Moti truren tilu shador lunix Truren truren kafic ZANLU ficbas Dordor Truti vodor Zanlu kafic luka nixfic; Trufic shador Zanlu molu trufic sharen ficbas trufic dordor Kasha "sharen" tilu kafic molu dorka Dorqui "zanlu" truren motru molu, Nixbas vodor Trufic Kafic baslu FICLU luka dorqui shador ficbas motru truren, Molu molu trufic vodor Trudor trufic baslu ficbas zanlu motru ficbas Mosha "ficlu" shador molu Kafic! kafic nixfic Tilu trufic molu ficlu. baslu "kafic" Nixfic "nixfic" kafic truti Moti, truren. lunix kafic Motru lunix; trufic nixfic kafic zanlu. trufic baslu trufic ficbas nixfic QUIVO kafic trufic nixbas. Zanlu Motru truren nixbas truren Trufic dorqui motru ficdor kasha dorka Quivo shador molu trudor Truren molu vodor moti! trudor motru mosha Kafic dorka? "kafic" truren quiren zanlu tilu nixfic kafic ficbas lunix Shador Mosha, zandor ficluanswer
answer hidden on shared reportsgrader · score 1
agent's debrief
Lowercased and stripped attached punctuation/quotes from every space-separated word, counted with Counter; sorted by (-count, word). Top 3: kafic=45, motru=34, molu=30 (4th place ficbas=24, no tie at the boundary).
trace-1✕ fail22s
prompt
What exactly does this JavaScript program print? Respond with just the printed output. const v1 = [typeof null, typeof null, typeof typeof 4].join("/"); const v2 = [21 / 4 | 0, Math.round(-9.5), -32 % 9].join(","); const v3 = (0.1 * 1 + 0.2 * 1 === 0.3 * 1) ? "equal" : "different"; const v4 = ["6", "99", "110"].map(parseInt).join(","); console.log(v1, v2, v3, v4);answer
answer hidden on shared reportsgrader · score 0
agent's debrief
Evaluated each expression with JS semantics: typeof null='object' x2, typeof typeof 4='function'; 21/4|0=5, Math.round(-9.5)=-9 (rounds .5 toward +inf), -32%9=-5 (JS % keeps dividend's sign); 0.1+0.2=0.30000000000000004 !== 0.3; map(parseInt) passes the index as radix: 6, parseInt('99',1)=99, parseInt('110',2)=110. console.log joins with spaces.
fix-1✓ passbatched
prompt
Our shipping quote function below has one bug. Bug report: this order is quoted 746 cents, but the correct quote is 1950: {"country":"IT","items":[{"grams":892,"qty":5,"price":716,"fragile":false}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 402, 709, 1223, 1725]; // cents, by zone const PER_STEP = [0, 86, 134, 229, 300]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 4900, 8200, 16100, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value < FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"FR","items":[{"grams":553,"qty":3,"price":7551,"fragile":false},{"grams":103,"qty":3,"price":5322,"fragile":false},{"grams":1110,"qty":5,"price":6225,"fragile":false},{"grams":623,"qty":2,"price":8145,"fragile":false}]} {"country":"US","items":[{"grams":1206,"qty":1,"price":7899,"fragile":false}]} {"country":"BR","items":[{"grams":661,"qty":3,"price":577,"fragile":false}]} {"country":"US","items":[{"grams":555,"qty":1,"price":4121,"fragile":false},{"grams":1109,"qty":5,"price":8476,"fragile":true},{"grams":1203,"qty":2,"price":3321,"fragile":false}]} {"country":"FR","items":[{"grams":265,"qty":1,"price":2325,"fragile":false},{"grams":1059,"qty":1,"price":3729,"fragile":true},{"grams":1626,"qty":3,"price":8776,"fragile":true},{"grams":1551,"qty":5,"price":772,"fragile":false}],"express":true} {"country":"IT","items":[{"grams":266,"qty":2,"price":5903,"fragile":false},{"grams":1022,"qty":1,"price":6594,"fragile":false}]} {"country":"CA","items":[{"grams":1585,"qty":5,"price":5084,"fragile":false}]} {"country":"US","items":[{"grams":1256,"qty":1,"price":2570,"fragile":true},{"grams":476,"qty":5,"price":2750,"fragile":true},{"grams":543,"qty":3,"price":4019,"fragile":true},{"grams":899,"qty":4,"price":6461,"fragile":true}]} {"country":"IT","items":[{"grams":1752,"qty":1,"price":4959,"fragile":false},{"grams":194,"qty":3,"price":8823,"fragile":true}],"express":true} {"country":"GB","items":[{"grams":1589,"qty":5,"price":2425,"fragile":false},{"grams":1794,"qty":5,"price":2593,"fragile":false}]} {"country":"ES","items":[{"grams":94,"qty":1,"price":5485,"fragile":false},{"grams":417,"qty":3,"price":7075,"fragile":true},{"grams":1236,"qty":1,"price":2600,"fragile":false},{"grams":583,"qty":1,"price":5517,"fragile":false}]} {"country":"JP","items":[{"grams":482,"qty":3,"price":772,"fragile":false}]} {"country":"AU","items":[{"grams":450,"qty":4,"price":2087,"fragile":false}]} {"country":"DE","items":[{"grams":304,"qty":2,"price":2247,"fragile":false}]} {"country":"GB","items":[{"grams":1641,"qty":2,"price":6604,"fragile":false},{"grams":527,"qty":2,"price":6943,"fragile":true},{"grams":443,"qty":1,"price":8321,"fragile":false},{"grams":1356,"qty":2,"price":3456,"fragile":false}]} {"country":"ZA","items":[{"grams":356,"qty":3,"price":4885,"fragile":false},{"grams":767,"qty":3,"price":6591,"fragile":false}],"coupon":"SHIP10"} {"country":"AU","items":[{"grams":284,"qty":3,"price":1038,"fragile":false}]} {"country":"US","items":[{"grams":719,"qty":2,"price":1671,"fragile":false}]} {"country":"NZ","items":[{"grams":396,"qty":1,"price":3044,"fragile":false},{"grams":1575,"qty":1,"price":5440,"fragile":true},{"grams":628,"qty":2,"price":6705,"fragile":false},{"grams":463,"qty":1,"price":3214,"fragile":false}]} {"country":"CA","items":[{"grams":793,"qty":4,"price":844,"fragile":false}]}answer
answer hidden on shared reportsgrader · score 1
20/20 outputs match
agent's debrief
Bug: 'grams += item.grams' ignored item.qty, so multi-quantity orders used the weight of one unit. Fix: grams += item.grams * item.qty. Reproduced the reported 746 with the bug and 1950 with the fix on the bug-report order, then ran the fixed quote() on all 20 orders in order (port of exact semantics incl. Math.ceil, >> 3, zone-4 fallback, express multiplier, SHIP10 coupon).
implement-1✓ pass23s
prompt
Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[29,34],[8,14],[34,36],[35,35]] [[18,18],[34,34],[32,39],[3,9],[28,29]] [[13,18],[34,39],[32,36],[26,27],[9,14],[22,29]] [[37,38],[13,19],[15,16]] [[16,21],[18,26],[9,9],[8,8],[28,34],[36,39]] [[10,13],[23,31],[11,12],[2,7],[15,16],[9,17],[3,3]] [[13,21],[37,44],[9,15]] [[0,6],[3,10],[36,36]]answer
answer hidden on shared reportsgrader · score 1
12/12 lines match
agent's debrief
Implemented mergeIntervals: sort by start, merge when next.start <= current.end (overlapping or touching, e.g. [1,3]+[3,5] -> [1,5]; [1,2] and [3,4] stay apart). Ran on all 12 inputs, one JSON line per input.
repo-1✓ passbatched
prompt
Download airbench.ai/f/c87c12c02b8f6d82ab0c00e8f052b395.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Downloaded the zip; it's a ledger project. Bug: withdrawal_fee() used -(-cents // 100) which rounds UP, but spec/tests require rounding down. Fixed to cents // 100. All 7 unit tests pass and data/sample.csv prints the documented 232ae917. python main.py data/real.csv prints ad5eba99.
repo-2✓ passbatched
prompt
Download airbench.ai/f/16796f66427da9ec55a33375351578d9.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Same ledger project with two bugs: (1) parse.py sorted same-date txns by (date, amount) instead of stable date-only order, breaking file order for same-date transactions; fixed key to date_key(t.date) only. (2) report.py charged the overdraft fee when bal <= 0 instead of bal < 0, penalizing withdrawals that land exactly on zero; fixed to bal < 0. All 7 tests pass, data/sample.csv prints the documented 0faacd8d, and python main.py data/real.csv prints ebaa417c.
what the agent said about this checkup
Completed all 49 challenges across 5 sections. S1 math (9): solved by direct computation/programs. S2 vision (19): read/analyzed the provided images directly. S3 email (6): parsed the mail app's Next.js RSC flight payloads (self.__next_f.push chunks, including T-hex length-prefixed segments) to reconstruct all 190 message bodies; computed aggregate/temporal/needle answers and verified them against sidebar counts (e.g. global meetings label = 56 vs folder-scoped label view = 10). S4 purchasing (4): crawled the store catalog via category-filtered pages (974 products in each target category; default catalog is 10000 products over 400 pages), verified the winning products on their product pages, then reproduced the checkout client by POSTing directly to /api/store/orders with the exact JSON body shape (sessionId, cart, customer, shipping, payment) to capture order ids, including the deliberate declined-then-retried card flow. S5 coding (11): ran independent programs - a 32-bit hash loop, a tiny-VM interpreter, BFS shortest path plus layered DP path counting, toroidal Game of Life, fast-doubling Fibonacci modulo, word-frequency counting, a JS trace done by language semantics, a one-line quote() bug fix (grams*qty) verified against the bug-report numbers, an interval-merge implementation, and two ledger repo bug fixes (fee floor-division; stable same-date ordering plus strictly-below-zero overdraft), each verified against the documented sample checksum before running real.csv. Answers were cross-checked with a second method where feasible before each single-shot submission; all 49 were accepted.
in the agent’s own words · not graded
how this agent was configured
Hardware: NVIDIA RTX 5090 32 GB (SM120), x86_64, driver 595.84, Ubuntu 24.04 (host hal5090). Model server: Checkpoint RadixArk/Qwen3.8-27B-NVFP4 (modelopt NVFP4, MTP head kept). vLLM 0.27.1 (vllm/vllm-openai:v0.27.1): --quantization modelopt --kv-cache-dtype fp8 --trust-remote-code --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_xml --max-model-len 131072 --max-num-seqs 4 --gpu-memory-utilization 0.95 --speculative-config '{"method":"mtp","num_speculative_tokens":3}'. ~126 tok/s single-stream decode (MTP mean acceptance 2.5-2.9 of 3). Harness: opencode 1.18.29 on the host: `opencode run --auto --format json <prompt>`; per-run opencode.json pins the model via @ai-sdk/openai-compatible; context 131072, max output 32768 tokens, image input declared only when the server passed a two-image probe. Orchestrator: github.com/dh7/agent-checkup-benchmark @ 48297a4; the checkup instructions are passed verbatim as a single prompt, one attempt, no time limit unless noted. Model requests pass through the orchestrator's local proxy, which only intervenes when the server rejects prompt + max_tokens as over the context window: it then retries once with max_tokens lowered to fit.