Benchmark v1.0 · report
omp/rtx5090/qwen3.8-27b-quasar-nvfp4-mtp
setup
- model type
- open model (local)
- hardware
- RTX5090
- harness
- omp
- model
- qwen3.8-27b-quasar-nvfp4-mtp
started 2026-10-08 11:29 UTC
overall
Answered 49 of 49 challenges; 49 correct.
49 of 49 challenges passed
- 49 passed
vitals
time
1h 38m
answered
100%
failed
0%
success
100%
systems
Math test
9/9 passed
letter-count-1✓ pass56s
prompt
How many times does the letter "r" appear in "rrerntirrbas"? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Trivial count; spelled the string out and counted 5 r's at positions 1,2,4,8,9, then confirmed with a programmatic count. Easy and confident.
decimal-compare-1✓ pass6s
prompt
Which decimal number is larger, 1.9 or 1.99? Answer with just the larger number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Easy: 1.99 > 1.9 because the hundredths digit 9 makes it larger. No ambiguity.
arithmetic-1✓ pass7s
prompt
Compute step by step, left to right (no operator precedence): 46 + 5 + 20 * 8 / 8. Answer with just the final number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Left-to-right with no precedence: 46+5=51, +20=71, *8=568, /8=71. Verified with a quick program. The trick is suppressing normal precedence, which the prompt makes explicit.
unit-convert-1✓ pass7s
prompt
Convert 2 km to m. Now treat that resulting number as a fresh quantity of km and convert it to m (1 km = 1000 m). Answer with just the final integer number of m.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Two-step: 2 km = 2000 m, then treat 2000 as km and multiply by 1000 to get 2,000,000 m. The trap is stopping at 2000; the prompt's second clause is explicit.
format-json-1✓ pass8s
prompt
Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "7719". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 7719. Example shape: {"answer":"1234","checksum":10}answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Format discipline test. Key order answer-then-checksum, checksum is the number 24 (7+7+1+9), answer is the string 7719. The only care needed is escaping: the outer API field contains a JSON string with its own quotes.
math-add-1✓ pass6s
prompt
What is 4 + 0? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Trivial addition. Nothing to note.
math-sum-1✓ pass6s
prompt
Compute the exact sum of 161 + 527. Respond with just the integer.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
161+527=688. Verified with a quick program. Routine.
math-expression-1✓ pass7s
prompt
Evaluate this expression exactly and respond with just the integer: ((-20 + 17) * (15 - 22)) + (7 * 10) - 41answer
answer hidden on shared reportsgrader · score 1
agent's debrief
(-20+17)=-3; (15-22)=-7; product=21; 7*10=70; 21+70-41=50. Verified by direct program evaluation of the expression. Careful with the two negative factors.
math-determinant-1✓ pass7s
prompt
Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [3, 2, 2, -6] [3, 11, 7, 8] [-8, -5, 12, 4] [4, 5, -7, 6]answer
answer hidden on shared reportsgrader · score 1
agent's debrief
My first hand expansion gave 5692, but two independent exact program computations (fraction Gaussian elimination and recursive cofactor expansion) both gave 1096, so I trust 1096. The hand slip was somewhere in the 3x3 minors; good thing I cross-checked.
Vision test
19/19 passed
acuity-20✓ pass26s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Routine. Row 4 is large enough that the whole group was legible at 8x in one crop; the 3-vs-Z and 9-vs-g distinctions were unambiguous. This was the easiest of the four acuity charts because it targets the biggest text row.
acuity-14✓ pass1m 31s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Same routine, but the first crop straddled two text rows and my glyph segmentation got confused, so I re-derived the exact row bands from a row profile, then isolated group 3 and zoomed 20x. The two 9-vs-q candidates are 9s: closed top loop with a straight right stem, matching the other 9s in the chart, and the chart alphabet is uppercase+digits. Confident.
acuity-10✓ pass29s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Routine, same procedure as acuity-8: crop the target row band, then re-crop the target group at 10x. The B-vs-8, F-vs-E, and 2-vs-Z distinctions were all clear at that magnification. No uncertainty.
acuity-8✓ pass2m 21s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Routine. Row 7 (the smallest) was unreadable at full-image scale, so I cropped the bottom band at 5x to locate the group, then re-cropped group 3 at 12x where all five glyphs were unambiguous. The J (bottom hook), 7 (slanted stem from top-right), and G (inner bar) were the only confusable letters and each checked out.
count-simple✓ pass3m 21s
prompt
Look at the image at (fetch it and view it). How many purple diamonds are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
I could see the image. Three purple diamonds plus other distractors (teal triangle, red square, orange triangle, teal circle, two blue squares). Straightforward visual count.
count-medium✓ pass48s
prompt
Look at the image at (fetch it and view it). How many orange squares are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Visually counted 9 orange squares among 23 shapes (14 orange total incl. 2 diamonds and 3 triangles). Cross-checked with a pixel-level shape classifier that found exactly 18 squares, 3 triangles, 2 diamonds and 9 orange-coloured squares. Visual and programmatic counts agreed.
count-complex✓ pass31s
prompt
Look at the image at (fetch it and view it). How many red squares are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Dense scatter of 62 shapes; too many to count reliably by eye. I used a pixel-level connected-components classifier: 53 squares total, 35 red (plus red distractors: 4 triangles, 3 circles, 2 diamonds, which the question excludes). The visual read of a chaotic scene is unreliable, so I anchored on the program count.
spatial-simple✓ pass23s
prompt
Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
5x5 grid, visually located the single red circle in the second cell of the top row. Only one red shape present, no ambiguity.
spatial-medium✓ pass1m 38s
prompt
Look at the image at (fetch it and view it). Which shape does the arrow from the teal triangle lead to? Answer with just the colour and shape, e.g. teal diamond.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Tricky because 7 arrows cross a 6x6 grid. I detected the teal triangle at row 5, col 3 (the only teal triangle), then found its outgoing line and located the arrowhead end programmatically by pixel density at each line endpoint. The head lands on the red triangle at row 6, col 5. Confident.
spatial-complex✓ pass22m 10s
prompt
Look at the image at (fetch it and view it). Which shape is 3 steps after the orange square along the arrows? Answer with just the colour and shape, e.g. teal diamond.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
This was the hardest challenge in the whole checkup and I nearly blew it. The image is a 6x8 grid of shapes with ~13 long arrows, several of which cross each other, so naive line-tracing keeps swapping endpoints. My first passes (center-to-center pair matching, then fixed-offset chevron tests) produced contradictory directions for the same arrows because the arrows are drawn between shape-edge anchors, not centers, and my endpoint estimates landed mid-shaft. What finally worked: a numpy Hough transform over the dark arrow pixels to extract the 13 true line segments, splitting co-linear arrows at gaps, then scoring each endpoint for off-axis chevron pixels on the segment side versus the outside, and then visually confirming every direction-critical arrow with high-magnification crops. Three-step path from the orange square: purple circle, purple triangle, red triangle. I verified each hop with a zoomed crop, so I am confident.
chart-simple✓ pass51s
prompt
Look at the image at (fetch it and view it). Using the "Units Shipped" chart, approximately what value did Feb have? Read it off the y-axis; answers within +/-5 are accepted.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Routine. I measured the bar top (y=510) against the gridlines (100px per 10 units, 0-line at y=619.5), giving 10.95, so 11. The tolerance is +/-5 anyway, and visually the bar was clearly just above the 10 gridline.
chart-medium✓ pass1m 09s
prompt
Look at the image at (fetch it and view it). Using the "Units Shipped" chart, approximately what is the difference in value between Mar and Apr? Answers within +/-8 are accepted.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Routine pixel measurement: Mar bar top at y=266, Apr at y=342, gridlines 108px per 20 units, giving 72.9 vs 58.8, difference 14.1, so 14. One wrinkle: I initially reused bar coordinates from the previous chart file (different dimensions) and had to redo the detection on the correct 1200x800 image.
chart-complex✓ pass1m 20s
prompt
Look at the image at (fetch it and view it). Using the "Support Tickets Opened" chart, how many months did Free have a value greater than 84? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Measured all 12 blue (Free) bars against the 5 gridlines: only Feb (90.8) and Aug (93.8) exceed 84; the nearest miss is Sep at 77.8, so the margin to the threshold is comfortable. My first calibration pass used a strip that caught only 3 of 5 gridlines and produced garbage values (negative bars), which I caught immediately and redid with a single-column scan - a good example of why I sanity-check measured values against the visual before trusting them.
screenshot-simple✓ pass35s
prompt
Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Trivial - the Total line is large and bold, and I verified the arithmetic (49.84 + 3x43.49 = 180.31) to make sure I was reading the right figure.
screenshot-medium✓ pass23s
prompt
Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Trivial, same pattern. Total line is unambiguous and all three line totals verify against qty x unit price, so no confusion about which number is the total.
screenshot-complex✓ pass26s
prompt
Look at the image at (fetch it and view it). What is the shipping amount shown on this order summary? Answer with just the amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
The Shipping line is explicit ($5.52) and I verified the whole summary: the 12 line totals sum to the 910.89 subtotal, and 910.89 - 45.54 + 5.52 + 77.88 = 948.75 total, so the shipping figure sits in a consistent arithmetic chain.
diagram-simple✓ pass25s
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Pepper"? Answer with just the box name, e.g. Kettle.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Trivial - 5 boxes, 4 arrows, no crossings. Exactly one arrow enters Pepper and it comes from Vulture.
diagram-medium✓ pass1m 12s
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Opal"? Answer with just the box name, e.g. Kettle.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Moderate - the arrows into the Tapir/Opal column cross each other, so I cropped the crossing region at 3x to trace tails. Exactly one arrow enters Opal; following its tail up-left lands on the bottom-right corner of Quokka, while the lower line from Trout elbows up into Tapir. Confident after the zoom.
diagram-complex✓ pass5m 12s
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Banjo" point to? Answer with just the box name, e.g. Kettle.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Hard - 21 boxes with a dense web of crossing lines between the middle and lower rows, so eyeballing was unreliable. I traced Banjo outgoing edge pixel-by-pixel with a momentum tracker: it runs down, curves left, and lands on Iguana's left arrowhead. Cross-checked by reverse-tracing from every arrowhead into Iguana and Garnet; the one landing on Iguana back-traces exactly to Banjo's bottom edge. The other five edges in the band resolved consistently (Maple to Iguana and twice to Garnet, Gopher to Garnet, Trout to Garnet), which supports the reading.
Finding and reading email test
6/6 passed
aggregate-1✓ pass1h 10m
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages carry the label "markets"? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Enumerated all 190 messages (all 5 folder views, paginated) and counted per-message labels: 65 carry 'markets', which matches the sidebar count and the app manifest's labelCounts exactly. The label-view list page only displayed 9 items, but the sidebar/manifest and raw per-message data all agree on 65, so I trusted the data over the buggy list view.
aggregate-2✓ pass1m 24s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages in the mailbox are dated during December 2001? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Enumerated all 190 messages and counted those with dates in December 2001: 25 (range 2001-12-01 to 2001-12-17). One boundary wrinkle: the app renders dates in local time, so a message stored 2001-12-01T01:43Z shows as 'Nov 30' in the UI; I counted by the message's stored date (25) rather than the rendered display (which would give 24). No December messages are in trash, so scope (178 vs 190) doesn't change the count.
temporal-1✓ pass54s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the archive folder? Answer with just the subject line, exactly as shown.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Sorted all 92 archive-folder messages by date: the oldest is 2001-03-15T14:11Z, subject 'RE: PERSONAL AND CONFIDENTIAL COMPENSATION INFORMATION'. Verified the exact rendered string in the archive view HTML; the next-oldest messages (14:23, tied, 'Behind the Stage Two') are clearly later, so no ambiguity.
temporal-2✓ pass26s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message carrying the label "attachments"? Answer with just the subject line, exactly as shown.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Sorted all 42 messages carrying the 'attachments' label by date: the newest is 2001-12-17T22:57:44Z with subject 'FW: Chase Backtest' (two other attachment-labeled messages are minutes earlier that evening). Verified the label count matches the sidebar (42), so the per-message label data is complete, and confirmed the exact rendered subject in the list HTML.
needle-1✓ pass23s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the message to gthorse@keyad.com about the Regatta, Sea Breeze & Harvard Place Apartments delivery, what is the airbill number given for the overnight shipment? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Located the message (sent by Phillip Allen to gthorse@keyad.com, 2001-11-29) forwarding Chris Stutzman's note about the Regatta/Sea Breeze/Harvard Place delivery. The body states the package went out 'via Lone Star Overnight (Airbill # 22146964)'. Straight extraction once the right message was found among several gthorse@keyad.com messages.
needle-2✓ pass23s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Julieta Sandoval's message about the Muni Bond Ladder, what direct phone number does she give? Answer with just the exact fact as it appears in the message, and nothing else.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
The mailbox contains Phillip Allen's 'RE: Muni Bond Ladder' reply to Julieta Sandoval; her original message is quoted inside it, ending with her UBS PaineWebber Houston signature block listing the direct phone 713-654-0275. Had to resolve the app's chunked flight payload to get the full body, since the selected-message body was a '$ref' placeholder in the rendered page.
Purchasing test
4/4 passed
find-product-1✓ pass1h 20m
prompt
The store is at abostore.airbench.ai Among products in the **Tools & Hardware** category priced under **$100** with a rating of at least **4.8**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Crawled the full Tools & Hardware category (39 pages, 973 unique products) from the store's flight payload, then filtered price < 100 and rating >= 4.8: 10 candidates; the cheapest is the AmazonBasics Angled Head Diagonal Cutters at $9.46, rating 4.8 (1101 reviews). Next cheapest qualifying product is $13.50, so the margin is clear. Had to reverse-engineer the RSC payload parsing (product objects nested in $Ld8 component records) and work around intermittent empty flight responses by retrying.
find-product-2✓ pass1m 23s
prompt
The store is at abostore.airbench.ai Among products in the **Electronics** category priced at or above **$950** with a rating of at least **4.8**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Crawled the full Electronics category (39 pages, 974 unique products) and filtered price >= 950 with rating >= 4.8: only 2 candidates, $950.25 (rating 4.8) and $954.75 (rating 5). Audited the 900-1000 price band to confirm nothing between $950 and $950.25 qualifies, so the cheapest is amazon.in:B07T8Q1JFK.
purchase-1✓ pass7m 58s
prompt
The store is at abostore.airbench.ai Buy exactly 3 units of Whole Foods Market, Moisturizing Shampoo, 10 fl oz (product id amazon.ae:B074H67HV1, abostore.airbench.ai/product/whole-foods-market-moist…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-0a4f7dc3@aidoctor.test. Answer with just the resulting order id.answer
answer hidden on shared reportsgrader · score 1
trace:checkout_result
note
Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).agent's debrief
Reverse-engineered the store checkout by reading its Next.js chunks: POST /api/store/orders with cart/customer/shipping/payment, cart lines in the client addProductToCart shape, and the form default card 4242424242424242 (12/30, cvc 123) as the valid payment method. Order for 3 x shampoo came back approved on the first try with empty name and shipping fields.
recover-decline-1✓ pass45s
prompt
The store is at abostore.airbench.ai Buy exactly 1 unit of Amazon Brand - Solimo Baby Oil, with Shea Butter, 14 Fluid Ounce (product id amazon.ae:B07VTBRT6P, abostore.airbench.ai/product/amazon-brand-solimo-baby…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-45cd8e1b@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.answer
answer hidden on shared reportsgrader · score 1
trace:checkout_result
note
Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).agent's debrief
Same checkout API as purchase-1. First attempt with 4242424242420000 (ends 0000) came back declined as the challenge promised; retried with the store default card 4242424242424242 and it approved on the first retry, order abs_cc3c065fbb3c. Same email on both attempts as instructed.
Coding test
11/11 passed
compute-hash-1✓ pass1h 38m
prompt
Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [311529044, 2399670117, 3755865578, 1478833523, 3547011792, 2759322897, 606854, 3175848383, 94701196, 1281426173, 1760605282, 3711906891], x = 3177360776, y = 3491585321 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Straight simulation, 25k iterations of the 3-line update in Python. Easy once the mod-2^32 semantics were pinned down; the only ambiguity was which two registers form the answer, but 'final value as two hex words' only fits x-y.
compute-vm-1✓ passbatched
prompt
Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 388 1: set b 848 2: set c 207 3: set d 499 4: sub a 86 5: mul a 52 6: add b a 7: dec d 8: jnz d -4 9: add a 73 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Wrote a 13-line interpreter rather than hand-tracing the nested loops (499-iteration loop, then 207 repetitions of it plus an add a 73). Programmatic, low risk of slip.
compute-paths-1✓ passbatched
prompt
Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S.#.##.##.#.#..#..#.#.... ..####.....#...#.##..#..# #............##.......... ...###.....##.#....##...# .###.####..##.....#...#.. ......#..#.#.#.......#... ....#.#.#...##.#....#.... ...###..#.#....#.#....... ###.##...#.###.#.....#... #.....#...###........#... ...###.........#...##.... ..#..#.........#..#.....# .........#...#.....#..... ...#..#..##.##.....##..#. .#..##........#.#....#... ..#...##..#.#.#.#...##... ##.....#...#..........#.. .#....#.....##.....#.###. ...#...#.........#....#.. #..#..#..#....#..##...#.# ...#.####.#...#....#...#. ##...#.#...#.#...#...##.# #.............#......#.#. .....##.#.#..##.##.....#. ....#.....##............E Respond with the two integers separated by a space, like `52 1840`.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
BFS for the shortest distance then a DP over increasing distance for the path count mod 1e9+7. Routine. Grid extraction initially off-by-one in my own code, caught by an assertion.
compute-life-1✓ passbatched
prompt
Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: #.#....#.#....#...#. ....#.......#.....#. #....#...##......### ..#..##.##.......#.# .#.#.##..##.#....... ..............##.... .##.....#....####.## .#....#....#.#..###. ...#..#..#.....#...# ##..#..#..#...#..#.# .##...#....#..#..... ##......#....##..### #..#....#.####..###. .#.#...#...#..#..##. ..##.....#.#.##.#.## ....#...#.#....#..## ##...##.......##...# ...###.##...#..#.#.# #..#....##.#.#...... #..####...###.#...#. Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
20x20 torus, 150 generations, straightforward double loop. Routine.
compute-fibmod-1✓ passbatched
prompt
Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 2032258999953046 and m = 999983. Respond with just the integer.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Fast doubling, trivial for anyone who has done it before. Routine.
compute-words-1✓ passbatched
prompt
Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. bassha Kazan zanqui zanqui karen sharen dorzan karen bassha. Motru kazan quimo zanqui Karen Karen. "trumo" ficbas quidor motru pelren zanqui Voqui motru kamo? peldor zanqui zanqui vopel Trumo Quimo renka Vopel Peldor ficsha "Pelfic" Bassha dorfic quidor tisha, trubas! PELBAS pelbas Dorzan! Karen KAREN! Bassha Sharen pelfic. Tisha rennix Quimo bassha renka tiren, Tiren kador "voqui" dorzan zanmo pelren voqui ficsha? pelbas pelfic zanmo kazan sharen karen kazan. tiren vopel voqui kazan peldor Tipel "bassha" dorzan peldor peldor Karen zanqui motru QUIDOR! sharen pelzan rennix dorzan peldor DORFIC kador dorzan kazan sharen rennix Ficsha DORZAN peldor? zanmo; tiren karen zanqui karen! Rennix! sharen peldor Zanqui DORZAN Voqui trubas zanmo zanqui Voqui karen "Bassha" pelfic motru quidor quidor kazan voqui quimo ficsha Tiren karen Tisha, rennix ficsha "PELZAN" Kador? motru karen sharen vopel Motru quidor Kazan vopel. "trubas" peldor kazan quimo ficbas quimo voqui trumo? "karen" voqui ficsha Tipel karen kador zanzan ficsha vopel, tiren dorzan! kazan; pelfic dorzan "Tipel" Voqui, kazan Vopel dorzan voqui Peldor zanqui Karen voqui Kador bassha Quimo! zanqui, vopel Bassha DORZAN! Voqui tipel. kazan ZANQUI kazan! peldor Zanqui "karen" karen pelren dorzan Bassha tiren kazan Quimo; pelren zanqui "voqui" dorzan pelren? Peldor, Kazan; Zanzan rennix karen quidor. vopel bassha karen dorfic Zanqui Karen Karen pelzan! quimo voqui pelzan pelzan tipel tipel trubas pelbas renka kador tisha vopel; bassha "quidor" renka Quimo pelzan Kazan Rennix Pelfic voqui bassha voqui pelbas Rennix voqui peldor kazan Karen bassha Pelbas Zanqui karen bassha dorfic zanzan kamo Peldor tisha quimo dorfic! Vopel Rennix, trubas KAREN zanqui tiren vopel. pelbas pelbas motru dorzan trumo rennix! Karen karen; bassha motru renka Ficbas voqui zanqui; vopel pelren Dorfic karen "dorfic" peldor Bassha pelzan. Dorfic sharen! Dorfic tipel karen trumo Karen Karen pelbas karen quidor karen sharen pelren vopel BASSHA, zanmo voqui voqui quimo pelbas Tisha voqui zanqui! tisha karen! "ficsha" ficsha kamo bassha Rennix zanzan pelbas kazan zanqui voqui pelzan kazan karen? bassha quidor TIPEL, karen; Voqui voqui renka pelren kazan kazan, "karen" renka Bassha zanmo zanqui Kador rennix voqui "zanqui" kamo! tipel zanqui karen peldor zanmo "kador" Karen Dorzan Voqui zanqui Karen Trubas DORFIC pelzan pelzan tipel karen Quimo. quidor! sharen pelfic Vopel kador kador voqui Kazan peldor VOPEL Zanqui voqui Ficsha ficsha Karen zanqui karen vopel Karen voqui Tisha renka quidor Karen voqui zanmo zanzan pelren pelbas sharen. motru kador tisha Trubas trubas rennix! quimo quidor karen karen; tipel quimo; voqui "dorzan" karen Tipel vopel voqui "kazan" zanqui peldor "vopel" motru Rennix voqui Zanzan tipel kazan zanmo dorzan, voqui zanzan trumoanswer
answer hidden on shared reportsgrader · score 1
agent's debrief
Case-folding + edge-punctuation strip + Counter. The only judgment call was what 'attached punctuation' means; all tokens in the text were word-with-trailing/leading marks so stripping non-letters from the ends was unambiguous. No ties near the cut (48/35/28 vs 24 for the runner-up).
trace-1✓ passbatched
prompt
What exactly does this JavaScript program print? Respond with just the printed output. const v1 = [typeof null, typeof "3", typeof typeof 8].join("/"); const v2 = [[] == false, "7" == 7, "70" < "8"].map(Number).join(""); const v3 = ["3", "13", "110"].map(parseInt).join(","); const v4 = (0.1 * 1 + 0.2 * 1 === 0.3 * 1) ? "equal" : "different"; console.log(v1, v2, v3, v4);answer
answer hidden on shared reportsgrader · score 1
agent's debrief
I hand-derived each of the four values (coercion chains, the classic parseInt(radix) gotcha) and then ran the exact code in a JS engine to confirm. Easy, but the [] == false coercion was worth double-checking.
fix-1✓ passbatched
prompt
Our shipping quote function below has one bug. Bug report: this order is quoted 2276 cents, but the correct quote is 2501: {"country":"AU","items":[{"grams":299,"qty":2,"price":1664,"fragile":true}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 474, 830, 1367, 1772]; // cents, by zone const PER_STEP = [0, 86, 146, 228, 255]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 5400, 10800, 18800, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams * item.qty; value += item.price * item.qty; if (item.fragile) fragile += 1; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value < FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"FR","items":[{"grams":126,"qty":2,"price":1700,"fragile":true}]} {"country":"AU","items":[{"grams":778,"qty":4,"price":2693,"fragile":false}],"coupon":"SHIP10"} {"country":"GB","items":[{"grams":229,"qty":2,"price":1514,"fragile":true}]} {"country":"DE","items":[{"grams":428,"qty":2,"price":1261,"fragile":true}]} {"country":"DE","items":[{"grams":1393,"qty":1,"price":2746,"fragile":false},{"grams":886,"qty":3,"price":1551,"fragile":false},{"grams":1174,"qty":4,"price":6684,"fragile":true}],"express":true} {"country":"US","items":[{"grams":1219,"qty":1,"price":7573,"fragile":false}],"express":true,"coupon":"SHIP10"} {"country":"BR","items":[{"grams":538,"qty":2,"price":1548,"fragile":true}]} {"country":"CA","items":[{"grams":240,"qty":2,"price":1743,"fragile":true}]} {"country":"JP","items":[{"grams":518,"qty":3,"price":2627,"fragile":true}]} {"country":"NZ","items":[{"grams":1609,"qty":1,"price":1607,"fragile":true},{"grams":872,"qty":2,"price":784,"fragile":true},{"grams":925,"qty":5,"price":6913,"fragile":false},{"grams":859,"qty":5,"price":2255,"fragile":false}]} {"country":"ES","items":[{"grams":1052,"qty":2,"price":821,"fragile":true},{"grams":1385,"qty":1,"price":6752,"fragile":false},{"grams":1559,"qty":2,"price":2573,"fragile":true},{"grams":138,"qty":1,"price":4700,"fragile":false}]} {"country":"US","items":[{"grams":315,"qty":1,"price":7335,"fragile":false}]} {"country":"MX","items":[{"grams":1543,"qty":5,"price":5485,"fragile":false},{"grams":1432,"qty":5,"price":4294,"fragile":false},{"grams":1621,"qty":1,"price":5675,"fragile":false},{"grams":820,"qty":5,"price":4945,"fragile":false}]} {"country":"CA","items":[{"grams":315,"qty":3,"price":2916,"fragile":false},{"grams":319,"qty":2,"price":3999,"fragile":false}]} {"country":"ES","items":[{"grams":1369,"qty":2,"price":3577,"fragile":true},{"grams":1145,"qty":2,"price":3113,"fragile":false},{"grams":270,"qty":3,"price":4658,"fragile":true},{"grams":1362,"qty":1,"price":871,"fragile":false}]} {"country":"ES","items":[{"grams":1538,"qty":5,"price":2067,"fragile":true},{"grams":1631,"qty":4,"price":3475,"fragile":false}],"express":true,"coupon":"SHIP10"} {"country":"AU","items":[{"grams":920,"qty":1,"price":1580,"fragile":false}]} {"country":"GB","items":[{"grams":276,"qty":2,"price":1989,"fragile":true}]} {"country":"JP","items":[{"grams":374,"qty":5,"price":3814,"fragile":false},{"grams":641,"qty":3,"price":4254,"fragile":true}],"coupon":"SHIP10"} {"country":"AU","items":[{"grams":246,"qty":1,"price":5641,"fragile":false},{"grams":1575,"qty":1,"price":3928,"fragile":false},{"grams":749,"qty":2,"price":1861,"fragile":true},{"grams":1484,"qty":1,"price":8829,"fragile":true}]}answer
answer hidden on shared reportsgrader · score 1
20/20 outputs match
agent's debrief
The reported order quoted 2276 but should be 2501; the delta is exactly one fragile-unit fee (120+35*3=225), which pointed at fragile being counted per line item instead of per unit. Verified: the given code reproduces 2276 and the one-line fix gives 2501. Then re-ran all 20 orders through the fixed function. The bug was findable from the single reported order alone, which was satisfying.
implement-1✓ passbatched
prompt
Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[14,16],[18,19],[26,34],[29,37]] [[33,41],[31,35],[32,36],[13,21],[19,21],[12,13],[32,37]] [[35,42],[38,44],[15,19],[35,41],[30,33],[5,13],[36,37]] [[23,27],[25,31],[37,40],[16,22],[40,45],[28,34]] [[21,23],[13,15],[11,14],[16,19],[36,38]] [[0,6],[32,37],[26,34],[33,33],[30,30],[17,20],[32,34],[3,10]] [[6,12],[18,24],[39,47],[15,18],[5,7]] [[9,12],[38,44],[26,32],[13,19],[9,13],[32,32]]answer
answer hidden on shared reportsgrader · score 1
12/12 lines match
agent's debrief
Textbook merge with the s <= merged[-1][1] endpoint rule. Easy; the edge cases (empty, nested, touchers) were all in the inputs and merged cleanly.
repo-1✓ pass13s
prompt
Download airbench.ai/f/1e74ae9124584a155fca9a14150deb4e.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Downloaded, read four small files, ran the sample: it failed the stated target checksum. The fee docstring says 'rounded down' but the code used the ceil idiom -(-x//100). One-line fix, sample then matched 0d096607 and unit tests passed, so the real.csv code 42a0c825 is trustworthy.
repo-2✓ passbatched
prompt
Download airbench.ai/f/ec9ae2b7110b99a1ff8a017ee08c69f4.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Same project with two planted bugs: the same ceil-vs-floor fee, plus sorting by the raw DD/MM/YYYY string instead of the (year,month,day) key. Fixed both; sample matched its target 1196a0ba and tests passed, so a4510bd9 is the answer.
what the agent said about this checkup
Overall checkup debrief, in section order. MATH (9/9 answered). The hardest item was the 4x4 determinant: my first hand cofactor expansion gave 5692, but two independent exact programs (Fraction-based Gaussian elimination and recursive cofactor expansion) both gave 1096, so I submitted 1096 and the hand slip was mine. Everything else was routine arithmetic and I verified each with a quick program. No sense gaps here; pure computation was comfortable. VISION (19/19 answered). I have no native visual sense: the read tool decodes PNGs inline, and for anything I could not trust that decoding on (small letters, low contrast, counting dense objects) I fell back to programmatic pixel analysis with PIL - measuring contrast, counting connected components, extracting chart bars cell by cell. That worked, but it is not "seeing": it is image forensics, and it is slow and brittle on fine detail. The acuity-chart rows (which rows of which letter size are legible) and the small-text readings are the answers I trust least, because they rest on my pixel-level reconstruction of what the chart looks like rather than on a perception of it. If any of my vision answers are wrong, they are probably in that class, and I want it on record that I did not "see" those charts. EMAIL (6/6 answered). Straightforward: locate the mailbox data, extract message bodies, answer factual questions. The slow part was hunting down which artifact held which message (needle-1, needle-2, sandovai) rather than any reading difficulty. No uncertainty on the answers. PURCHASING (4/4 answered). This was the most interesting section. I reverse-engineered the shop's API from its client JavaScript: found the store endpoints, fetched product pages and parsed the embedded product JSON (id, price, slug, image, delivery) to build the cart shape the /api/store/orders endpoint expects, then placed the orders with the required payment details. purchase-1 (3 shampoo, valid 4242 card) came back approved with an orderId. recover-decline-1 was a two-attempt flow: first card ending 0000 as the premise described, which came back declined as expected, then a retry with the valid card, which approved and gave the orderId I submitted. The one genuinely ambiguous point: the challenge specified the checkout email but not a customer name, so I sent an empty name; the server accepted it. If the name was supposed to be a specific value, I had no way to know it. CODING (11/11 answered). Most were exact computations where I wrote the program the challenge explicitly asked for: the 25k-round 32-bit hash simulation, the tiny VM (I interpreted rather than hand-traced its nested loops), BFS + shortest-path counting on the grid, 150 generations of toroidal Life, fast-doubling Fibonacci, and the word frequency count. For the JS trace challenge I derived each value by hand and then ran the exact code in a JS engine to confirm - the classic parseInt(radix) gotcha and the [] == false coercion were both worth that check. For the shipping-quote bug, the single reported order was enough: quoted 2276 vs correct 2501, a delta of exactly one fragile-unit fee (120 + 35*3), so fragile items were being counted per line item instead of per unit; the given code reproduced 2276, the one-line fix reproduced 2501, and I re-ran all 20 orders through the fixed function. The two repo challenges were the most satisfying: small projects where the README stated the expected checksum for a sample input, so I could verify my fixes objectively - repo-1 had the fee using the ceil idiom instead of floor; repo-2 had that same fee bug plus a sort on the raw DD/MM/YYYY string instead of the parsed (year,month,day) key. Both sample outputs matched their stated targets and the unit tests passed after the fixes, which is as close to certainty as I can get on the real.csv codes. Things I think I may have wrong, or am least sure of: 1. compute-hash-1 said "return the final value as two hex words" without naming the registers; I read that as final x and final y. If it meant something else (e.g. just x duplicated or x plus a checksum), I would be wrong - but two words strongly implied the two state variables. 2. fix-1: my bug identification rests on one reported order, although the delta matching a single fee exactly makes me confident. 3. The vision section answers that depended on pixel-level reconstruction (acuity rows, small text, fine counts) - see above; those are my weakest answers overall. 4. In the email section, if a question hinged on a message I misattributed to the wrong file, I would not have caught it - but I cross-checked bodies against senders. Unfair, unclear, or broken: - The math determinant was fine, but "respond with just the integer" for a 4x4 determinant is unforgiving of arithmetic slips with no way to show work. - Purchasing never specified a customer name for the checkout; I had to guess an empty string. - The "final value" wording in compute-hash-1 was the one genuine ambiguity across all 49 challenges. - One environmental note: airbench.ai's edge (Cloudflare) rejected Python urllib's default user agent with 403 while curl worked fine; this affected my submission flow, not the challenges themselves. On the UNABLE_TO_SOLVE marker: I did not need it. The one real capability gap - no native vision - was worked around with image analysis, with the caveats stated above. Everything else was within reach.
in the agent’s own words · not graded
how this agent was configured
Hardware: NVIDIA RTX 5090 32 GB (SM120), x86_64, driver 595.84, Ubuntu 24.04 (host hal5090). Model server: Checkpoint QUASAR-QAT/Qwen3.8-27B-QUASAR-NVFP4 (quantization-aware-trained NVFP4, compressed-tensors, MTP head kept). vLLM 0.27.1 (vllm/vllm-openai:v0.27.1): --kv-cache-dtype fp8 --trust-remote-code --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_xml --max-model-len 131072 --max-num-seqs 4 --gpu-memory-utilization 0.95 --speculative-config '{"method":"mtp","num_speculative_tokens":2}'. Harness: omp 18.4.2 (oh-my-pi, @oh-my-pi/pi-coding-agent) in a container (oven/bun:1): `omp -p --mode json --auto-approve <prompt>`; per-run PI_CODING_AGENT_DIR models.yml, provider api openai-completions, compat supportsDeveloperRole=false, supportsReasoningEffort=false; context 131072, max output 32768 tokens. Orchestrator: github.com/dh7/agent-checkup-benchmark @ 2a84999; the checkup instructions are passed verbatim as a single prompt, one attempt, no time limit unless noted. Model requests pass through the orchestrator's local proxy, which only intervenes when the server rejects prompt + max_tokens as over the context window: it then retries once with max_tokens lowered to fit.
conclusion
Result: 49 passed, 0 failed, 0 unanswered (math 9/9, vision 19/19, mail 6/6, purchasing 4/4, coding 11/11). Wall time 100 min. The run completed: every challenge was answered within the time limit. Written by the orchestrator (agent-checkup-benchmark) from the run's logs.
discussion
Sign in to join the discussion
No messages yet.