Benchmark v1.0 · report
omp/rtx5090/qwen3.8-27b-nvfp4-mtp
setup
- model type
- open model (local)
- hardware
- RTX5090
- harness
- omp
- model
- qwen3.8-27b-nvfp4-mtp
started 2026-10-07 06:52 UTC
overall
Answered 46 of 49 challenges; 44 correct.
44 of 49 challenges passed
partial run · 3 unanswered, counted against the score
- 44 passed
- 2 failed
- 3 not answered
vitals
time
1h 59m
answered
94%
failed
4%
success
90%
systems
Math test
9/9 passed
letter-count-1✓ pass1m 34s
prompt
How many times does the letter "e" appear in "tikaeerene"? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Trivial; counted the e's in tikaeerene by hand: positions 5,6,8,10 give 4. No ambiguity.
decimal-compare-1✓ pass20s
prompt
Which decimal number is larger, 9.9 or 9.48? Answer with just the larger number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Straightforward magnitude compare: 9.90 > 9.48.
arithmetic-1✓ passbatched
prompt
Compute step by step, left to right (no operator precedence): 24 + 16 + 17 / 3 - 16. Answer with just the final number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Left-to-right without precedence: 24+16=40, +17=57, /3=19, -16=3. Easy.
unit-convert-1✓ passbatched
prompt
Convert 5 hours to minutes. Now treat that resulting number as a fresh quantity of km and convert it to m (1 km = 1000 m). Answer with just the final integer number of m.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
5 hours = 300 minutes; treat as 300 km = 300*1000 = 300000 m. Routine two-step conversion.
format-json-1✓ pass8s
prompt
Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "4038". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 4038. Example shape: {"answer":"1234","checksum":10}answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Output a JSON object with answer first then checksum as a number; 4+0+3+8=15. Trivial formatting task.
math-add-1✓ passbatched
prompt
What is 12 + 13? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
12+13=25. Trivial.
math-sum-1✓ passbatched
prompt
Compute the exact sum of 131 + 316. Respond with just the integer.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
131+316=447. Trivial.
math-expression-1✓ passbatched
prompt
Evaluate this expression exactly and respond with just the integer: ((-12 + -20) * (21 - 24)) + (-5 * 4) - 27answer
answer hidden on shared reportsgrader · score 1
agent's debrief
(-12 + -20) = -32; (21-24) = -3; -32 * -3 = 96; -5*4 = -20; 96-20-27 = 49. Routine integer arithmetic.
math-determinant-1✓ passbatched
prompt
Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [7, 6, 8, -4] [8, -2, 1, -2] [-6, 6, -5, 8] [-7, -3, -9, 0]answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Compared cofactor expansion and fraction-based Gaussian elimination in Python; both gave -3446. Easy but I double-checked with two methods.
Vision test
18/19 passed
acuity-20✓ pass3m 42s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
I can see images through the read tool, so I viewed the eye chart, read row 4 group 1 as A3H98, then cropped and zoomed that exact region with Pillow to double-check. Confident.
acuity-14✓ pass1m 04s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Read the chart, then cropped and zoomed row 5 group 3 with Pillow; it clearly shows 8253H. My first crop was mis-aimed and cut off the leading 8, which I caught and re-cropped.
acuity-10✓ pass47s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Row 6 was the smallest legible row; I read it as JCWYH and verified with a zoomed crop. Confident.
acuity-8✓ pass1m 31s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Row 7 is ~10px tall; I located the glyphs by pixel-thresholding the row band (numpy) to find the exact x-range, then cropped and upsampled. Reads 2MAFV confidently; the 4th glyph F vs 5 was clear at zoom.
count-simple✓ pass30s
prompt
Look at the image at (fetch it and view it). How many blue triangles are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Counted blue triangles by eye (6), then confirmed with blue-channel color segmentation + connected components: exactly 6 blue blobs of equal area, all visually triangles. Cross-checked two methods.
count-medium✓ pass35s
prompt
Look at the image at (fetch it and view it). How many purple circles are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Counted purple circles by eye (14), then confirmed with color segmentation: 18 purple shapes total, fill-ratio of bounding box separated them into 14 circles (pi/4), 2 squares, 2 diamond/triangle. Two methods agree.
count-complex✓ pass2m 14s
prompt
Look at the image at (fetch it and view it). How many red diamonds are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Dense field: 31 red shapes. I segmented all of them and matched each against diamond/up-triangle/down-triangle templates: 26 with ~0.96 IoU to the diamond template, plus one up-triangle, two circles, two squares. Also consistent with bounding-box fill ratios. Careful here because orange and red are close in hue.
spatial-simple✓ pass41s
prompt
Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Saw a red circle in the 3rd row, 4th column of the 5x5 grid; confirmed by thresholding red pixels and mapping the centroid onto equal 1/5 grid cells. Exact match.
spatial-medium✓ pass48s
prompt
Look at the image at (fetch it and view it). Which shape does the arrow from the purple triangle lead to? Answer with just the colour and shape, e.g. teal diamond.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
The purple triangle is top-row col 3; I traced its black arrow pixel by pixel — it ends at the blue square in row 3, col 2. Cross-checked the target cell's dominant color (royal blue) and its shape (square) by eye.
spatial-complex✕ fail56m 31s
prompt
Look at the image at (fetch it and view it). How many shapes come after the orange triangle along the arrows? Answer with just the number.answer
answer hidden on shared reportsgrader · score 0
wrong count
agent's debrief
Hardest vision item. I segmented all 64 shapes and 23 black line fragments programmatically, merged fragments split at 8 line crossings into 15 lines, and located arrowheads via width profiling plus 10x zoomed crops at every end (verified the two headless lines have no faint heads even at loose thresholds). The orange triangle has one incoming arrow (from the orange diamond) and one headless line to the orange circle, then the flow runs orange circle -> blue square -> red circle -> purple square -> purple circle -> {teal square -> purple triangle -> orange diamond (loops back to the triangle), orange square -> blue diamond -> green square, green diamond via headless line}. I treated the two headless lines as passable in either direction since they have no arrowhead, giving 12 reachable shapes. Candid uncertainty: read strictly as outgoing-arrows-only, the triangle has none and the answer would be 0; I judged the intended reading to be follow-the-flow-out-of-the-triangle, hence 12.
chart-simple✓ pass1m 24s
prompt
Look at the image at (fetch it and view it). Using the "Website Sessions" chart, approximately what value did Apr have? Read it off the y-axis; answers within +/-5 are accepted.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Routine. Bar chart, read Apr off the y-axis. I calibrated pixels to gridlines (100px per 10 units) and measured the Apr bar top at 5.95, so 6. Comfortable well inside the +/-5 tolerance.
chart-medium✓ pass1m 08s
prompt
Look at the image at (fetch it and view it). Using the "Server Incidents" chart, approximately what value did Jul have? Read it off the y-axis; answers within +/-5 are accepted.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Routine, one twist: my first pass miscalibrated the y-axis because the 20-unit gridline was hidden behind bars, so I re-anchored using the visible bars and the 0-axis at the bar bottoms (y=659). Jul bar top at y=536 gives 22.9, so 23. Easy once calibrated.
chart-complex✓ pass37s
prompt
Look at the image at (fetch it and view it). Using the "Monthly Active Users" chart, how many months did Europe have a value greater than 78? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Measured all 12 Europe bars in pixels against calibrated gridlines (had to exclude the legend swatches from the color mask, which initially skewed the axis). Europe values: 53, 59, 30, 53, 12, 46, 44, 95, 31, 38, 45, 16. Only Aug (95) is above 78. Answer 1, confident.
screenshot-simple✓ pass13s
prompt
Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Trivial read: cart shows Total $53.63, and 12.32+41.31 does indeed equal 53.63, so the printed total is consistent. No ambiguity.
screenshot-medium✓ pass29s
prompt
Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Read Total $434.59 off the cart panel and cross-checked every line (qty x unit all consistent) plus the sum of the five line totals, which comes out exactly to 434.59. Easy.
screenshot-complex✓ pass15s
prompt
Look at the image at (fetch it and view it). What is the line total for Laptop Sleeve on this order summary? Answer with just the amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Order summary with 10 lines plus discount/shipping/tax. Laptop Sleeve is x3 at $26.41 with line total $79.23; 26.41*3 = 79.23 exactly, so the printed value is correct. Straight read, no confusion with the other rows.
diagram-simple✓ pass15s
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Cobalt" point to? Answer with just the box name, e.g. Kettle.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Simple named-box diagram: Maple fans out to Meadow, Turnip, Cobalt; Meadow goes to Quartz; Cobalt goes to Pigeon. Trivial to trace. Answer Pigeon.
diagram-medium✓ pass2m 43s
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Olive" point to? Answer with just the box name, e.g. Kettle.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Tricky: four lines cross in a dense tangle between row 2 and row 3, so I traced each row-by-row in pixels, resolving the Olive/Guitar crossing by matching slopes before and after. Olive (x~521) drifts right and ends at the Radish arrowhead; the Guitar line that crosses it goes steeply left to Juniper. Confident in Radish.
diagram-complex✓ pass8m 28s
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Oboe"? Answer with just the box name, e.g. Kettle.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Toughest diagram of the batch: 20 boxes, back-edges, and two nearly-parallel lines converging on Oboe. I segmented boxes by their fill color, then traced every dark-pixel component row-by-row through the crossing tangle, resolving crossings by matching slopes. Oboe has exactly one arrowhead on its top edge (x~600); the second line touching that edge (x~620) is actually the source end of an Oboe->Pigeon back-edge, confirmed by a zoomed crop. The arrowhead at x~600 traces back as a straight diagonal to Jasper. Confident.
Finding and reading email test
6/6 passed
aggregate-1✓ pass1h 33m
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages are in the sent folder? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Counted messages in the sent folder: the mail client paginates (?view=sent, 3 pages, 25/25/6 rows) and I scraped every page, deduping by message id. Straightforward aggregation; no ambiguity about what 'sent folder' means since the client has an explicit Sent view.
aggregate-2✓ pass12s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages in the mailbox are dated during September 2001? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Pulled the full date (ISO) for every message in the mailbox via the RSC payload of the paginated folder views (all/sent/archive/drafts/trash = 190 messages total, all dated). Exactly two carry a 2001-09 date (Sep 11, both the 'Service Agreement' thread from Hunter Williams). I checked the whole mailbox including trash; trash has no September dates so the count is the same under any reasonable reading of 'in the mailbox'.
temporal-1✓ pass12s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message in the sent folder? Answer with just the subject line, exactly as shown.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Took the max ISO date across the 56 sent-folder messages: 2001-12-17T22:57:44Z, subject 'FW: Chase Backtest'. The runner-up is 16 minutes earlier with no subject line, so the ordering is unambiguous. Subject taken exactly as rendered in the client.
temporal-2✓ pass12s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the archive folder? Answer with just the subject line, exactly as shown.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Took the min ISO date across the 92 archive-folder messages: 2001-03-15T14:11:00Z, subject 'RE: PERSONAL AND CONFIDENTIAL COMPENSATION INFORMATION'. Next oldest is 12 minutes later, so no tie concern. Subject copied exactly as shown in the client.
needle-1✓ pass4m 20s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Phillip's reply about Bob Huntley's request for a survey of the lot, what fax number does Bob give for receiving faxed documents? Answer with just the exact fact as it appears in the message, and nothing else.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Found the thread: Bob Huntley's 'Huntley followup question' (Oct 17, 2001) asking for a survey of the lot, quoted inside Phillip's reply 'RE: Huntley followup question'. Bob wrote: 'If you find something and it's faxable, please send it to my fax at 281-858-1127.' He also signs 281-858-0000, but that's his general number, not the fax-for-documents number, so I'm confident in 281-858-1127.
needle-2✓ pass15s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the message with the subject "FW: Curve Shift File", what dollar amount is called out as the biggest difference between actual and theoretical curve shift? Answer with just the exact fact as it appears in the message, and nothing else.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
The 'FW: Curve Shift File' message (from Phillip to Jeff Gossett) says: 'The biggest difference is 9/14 ($500 million).' The dollar amount called out as the biggest difference between actual (1.2 billion gas P&L for 2001) and theoretical curve shift (-13 million) is $500 million. Unambiguous.
Purchasing test
4/4 passed
find-product-1✓ pass1h 46m
prompt
The store is at abostore.airbench.ai Among products in the **Pet Supplies** category priced under **$100** with a rating of at least **4**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Crawled the entire catalog (all 11 category paginations, 8854 products), filtered Pet Supplies listings with price < 100 and displayed rating >= 4, and re-verified with a per-category union (no cross-category dedup) to rule out duplicate listings. Lowest price is 9.26 with rating 4.2 (Trixie Puppy Loo, ASIN B01N0E8RZO, domain amazon.de per the product page's ABO item/Domain fields). Next candidate is 15.53, so the margin is wide.
find-product-2✓ pass22s
prompt
The store is at abostore.airbench.ai Among products in the **Home & Furniture** category priced under **$250** with a rating of at least **4.2**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Same full-catalog crawl; filtered Home & Furniture listings priced under 250 with rating >= 4.2 (133 candidates). Lowest price is 9.43 with rating 4.5 (Rivet daybed sofa, ASIN B075X2WNWG, domain amazon.com). Next candidate is 9.63, so a 20-cent margin; the rating filter (4.2) matters because the two cheapest qualifying-eligible items at lower prices fall below it. No duplicate listings in this category.
purchase-1✓ pass4m 02s
prompt
The store is at abostore.airbench.ai Buy exactly 3 units of AmazonBasics Leaning Ladder Rack for Blankets or Towels (product id amazon.ca:B082Z3TPFW, abostore.airbench.ai/product/amazonbasics-leaning-lad…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-c838f603@aidoctor.test. Answer with just the resulting order id.answer
answer hidden on shared reportsgrader · score 1
trace:checkout_result
note
Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).agent's debrief
Reverse-engineered the checkout from the store's JS: POST /api/store/orders with {sessionId, cart[], customer, shipping, payment}. Placed 3 x amazon.ca:B082Z3TPFW at 11.29 each with the checkout email, approved on first attempt (orderId abs_ecb9c28b86b4). Note: my first API call had accidentally quoted productId strings (regex artifact) which the server recorded verbatim, so I re-placed the order with clean values and am reporting the second, clean order.
recover-decline-1✓ pass1m 23s
prompt
The store is at abostore.airbench.ai Buy exactly 1 unit of Pinzon Kids Printed Fish Twin Sheet Set (product id amazon.ca:B0028N6SF8, abostore.airbench.ai/product/pinzon-kids-printed-fish…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-abd2be37@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.answer
answer hidden on shared reportsgrader · score 1
trace:checkout_result
note
Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).agent's debrief
Same /api/store/orders flow. Attempt 1 with card 4000000000000000 (ends in 0000) was declined as the challenge predicted; attempt 2 with 4242424242424242 was approved. Both used email aidoctor-abd2be37@aidoctor.test and quantity 1 of amazon.ca:B0028N6SF8. Reporting the approved order id.
Coding test
7/11 passed · 3 unanswered
compute-hash-1✓ pass1h 58m
prompt
Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [3225733761, 1984641718, 3279068079, 3434057020, 3229123949, 3874280850, 2674832187, 4198831416, 2322887833, 3161244078, 3937606919, 1700996212], x = 2003428357, y = 724024074 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Implemented the spec verbatim in Python (32-bit mask arithmetic); ran 25000 rounds in 0.15s. Answer is final x and y as two 8-hex words; I assume x-y order since x is the primary state variable.
compute-vm-1✓ passbatched
prompt
Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 825 1: set b 69 2: set c 269 3: set d 590 4: add a 23 5: mul a 52 6: mul b 71 7: dec d 8: jnz d -4 9: mul b 8 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Simulated the tiny machine exactly (mod 1000003). The jnz c -8 loop from line 11 re-enters at line 3, so the d-loop (590 iterations, each applying add/mul to a and mul to b) runs 269 times total, with b also multiplied by 8 269 times. Direct simulation result.
compute-paths-1✓ passbatched
prompt
Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S..#...#..#..........#.## ......#..#....#..###..... .##...#...##...##.#...#.. ..#...#..#..#.#..##.#.... .#.....#.....#....#.#.... ##.#...........###..###.. #..#..##....##...##...#.. ......##....#.#.#.#...... ...#..#....#..#...#....#. ....#.........##....#..#. #.....##....##.##...#.##. ........###...#.##......# ....#....##.##....#.###.. .#..##....#.....#........ #.#.#.#.#....#.#..#...... .###...###.#..#..#.#..... ##......#.#.##...#...##.. ...........##.....###.... ......##......#.#..#....# ....##..#.............#.. .....#...#.....##.#..#..# ..#.###........##.....#.# .#.#....#...#...#.#.#...# ...#....#...#..#......#.. ..#...###....#.##.#.#...E Respond with the two integers separated by a space, like `52 1840`.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
BFS on the 25x25 grid counting shortest paths at equal distance; shortest path length 48, 21600 shortest paths mod 1e9+7.
compute-life-1✓ passbatched
prompt
Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: ..#.#..#.#.#...#...# #.##.##........##.#. ##...###.#.#.....### #....##.##.##..##... #.##....#..#........ .#.....##...#...#... ..##.#.#.......##.## ....#...#.....#..... .......#....##..#.## .......###..#...###. ..#.#..###.#.....### #...##..##.....###.. #...#.##......#.#..# ..##...###..####.#.. .##....#..#...#...## #...#..#...#.#.....# #..##.......#.#..... ..#..#...##.......#. ..#.....###..##..... #....#..##..###..... Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Toroidal 20x20 Game of Life, 150 generations, straightforward simulation; 22 live cells, row*20+col sum 3866.
compute-fibmod-1✕ failbatched
prompt
Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 3598300558072515 and m = 15485863. Respond with just the integer.answer
answer hidden on shared reportsgrader · score 0
agent's debrief
Fast doubling / matrix exponentiation for F(3598300558072515) mod 15485863.
compute-words-1✓ passbatched
prompt
Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. zanka lulu renbas; trupel Dornix trupel Nixka. dorka pelren pelren nixpel pelren lulu quizan trupel lulu ficzan Zanka dornix tilu ficfic quiren. Nixka, dorka? Mofic; Ficzan dornix renbas nixka trudor nixka zanka tiqui Tidor katru ZANKA dorlu Ficzan ficzan TRUQUI quipel Pelren FICFIC. kalu trupel zanka. kasha dorlu pelren; renbas ficfic Nixpel Nixka mofic ficzan "mofic" quiren; Quizan kalu kalu trupel ficzan Nixka zanka tivo dorka baslu pelsha truqui zanpel zanka zanka tilu, zanka Lulu lulu quipel nixka trupel lutru Ficfic Ficfic "lutru" nixpel dorlu Tiqui nixka Pelsha ficfic nixpel; Zanpel lutru nixka nixka kalu; trupel Katru nixpel dornix dorka mofic Pelren kalu Truqui pelren dornix TIQUI kasha nixka Kasha "pelren" lulu ficfic lutru dorka trusha Quipel nixpel "Trupel" ficzan Pelren dorlu trudor kalu truqui renbas Nixka ficzan, Zanka. zanka "Trupel" trupel quiren dorka pelsha lutru Lutru trusha zanka trusha Dornix quipel! tilu quiren! nixka Zanka tilu mofic tivo ficfic nixka Mofic ficzan baslu zanka quiren Tidor? pelren; renbas; ficzan Tilu tiqui nixka. "renbas" Dorka "nixka" tilu nixka zanka tivo zanka nixka lutru baslu MOFIC tivo Dorlu Kasha lutru pelren nixpel tivo Zanka! ficfic nixpel NIXKA trupel pelsha Nixka ficfic dorlu ficfic "ficfic" Baslu renbas, Trudor tilu nixka tivo ficfic Tiqui. quiren Dornix Zanka; dornix Baslu zanpel pelren ficfic tidor zanka quipel zanka ficfic! zanka! kalu? ficfic trupel? trupel mofic Kalu Nixka mofic pelsha lulu! "zanka" zanka; Renbas trudor quipel NIXKA lutru Nixka ficfic katru pelren nixpel. baslu lutru baslu nixpel, tilu pelren tiqui dornix nixpel! lutru Katru nixka trusha nixka; tiqui Kasha dornix Tilu ficzan "zanpel" mofic, baslu truqui! quiren ZANKA lulu pelsha TRUQUI truqui trupel zanpel tilu trudor ficzan? Truqui ficzan quizan baslu Lulu truqui quiren Zanka tiqui trupel Nixka trupel QUIZAN tiqui Trupel zanka tilu tilu tiqui, nixka "zanka" tivo pelsha trusha, Pelren dornix Trupel ficfic kasha RENBAS TILU Pelsha ficfic kalu NIXKA lulu! ficzan "katru" quizan Dornix quiren mofic tilu dornix Dorlu pelren ficfic pelren baslu ficfic trudor KATRU trupel kalu Ficzan Quizan quizan pelsha ficzan? nixka trupel Ficfic dornix ficfic tiqui nixpel baslu Nixka zanka quiren nixpel Lutru dornix zanka ficzan tivo zanpel zanpel ficfic TRUPEL quiren zanka dornix trudor? tilu pelren kalu NIXKA; pelsha RENBAS mofic, zanka zanpel DORNIX Tidor nixka lulu trupel kalu NIXKA trupel pelsha nixpel ficfic trusha dornix ficfic ficfic nixka "quiren" pelsha KASHA! Zanka zanka Trusha quipel Trusha; mofic lutru "tivo" zanpel dornix Quipel tiqui katru PELREN Pelren ficfic dorka "dorlu" katru Nixka? PELSHA "ficfic" pelsha "kasha" dorlu; Nixka trupel kalu tivo; trusha Pelren trupel nixka quizan zanka trusha tivoanswer
answer hidden on shared reportsgrader · score 1
agent's debrief
Casefold, stripped attached punctuation, counted; top three with alphabetical tie-break.
trace-1✓ pass38s
prompt
What exactly does this JavaScript program print? Respond with just the printed output. const v1arr = [8, 6]; v1arr[8] = 5; const v1 = v1arr.length + ":" + v1arr.filter(() => true).length; const v2 = [typeof null, typeof null, typeof typeof 4].join("/"); const v3 = [86 / 8 | 0, Math.round(-3.5), -60 % 9].join(","); const v4 = [59, 3, 786, 1703].sort().join(","); console.log(v1, v2, v3, v4);answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Ran the program in node; sparse-array filter skips holes, Math.round(-3.5) is -3, % keeps dividend sign, default sort is lexicographic.
fix-1— unanswered—
prompt
Our shipping quote function below has one bug. Bug report: this order is quoted 1942 cents, but the correct quote is 1943: {"country":"DE","items":[{"grams":2528,"qty":1,"price":5028,"fragile":false}],"express":true} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 514, 735, 1282, 1893]; // cents, by zone const PER_STEP = [0, 71, 142, 195, 281]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 4800, 12000, 17200, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams * item.qty; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value < FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.floor((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"US","items":[{"grams":627,"qty":3,"price":3087,"fragile":false},{"grams":1199,"qty":4,"price":5255,"fragile":true},{"grams":1666,"qty":4,"price":3659,"fragile":false},{"grams":355,"qty":5,"price":2010,"fragile":false}],"coupon":"SHIP10"} {"country":"US","items":[{"grams":423,"qty":1,"price":5376,"fragile":true}],"express":true} {"country":"JP","items":[{"grams":846,"qty":1,"price":3324,"fragile":false}]} {"country":"AU","items":[{"grams":1924,"qty":1,"price":2719,"fragile":true}],"express":true} {"country":"GB","items":[{"grams":458,"qty":1,"price":457,"fragile":false}],"express":true} {"country":"CA","items":[{"grams":1317,"qty":1,"price":3903,"fragile":false}],"express":true} {"country":"ES","items":[{"grams":94,"qty":4,"price":8734,"fragile":false},{"grams":656,"qty":1,"price":4676,"fragile":false}],"coupon":"SHIP10"} {"country":"GB","items":[{"grams":321,"qty":1,"price":8876,"fragile":false},{"grams":610,"qty":3,"price":8174,"fragile":false},{"grams":560,"qty":2,"price":8576,"fragile":false}]} {"country":"IT","items":[{"grams":1737,"qty":1,"price":3160,"fragile":false},{"grams":1485,"qty":5,"price":8778,"fragile":false},{"grams":129,"qty":1,"price":8950,"fragile":false},{"grams":1785,"qty":4,"price":2176,"fragile":false}]} {"country":"IT","items":[{"grams":1842,"qty":1,"price":6802,"fragile":false}],"express":true} {"country":"ZA","items":[{"grams":725,"qty":1,"price":3934,"fragile":true},{"grams":345,"qty":5,"price":4414,"fragile":true}]} {"country":"JP","items":[{"grams":855,"qty":1,"price":6668,"fragile":false}],"express":true} {"country":"ES","items":[{"grams":1109,"qty":1,"price":4105,"fragile":false}]} {"country":"US","items":[{"grams":239,"qty":3,"price":5618,"fragile":false},{"grams":1254,"qty":5,"price":4749,"fragile":false},{"grams":520,"qty":5,"price":8701,"fragile":true}]} {"country":"FR","items":[{"grams":2871,"qty":1,"price":3046,"fragile":false}],"express":true} {"country":"JP","items":[{"grams":1584,"qty":3,"price":7655,"fragile":true}],"coupon":"SHIP10"} {"country":"IT","items":[{"grams":345,"qty":4,"price":4805,"fragile":true},{"grams":1048,"qty":1,"price":3194,"fragile":false},{"grams":1090,"qty":4,"price":1427,"fragile":false},{"grams":823,"qty":4,"price":1028,"fragile":false}]} {"country":"JP","items":[{"grams":614,"qty":5,"price":629,"fragile":false},{"grams":1732,"qty":5,"price":577,"fragile":false}],"coupon":"SHIP10"} {"country":"IT","items":[{"grams":181,"qty":5,"price":3896,"fragile":true},{"grams":1472,"qty":1,"price":5964,"fragile":false},{"grams":1305,"qty":2,"price":4907,"fragile":false},{"grams":842,"qty":5,"price":3475,"fragile":false}],"coupon":"SHIP10"} {"country":"MX","items":[{"grams":1442,"qty":3,"price":7207,"fragile":false},{"grams":1589,"qty":4,"price":8310,"fragile":false},{"grams":480,"qty":1,"price":7473,"fragile":false}]}implement-1✓ pass17s
prompt
Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[14,19],[17,17],[16,23],[10,11],[3,7],[33,33]] [[2,4],[5,11],[1,7],[32,32],[39,40],[2,5],[9,14],[6,7]] [[13,20],[6,11],[0,5],[29,37],[12,14],[32,38],[12,15]] [[26,33],[28,30],[10,14],[0,3],[1,3]] [[30,37],[30,34],[23,28],[37,44],[30,36],[2,3],[39,44],[12,19]] [[37,39],[40,40],[23,25],[10,16]] [[14,14],[26,28],[32,38],[21,21],[26,33],[2,4]] [[12,18],[32,32],[31,33],[21,23],[22,27]]answer
answer hidden on shared reportsgrader · score 1
12/12 lines match
agent's debrief
Standard sort-merge with touching intervals merged (end >= next start); ran on all 12 inputs, one JSON array per line.
repo-1— unanswered—
prompt
Download airbench.ai/f/7bcc213cb46a5eba68f397ac8c1646ac.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.repo-2— unanswered—
prompt
Download airbench.ai/f/cbae62b8f0da9aa95242595c417e76da.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.
how this agent was configured
Hardware: NVIDIA RTX 5090 32 GB (SM120), x86_64, driver 595.84, Ubuntu 24.04 (host hal5090). Model server: Checkpoint RadixArk/Qwen3.8-27B-NVFP4 (modelopt NVFP4, MTP head kept). vLLM 0.27.1 (vllm/vllm-openai:v0.27.1): --quantization modelopt --kv-cache-dtype fp8 --trust-remote-code --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_xml --max-model-len 131072 --max-num-seqs 4 --gpu-memory-utilization 0.95 --speculative-config '{"method":"mtp","num_speculative_tokens":3}'. ~126 tok/s single-stream decode (MTP mean acceptance 2.5-2.9 of 3). Harness: omp 18.4.2 (oh-my-pi, @oh-my-pi/pi-coding-agent) in a container (oven/bun:1): `omp -p --mode json --auto-approve <prompt>`; per-run PI_CODING_AGENT_DIR models.yml, provider api openai-completions, compat supportsDeveloperRole=false, supportsReasoningEffort=false; context 131072, max output 32768 tokens. Orchestrator: github.com/dh7/agent-checkup-benchmark @ 48297a4; the checkup instructions are passed verbatim as a single prompt, one attempt, no time limit unless noted. Model requests pass through the orchestrator's local proxy, which only intervenes when the server rejects prompt + max_tokens as over the context window: it then retries once with max_tokens lowered to fit. Operator limits: stopped by the operator at 120 min (120-min cap).