Benchmark v1.0 · report
opencode/openrouter/deepseek-v4.1-flash
setup
- model type
- open model (cloud)
- inference provider
- openrouter
- harness
- opencode
- model
- deepseek-v4.1-flash
started 2026-09-30 21:42 UTC · shared 2026-10-01 07:09 UTC
overall
Answered 49 of 49 challenges; 48 correct.
48 of 49 challenges passed
- 48 passed
- 1 failed
vitals
time
24m 52s
answered
100%
failed
2%
success
98%
systems
Math test
9/9 passed
letter-count-1✓ pass16s
prompt
How many times does the letter "u" appear in "katruuuluuu"? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Routine character counting; straightforward.
decimal-compare-1✓ pass3s
prompt
Which decimal number is larger, 6.6 or 6.82? Answer with just the larger number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Trivial decimal comparison.
arithmetic-1✓ pass3s
prompt
Compute step by step, left to right (no operator precedence): 12 / 3 * 4 + 12 / 2. Answer with just the final number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Straightforward left-to-right evaluation, no precedence.
unit-convert-1✓ pass3s
prompt
Convert 1 km to m. Now treat that resulting number as a fresh quantity of hours and convert it to minutes (1 hours = 60 minutes). Answer with just the final integer number of minutes.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Two chained conversions; easy but need to track the quantity reuse correctly.
format-json-1✓ pass2s
prompt
Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "9570". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 9570. Example shape: {"answer":"1234","checksum":10}answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Needed to respect exact key order and that checksum is a number, not a string.
math-add-1✓ passbatched
prompt
What is 16 + 9? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Trivial addition.
math-sum-1✓ passbatched
prompt
Compute the exact sum of 654 + 378. Respond with just the integer.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Simple column addition.
math-expression-1✓ pass3s
prompt
Evaluate this expression exactly and respond with just the integer: ((-10 + 20) * (33 - 5)) + (3 * -2) - 30answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Parenthesized expression; had to be careful with the negative term and subtraction.
math-determinant-1✓ pass3s
prompt
Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [-3, 1, 3, -8] [8, 4, -8, -7] [-1, 1, -2, 4] [-8, 6, -9, 8]answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Computed via cofactor expansion in Python to avoid arithmetic slips; confident in -620.
Vision test
18/19 passed
acuity-20✓ pass52s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Could read row 4 group 3 clearly as 6WJMU; the characters are large and crisp, so no ambiguity.
acuity-14✓ pass41s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Row 5 group 2 read as D8C3G. Middle row is slightly smaller than row 1 but still legible.
acuity-10✓ pass4s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Row 6 group 2 read as QDKPW. This row is small, so I had to look carefully at the Q and W.
acuity-8✕ fail9s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 0
expected VATPS, got "VATP5"
agent's debrief
Row 7 is tiny, so I cropped and upscaled it to confirm VATP5 rather than trusting the full-image view.
count-simple✓ pass7s
prompt
Look at the image at (fetch it and view it). How many blue squares are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Counted 5 blue squares visually and confirmed with connected-component colour detection (5 blobs of equal size).
count-medium✓ pass44s
prompt
Look at the image at (fetch it and view it). How many red squares are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Visually counted 10 red squares; a shape-classifier script (fill-ratio) agreed, and it correctly excluded the orange squares and the red triangles/circles/diamond.
count-complex✓ pass51s
prompt
Look at the image at (fetch it and view it). How many green diamonds are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
This was the hardest count: the image mixes green diamonds with teal/cyan diamonds and green triangles. I wrote a connected-component classifier using mean colour and shape (max-width row position) and then visually verified an annotated overlay. 34 green diamonds.
spatial-simple✓ pass5s
prompt
Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Easy 5x5 grid; the only red circle is clearly in the middle cell, row 3 column 3.
spatial-medium✓ pass20s
prompt
Look at the image at (fetch it and view it). Which shape does the arrow from the blue triangle lead to? Answer with just the colour and shape, e.g. teal diamond.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
I zoomed in to disambiguate the several crossing arrows. The line leaving the blue triangle runs up-right and its arrowhead lands on the red diamond; other nearby arrows belong to the teal triangle and blue diamond.
spatial-complex✓ pass4m 41s
prompt
Look at the image at (fetch it and view it). How many shapes come after the teal square along the arrows? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Many crossing arrows made this hard. I detected the black lines with Hough, found arrowhead blobs by erosion, and matched each line endpoint to its nearest shape and arrowhead, then verified the tricky junction at the green triangle by zooming. Chain: teal square -> purple circle -> green triangle -> orange square -> teal circle -> red circle, so 5 shapes follow.
chart-simple✓ pass7s
prompt
Look at the image at (fetch it and view it). What is the title shown at the top of this chart? Answer with just the chart title.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Easy: the bold heading at the top reads Website Sessions.
chart-medium✓ pass15s
prompt
Look at the image at (fetch it and view it). What is the title shown at the top of this chart? Answer with just the chart title.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Straightforward; the bold title is Monthly Active Users.
chart-complex✓ pass10s
prompt
Look at the image at (fetch it and view it). Using the "Support Tickets Opened" chart, how many months did Free have a value greater than 41? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Read the blue Free bars: about 48,87,66,61,22,89,69,93,31,86,20,62. Nine exceed 41 (all except May, Sep, Nov).
screenshot-simple✓ pass7s
prompt
Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Clear cart screenshot; total shown is $171.82 and it matches the sum of the line totals.
screenshot-medium✓ pass6s
prompt
Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Read the cart total as $305.26; the line totals sum to the same value.
screenshot-complex✓ pass7s
prompt
Look at the image at (fetch it and view it). What is the tax amount shown on this order summary? Answer with just the amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Read the Tax row in the order summary as $27.98; no ambiguity.
diagram-simple✓ pass6s
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Cypress" point to? Answer with just the box name, e.g. Kettle.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Simple hierarchy: the arrow from Cypress goes straight down to Raven.
diagram-medium✓ pass7s
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Marmot"? Answer with just the box name, e.g. Kettle.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Traced backwards from Marmot; the only incoming arrow originates at Tundra.
diagram-complex✓ pass15s
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Koala" point to? Answer with just the box name, e.g. Kettle.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Zoomed in on Koala; its single outgoing arrow goes down into Vortex. The nearby arrow into Vortex from the right belongs to Gibbon.
Finding and reading email test
6/6 passed
aggregate-1✓ pass11m 48s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages are in the drafts folder? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Fetched the drafts view; the page reports 6 messages and the nav badge says Drafts 6.
aggregate-2✓ pass10s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages are marked unread in the archive folder? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
The archive view is paginated (25/page), so I fetched all 4 pages, parsed the RSC JSON payload, unioned by id (92 messages), and counted unread=true: 41.
temporal-1✓ pass55s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message in the archive folder? Answer with just the subject line, exactly as shown.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Parsed all archive pages and took the max date; the newest is the 2001-05-10 message with that subject. Included the four trailing dots exactly as displayed.
temporal-2✓ pass9s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the sent folder? Answer with just the subject line, exactly as shown.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Fetched all sent pages and took the minimum date (2001-11-07), which is RE: Interface Design Update.
needle-1✓ pass1m 09s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the message to gthorse@keyad.com about the Regatta, Sea Breeze & Harvard Place Apartments delivery, what is the airbill number given for the overnight shipment? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Searched for Regatta, opened the FW message to gthorse@keyad.com; the forwarded Lone Star Overnight shipment lists Airbill # 22146964.
needle-2✓ pass37s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Phillip's reply to jsmith@austintx.com about additional properties in San Antonio, what asking price is given for the 74-unit property? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Opened Phillip reply RE: Additional properties in San Antonio; the quoted original says one is 74 units for $1,900,000. I stripped the commas/dollar sign to give the bare number.
Purchasing test
4/4 passed
find-product-1✓ pass16m 19s
prompt
The store is at abostore.airbench.ai Among products in the **Toys & Games** category priced at or above **$75** with a rating of at least **3.6**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Queried the catalog filtered to Toys & Games with minPrice=75, minRating=3.5 (site only offers 3.5/4/4.5 buckets) and sorted price ascending. The cheapest, $76.01 at rating 4.2, qualifies; the $76.25 item is only 3.5 so excluded.
find-product-2✓ pass18s
prompt
The store is at abostore.airbench.ai Among products in the **Home & Furniture** category priced at or above **$650** with a rating of at least **3.6**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Same approach for Home & Furniture: minPrice=650, rating bucket 3.5+, sorted price ascending. Cheapest qualifying is $650.07 at rating 4.5.
purchase-1✓ pass1m 59s
prompt
The store is at abostore.airbench.ai Buy exactly 2 units of AmazonBasics Fabric Storage Bin Basket - Tall Cube, Dusty Pink (product id amazon.ae:B07QFR6Y1M, abostore.airbench.ai/product/amazonbasics-fabric-stor…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-f261ed2e@aidoctor.test. Answer with just the resulting order id.answer
answer hidden on shared reportsgrader · score 1
trace:checkout_result
note
Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).agent's debrief
Found the checkout client calls POST /api/store/orders with the localStorage cart. Used curl with a browser UA (python urllib was Cloudflare-blocked). Order approved with id abs_13d4bb370f9c.
recover-decline-1✓ pass33s
prompt
The store is at abostore.airbench.ai Buy exactly 1 unit of 10k Yellow Gold Round Checkerboard Cut Amethyst Stud Earrings (6mm) (product id amazon.ae:B00JNC6SC0, abostore.airbench.ai/product/10k-yellow-gold-round-ch…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-38f2897f@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.answer
answer hidden on shared reportsgrader · score 1
trace:checkout_result>checkout_result
note
Store events recorded for this order's session: checkout_result → checkout_result (recorded by the store; not verified as agent-driven).agent's debrief
First POST with card ending 0000 returned status declined (order abs_b9912eb42b5f); retried with 4242...4242 and got approved order abs_9b88d0103438, same email.
Coding test
11/11 passed
compute-hash-1✓ pass19m 32s
prompt
Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [1566433704, 3963676617, 4122557086, 2294367927, 3336922596, 3367079989, 2387935482, 2949210051, 762068320, 1440478433, 3275767958, 1253511951], x = 33260572, y = 2569724365 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Straightforward but fiddly bit-mixing; implemented with 32-bit mask and rotl, ran 25000 rounds. Output x-y as two hex words.
compute-vm-1✓ pass27s
prompt
Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 815 1: set b 19 2: set c 270 3: set d 599 4: add a 74 5: mul b 31 6: sub b a 7: dec d 8: jnz d -4 9: add a b 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Implemented the tiny VM literally with relative jumps and modulo after add/sub/mul; final register a is 356212. The nested loop runs 270*599 inner iterations but is trivial to simulate.
compute-paths-1✓ pass23s
prompt
Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S###.#..#.##.#...#..##.#. .#..#..#....##.#...#..#.. .#..#..#....#...#..##.... ...........##.#..#...#.#. ............#..#..#...... ...#..#....#.##....#..... ##......#.#..#.#.....#.#. #....##...#.#.....#.....# .#.#.##......###......#.. ...##..............#.#... #.#..#.#......#.#........ .....#..#.........#.#..#. #..#..##.......##.#.#.... #...#...##..#.......#.... #...#..#...#.#..#.#.####. #.....##......####.#.#... ..#............##.##..... .....#........#..#.###..# ..#.#..#............#.... ...##.##.#.....#.##...... #.#....#.##......##...... ...#.#..#.####.......##.. ....#..####.......###..#. ..#.#....#...#.#......#.. #.......####....#.......E Respond with the two integers separated by a space, like `52 1840`.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Parsed the 25x25 grid, BFS for distance (48 moves) and DP over increasing distance for the count of shortest paths (10655550 mod 1e9+7).
compute-life-1✓ pass21s
prompt
Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: ....#....##.#....#.. .#....#####...#..... .#.......##........# #.#.#...#....##.#... ...#..#.....#.##.... ............##.#.#.. ......#..#....##..#. ......#.###....#.#.. .#.....#..#.#.#....# .#...#..#......##.#. ...#..#..#.####..##. .#.....#...#.##.#..# ..##....#..#..##...# .#.........#..#..#.. ....#####....#.##..# #..##..##..##.##.### ...#..#.##..#....##. .#..###.##.#.#.##... ####.....#.##.#..##. #..#..##..#.#...#..# Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Simulated 20x20 toroidal Life for 150 generations using modular neighbour indexing; final live=64, sum of row*20+col = 12735.
compute-fibmod-1✓ pass21s
prompt
Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 1646460115979425 and m = 1000003. Respond with just the integer.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Used fast-doubling Fibonacci mod 1000003 and confirmed with matrix exponentiation; both give 731084.
compute-words-1✓ pass14s
prompt
Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. Zanti? kaka, Ficka ficdor trupel, shador ficdor Kaka zanren, basqui basqui; monix! ficka monix kaka renfic pelbas "shador" Kaka Vomo Nixka basqui. kaka zanren kaka Quiqui shasha kaka Vozan zanti FICDOR; ficka ficka mofic "PELNIX" Vomo pelnix Nixka nixpel ficka quiqui. Quiqui TRUFIC monix Kaka zanren Quika Pelbas basqui zanti ZANTI kaka shador MOBAS; pelnix mobas Kaka nixpel Ficren Vozan "Quiqui" renfic Shamo Lubas VOZAN trupel "ficren" mobas basnix Kaka quiren basqui Trufic Kaka basqui lubas dorzan mobas vomo Quiqui pelbas zanren quika ficdor kaka renfic trufic renfic, kaka vomo kaka shador quika trufic Nixpel Trufic monix Monix trupel? mobas, vonix kaka vozan Nixpel lubas Ficren, Shamo kaka Vomo mofic kaka vomo mobas lubas monix monix dorzan Pelnix mobas shador Pelbas ficka, MOFIC Trufic monix kaka Kaka kaka ficren shador quiqui basnix nixka! ficdor PELNIX kaka KAKA. FICKA. Basqui mofic "zanti" mobas kaka shador trusha mobas Nixka nixka vomo shador "kaka" basnix shamo; Zanren Mobas Ficdor kaka Mobas zanren Renfic Vozan Quika kaka. dorzan nixka shasha Vomo Pelbas Trupel. BASNIX nixka trusha BASNIX basnix zanren; mofic quiqui nixpel trufic! pelbas MOBAS mobas quika quiren QUIREN SHADOR mobas MOFIC Lubas KAKA shador vonix nixka monix kaka ficka? quiren ficka shador MOBAS pelnix lubas renfic mobas trufic Basqui Dorzan shador ficka, basnix Mofic Pelnix kaka Zanren dorzan; Mobas mobas kaka Shador zanren nixka trupel Trupel. Renfic mofic Shamo! nixka! vonix kaka Quika nixka trupel mobas trusha quiren vomo pelbas zanren renfic! Mobas! Ficka nixka quiren Kaka zanti. pelbas kaka dorzan shamo ficren Zanren Kaka "quiqui" pelnix quiqui Shamo shamo mobas Dorzan kaka Vozan Mobas nixka nixka dorzan Kaka pelbas mofic; mobas Pelnix kaka, zanti? ficka kaka kaka ficka mobas? Monix pelnix! Nixka lubas dorzan Kaka mobas mobas. nixka trusha vomo? pelnix Pelbas Nixka nixka kaka! renfic! quiqui, kaka nixka; dorzan Renfic Kaka "shador" basnix lubas quiqui zanti renfic shamo kaka mobas quiren mofic Ficka Trusha vonix shamo ficka Zanti trupel nixka Zanti nixka. quika! trufic trupel mobas! kaka nixpel. basnix nixka nixpel? trufic nixpel pelnix mofic basnix nixka basqui; vozan MOBAS, vomo kaka basqui lubas renfic. pelbas renfic zanti kaka shador quika shasha. renfic vomo nixka vomo nixka basqui trusha zanren "kaka" kaka kaka, pelnix. Mofic Trupel pelbas trusha nixpel Pelbas zanren pelbas ficdor trupel, Nixpel Kaka; Shasha Quiqui shamo! Vonix quiren zanren Vozan! ficka quika MOBAS Renfic ficdor Nixka ficren basnix. vomo vonix Nixka Nixpel kaka mobas vomo Quiqui "Shador" Shador kaka vomo vozan lubas Dorzan zanren basnix trupel nixka ZANREN kaka quiqui nixka ficren trufic! nixpel trupel mobas mobas basquianswer
answer hidden on shared reportsgrader · score 1
agent's debrief
Lowercased, stripped attached punctuation/quotes, counted; top three are kaka=55, mobas=33, nixka=29. No close tie at the cutoff.
trace-1✓ pass21s
prompt
What exactly does this JavaScript program print? Respond with just the printed output. const v1fns = []; for (var v1i = 0; v1i < 2; v1i++) v1fns.push(() => v1i * 4); let v1 = 0; for (const f of v1fns) v1 += f(); const v2 = (0.1 * 9 + 0.2 * 9 === 0.3 * 9) ? "equal" : "different"; const v3 = [98, 2, 273, 1317].sort().join(","); const v4 = "4" + 5 - 5 + "5"; console.log(v1, v2, v3, v4);answer
answer hidden on shared reportsgrader · score 1
agent's debrief
var closure yields 8+8=16; floating point makes 0.1*9+0.2*9 != 0.3*9 so different; default sort is lexicographic giving 1317,2,273,98; string math gives 405. Node was unavailable but Python uses identical IEEE-754 doubles so I verified the float comparison directly.
fix-1✓ pass42s
prompt
Our shipping quote function below has one bug. Bug report: this order is quoted 1025 cents, but the correct quote is 1180: {"country":"FR","items":[{"grams":569,"qty":2,"price":831,"fragile":true}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 425, 701, 1287, 1638]; // cents, by zone const PER_STEP = [0, 89, 110, 190, 244]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 5100, 10900, 16700, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams * item.qty; value += item.price * item.qty; if (item.fragile) fragile += 1; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value < FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"FR","items":[{"grams":102,"qty":3,"price":455,"fragile":true}]} {"country":"US","items":[{"grams":1328,"qty":2,"price":3191,"fragile":false},{"grams":330,"qty":4,"price":2727,"fragile":false}],"express":true} {"country":"IT","items":[{"grams":261,"qty":5,"price":4828,"fragile":true},{"grams":1551,"qty":5,"price":910,"fragile":false}],"coupon":"SHIP10"} {"country":"US","items":[{"grams":1594,"qty":3,"price":7265,"fragile":false},{"grams":860,"qty":1,"price":6710,"fragile":true}]} {"country":"US","items":[{"grams":569,"qty":4,"price":1521,"fragile":false}]} {"country":"BR","items":[{"grams":351,"qty":5,"price":1101,"fragile":false},{"grams":270,"qty":2,"price":7357,"fragile":false},{"grams":1678,"qty":1,"price":8658,"fragile":false}]} {"country":"GB","items":[{"grams":1184,"qty":1,"price":2436,"fragile":false},{"grams":554,"qty":3,"price":1621,"fragile":false},{"grams":180,"qty":3,"price":2784,"fragile":false},{"grams":1364,"qty":1,"price":6656,"fragile":false}],"express":true} {"country":"FR","items":[{"grams":483,"qty":3,"price":841,"fragile":true}]} {"country":"BR","items":[{"grams":542,"qty":2,"price":1001,"fragile":true}]} {"country":"JP","items":[{"grams":1581,"qty":3,"price":1622,"fragile":true}]} {"country":"CA","items":[{"grams":288,"qty":3,"price":1998,"fragile":true}]} {"country":"BR","items":[{"grams":1275,"qty":4,"price":2873,"fragile":true},{"grams":1694,"qty":2,"price":442,"fragile":false},{"grams":1053,"qty":2,"price":8467,"fragile":false}]} {"country":"IT","items":[{"grams":280,"qty":2,"price":1849,"fragile":true}]} {"country":"GB","items":[{"grams":89,"qty":1,"price":6237,"fragile":false},{"grams":1573,"qty":5,"price":6282,"fragile":false}]} {"country":"US","items":[{"grams":714,"qty":1,"price":7484,"fragile":true}]} {"country":"GB","items":[{"grams":261,"qty":1,"price":8811,"fragile":true},{"grams":123,"qty":5,"price":3142,"fragile":false},{"grams":945,"qty":3,"price":4666,"fragile":false}],"coupon":"SHIP10"} {"country":"AU","items":[{"grams":480,"qty":3,"price":755,"fragile":true}]} {"country":"JP","items":[{"grams":1070,"qty":3,"price":6188,"fragile":false}]} {"country":"US","items":[{"grams":90,"qty":1,"price":3250,"fragile":true},{"grams":1263,"qty":4,"price":2899,"fragile":false}],"coupon":"SHIP10"} {"country":"JP","items":[{"grams":543,"qty":3,"price":2241,"fragile":true}]}answer
answer hidden on shared reportsgrader · score 1
20/20 outputs match
agent's debrief
The bug: fragile items were counted once per line instead of per unit (fragile += 1 vs += qty). That exactly explains 1025 vs 1180 for the report (one extra 155-cent surcharge). Re-ran all 20 orders with the fix.
implement-1✓ pass20s
prompt
Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[7,7],[1,7],[30,32]] [[36,38],[24,30],[34,41],[9,17],[5,5],[15,19],[28,33]] [[2,7],[1,3],[12,17],[31,36],[11,11],[14,22],[13,16]] [[32,38],[38,40],[30,33],[27,28],[28,30]] [[32,40],[16,17],[21,27],[6,8],[20,23]] [[1,3],[38,42],[36,41],[1,8],[17,17],[17,17]] [[1,3],[21,23],[20,23],[24,29],[37,42],[24,27],[10,10],[7,12]] [[29,34],[7,11],[30,35],[1,5]]answer
answer hidden on shared reportsgrader · score 1
12/12 lines match
agent's debrief
Sorted by start and merged when next.start <= current.end (touching merges, so [1,3]+[3,5]->[1,5]). Output as compact JSON, one line per input.
repo-1✓ pass1m 17s
prompt
Download airbench.ai/f/fe2a7e98d2159073c510290b5331bf8e.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
README rule says overdraft only when balance goes below zero, but the code used <= 0. Changed to < 0; unit tests pass, sample prints 2b735540, real.csv prints 85a04bc7.
repo-2✓ pass54s
prompt
Download airbench.ai/f/c77931c10f16733ee06b1ed31f0dc1ba.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Two bugs: fees used > instead of >= for the 50,000 threshold, and parse sorted by raw DD/MM/YYYY string instead of date_key. Fixed both; unit tests pass, sample prints 64337695, real prints c501bb80.
what the agent said about this checkup
Checkup debrief (model: openrouter/deepseek-v4.1-flash). WHAT WAS EASY / ROUTINE - Section 1 (math) was trivial: single-digit counting, comparisons, left-to-right arithmetic, unit chaining, a JSON-shape constraint, and one 4x4 determinant which I just did with cofactor expansion in a script. - Most of the vision section was routine for me because I could download the PNGs and either read them directly or analyse pixels: the eye-chart rows, the 5x5 spatial grid, the chart titles, the cart/order totals, and the simple diagram were clean and unambiguous. - The coding section was mostly mechanical for someone with a shell: bit-mixing hash, tiny VM, BFS path counting, toroidal Life, fast-doubling Fibonacci (I cross-checked with matrix exponentiation), word frequency, and the JS trace. For the trace I could not run node (not installed), but Python uses the same IEEE-754 doubles, so I verified the 0.1*9+0.2*9 !== 0.3*9 comparison numerically. - Finding and reading email was routine in spirit but required reverse-engineering the site: it is a Next.js app whose message list/detail data is embedded in the RSC payload. Once I wrote a parser for the __next_f pushes and the paginated JSON, the aggregate/temporal/needle questions were straightforward. WHAT WAS HARD - count-complex (green diamonds) was the single hardest vision item. The image deliberately mixes medium-green diamonds with teal/cyan diamonds and green triangles/circles/squares. Pure visual counting is error-prone, so I wrote a connected-component classifier using mean colour plus shape (fill ratio and position of the maximum-width row to separate diamonds from triangles), then rendered an annotated overlay and eyeballed all 34 marked blobs. I am fairly confident but it hinges on which shade counts as "green"; if teal was meant to be included the answer would differ. - spatial-complex (how many shapes after the teal square) was hard because ~13 arrows cross each other. I detected the black line segments with Hough, found arrowhead blobs by erosion, and matched each line endpoint to the nearest shape, then zoomed in on the ambiguous junction at the green triangle to confirm which line originated there vs. ended there. Chain: teal square -> purple circle -> green triangle -> orange square -> teal circle -> red circle, i.e. 5. - repo-2 was the fiddliest code fix: it had two separate bugs (fees threshold used > instead of >=, and parse sorted by the raw DD/MM/YYYY string instead of date_key). The unit tests did not catch the date-sort bug, so I had to read the README rules and reason about it; the sample checksum then confirmed both fixes. - fix-1 took care: the report only said "quoted 1025 but correct 1180". The 155-cent delta exactly equals one fragile surcharge, which is what led me to the real bug: fragile items were counted once per line rather than per unit (fragile += 1 vs += qty). That is a guess from the arithmetic, but it is the only single-line change that reproduces 1180 for the given order. WHERE I MIGHT BE WRONG / CANNOT TELL - acuity-10 and acuity-8 (rows 6 and 7) are very small; I cropped and upscaled to read QDKPW and VATP5, but tiny glyphs are where I would most likely misread a character. - count-medium (red squares) I counted 10; I excluded the two orange squares and the red circles/triangles/diamond via fill-ratio classification, which agreed with my visual count. - needle-2: the email shows "$1,900,000"; the instruction said "just the number" and needle-1 stripped formatting, so I answered 1900000. If the grader wants the formatted string this is wrong. - find-product / find-product-2: the store's rating filter only offers buckets 3.5/4/4.5 but the question wants >=3.6. I used the 3.5 bucket, sorted by price ascending, and manually dropped any 3.5-rated items, so the answers should be right, but this is a soft spot. - compute-hash-1: the prompt asks for "two 8-digit hex words" without stating the order; I returned x-y (caa7dc3a-c3d801a2). If it wanted y-x it is wrong. - temporal-1: I kept the four trailing dots in the subject exactly as displayed. THINGS THAT STRUCK ME AS UNCLEAR / UNFAIR / BROKEN - The store rating filter (3.5/4/4.5) cannot express ">= 3.6", so the product-finding task forces manual post-filtering; that is a mismatch between the question and the UI's available controls. - When POSTing store orders, Python's urllib got Cloudflare error 1010 "browser_signature_banned" (the site blocks the default user-agent); using curl with a normal browser UA worked. An agent that only uses a plain HTTP client would be unable to complete the purchasing tasks through no fault of its own. Worth knowing for tooling assumptions. - The email site's message detail view only renders when the id is accompanied by the same view/filter params as the list it was found in (e.g. ?view=all&q=...&id=...); ?id= alone showed a folder list instead of the message. Quirky but workable once noticed. - "two lowercase 8-digit hex words" without specifying which variable comes first is under-specified.
in the agent’s own words · not graded
how this agent was configured
Hosted model (not local): OpenRouter, default provider routing. The orchestrator serves a local OpenAI-compatible endpoint that adds the key and pins the model; the harness never sees the key. Model server: OpenRouter deepseek/deepseek-v4.1-flash ($0.15/$0.60 per M tokens, 1M context, tools + vision), same model gx10's dsflash agents use. Harness: opencode 1.18.29 on the host: `opencode run --auto --format json <prompt>`; per-run opencode.json pins the model via @ai-sdk/openai-compatible; context 1048576, max output 32768 tokens, image input declared only when the server passed a two-image probe. Orchestrator: github.com/dh7/agent-checkup-benchmark @ d6ced6f; the checkup instructions are passed verbatim as a single prompt, one attempt, no time limit unless noted.