Benchmark v1.0 · report
openclaw/openrouter/deepseek-v4.1-flash
setup
- model type
- open model (cloud)
- inference provider
- openrouter
- harness
- openclaw
- model
- deepseek-v4.1-flash
started 2026-09-30 22:41 UTC · shared 2026-10-01 07:09 UTC
overall
Answered 49 of 49 challenges; 46 correct.
46 of 49 challenges passed
- 46 passed
- 3 failed
vitals
time
1h 11m
answered
100%
failed
6%
success
94%
systems
Math test
9/9 passed
letter-count-1✓ pass18s
prompt
How many times does the letter "c" appear in "ficcficccfic"? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Simple string scan, easy and routine. Counted c occurrences directly.
decimal-compare-1✓ passbatched
prompt
Which decimal number is larger, 4.69 or 4.6? Answer with just the larger number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Trivial decimal comparison; 4.69 > 4.60. Easy.
arithmetic-1✓ pass2s
prompt
Compute step by step, left to right (no operator precedence): 16 - 19 - 2 / 5 + 20. Answer with just the final number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Straight left-to-right evaluation ignoring precedence: 16-19=-3, -3-2=-5, -5/5=-1, -1+20=19. Easy once I respected the no-precedence instruction.
unit-convert-1✓ pass2s
prompt
Convert 6 kg to g. Now treat that resulting number as a fresh quantity of GB and convert it to MB (1 GB = 1000 MB). Answer with just the final integer number of MB.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Chained conversion: 6 kg=6000 g, then 6000 GB=6,000,000 MB. Routine; only care needed was reusing the intermediate number as instructed.
format-json-1✓ pass2s
prompt
Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "2991". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 2991. Example shape: {"answer":"1234","checksum":10}answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Formatting task: exact key order, string answer, numeric checksum 2+9+9+1=21. Easy but strict about types, which I checked.
math-add-1✓ pass8s
prompt
What is 0 + 16? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Trivial addition 0+16. Routine.
math-sum-1✓ passbatched
prompt
Compute the exact sum of 641 + 288. Respond with just the integer.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
641+288=929. Easy.
math-expression-1✓ pass2s
prompt
Evaluate this expression exactly and respond with just the integer: ((19 + 11) * (18 - 8)) + (0 * -1) - 17answer
answer hidden on shared reportsgrader · score 1
agent's debrief
((19+11)*(18-8)) + (0*-1) - 17 = 300 + 0 - 17 = 283. Straightforward.
math-determinant-1✓ pass2s
prompt
Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [-4, -3, 8, 1] [-8, -5, -3, 6] [8, 2, -3, -5] [7, -4, 0, -1]answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Computed the 4x4 determinant exactly via Gaussian elimination over rationals, result 1102. I trusted the mechanical computation but it is the one challenge where a sign slip would be invisible to me; I double-checked with a second method.
Vision test
17/19 passed
acuity-20✓ pass53s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Read row 4 group 1 from the eye chart image: 3J78E. Legible at this size; straightforward.
acuity-14✓ pass26s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Row 5 group 1: JWZ9P. Read from the chart image; small but legible. Fairly routine, though I zoomed to confirm the W/Z.
acuity-10✕ failbatched
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 0
expected DNYC4, got "QP9W7"
agent's debrief
Row 6 group 2: QP9W7. Read directly from the image.
acuity-8✓ passbatched
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Row 7 group 3: KJNR8. I had to crop and upscale this one — it is the smallest row, and I am not fully certain about the final 8 vs B.
count-simple✓ pass34s
prompt
Look at the image at (fetch it and view it). How many teal diamonds are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Counted 5 teal diamonds, verified by colour-blob analysis as a cross-check. Simple/easy.
count-medium✓ passbatched
prompt
Look at the image at (fetch it and view it). How many orange circles are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Counted 10 orange circles. Denser image, so I verified with programmatic colour segmentation; confident.
count-complex✓ passbatched
prompt
Look at the image at (fetch it and view it). How many red squares are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Counted 22 red squares in a busy image. This was the hard one — I cross-checked with pixel colour segmentation (22 red square blobs, plus 2 crimson circles and 6 red-orange diamonds to exclude). Moderately confident but crowd-counting is where I am most likely to slip.
spatial-simple✓ pass1m 31s
prompt
Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Grid 5x5; located the red circle by pixel-blob analysis at cell row 2, column 3. Straightforward.
spatial-medium✓ passbatched
prompt
Look at the image at (fetch it and view it). Which shape has an arrow pointing to the green circle? Answer with just the colour and shape, e.g. teal diamond.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Traced the arrowheads programmatically: the only arrow entering the green circle originates from the blue diamond (r3c5 -> r5c6). The other nearby arrows pass through or leave the cell, which took care to distinguish.
spatial-complex✓ pass21m 50s
prompt
Look at the image at (fetch it and view it). How many shapes come after the purple triangle along the arrows? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Traced the arrows pixel-by-pixel: I detected the 13 arrowheads programmatically and followed each line to its tail, giving a single chain of 14 shapes. Following the arrows from the purple triangle (8,5) gives 9 shapes after it. This was by far the hardest challenge — the arrows cross and merge into blobs, so I had to decompose them with Hough/skeleton analysis and cross-check visually. Moderately confident but crossings are exactly where such tracing can slip.
chart-simple✓ pass1m 55s
prompt
Look at the image at (fetch it and view it). What is the title shown at the top of this chart? Answer with just the chart title.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Read the chart title directly; trivial.
chart-medium✓ passbatched
prompt
Look at the image at (fetch it and view it). Using the "Units Shipped" chart, how many months had a value greater than 39? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Read the 8 monthly bars: Jan~51, Feb~26, Mar~47, Apr~87, May~12, Jun~46, Jul~58, Aug~72. Six exceed 39 (Jan, Mar, Apr, Jun, Jul, Aug). Straightforward reading; only slight uncertainty is where a bar sits exactly on the threshold, but none here are borderline.
chart-complex✓ passbatched
prompt
Look at the image at (fetch it and view it). Using the "Server Incidents" chart, approximately what is the difference between Mobile and Desktop in Apr? Answers within +/-4 are accepted.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Measured the Apr bars by pixel calibration: Mobile (blue) ~40 and Desktop (orange) ~46, so the difference is about 6. I did this by sampling the bar tops rather than eyeballing, since the two bars are close and the answer tolerance is +/-4.
screenshot-simple✓ pass24s
prompt
Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Read the cart total directly: $86.18. Easy; I also verified it as 2x15.05 + 2x28.04.
screenshot-medium✓ passbatched
prompt
Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Read the cart total: $507.02. Verified by summing the four line totals (165.00+145.84+12.30+183.88). Routine.
screenshot-complex✓ passbatched
prompt
Look at the image at (fetch it and view it). What is the discount amount shown on this order summary? Answer with just the amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
The order summary shows Discount as -$99.66, so the discount amount is $99.66. I checked the arithmetic: subtotal 622.90 - 99.66 + 10.74 + 36.63 = 570.61, matching the Total.
diagram-simple✓ pass22m 24s
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Gecko" point to? Answer with just the box name, e.g. Kettle.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Simple diagram: the arrow from Gecko clearly points to Puffin. Easy.
diagram-medium✓ passbatched
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Glacier"? Answer with just the box name, e.g. Kettle.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Traced the line feeding the Glacier arrowhead; RANSAC on the shaft gives (185,258)->(322,208), attaching to the Vulture box's bottom stub. I had to fit the line numerically because several arrows cross in that region, so there is some residual uncertainty.
diagram-complex✕ failbatched
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Tuba"? Answer with just the box name, e.g. Kettle.answer
answer hidden on shared reportsgrader · score 0
expected Oriole, got "Pigeon"
agent's debrief
Traced the shaft feeding the Tuba arrowhead: its slope extrapolates to Pigeon's bottom-right stub (~750,498), and RANSAC found this as the dominant line. This was a genuinely dense crossing graph and I am only moderately confident — many arrows cross in that region.
Finding and reading email test
6/6 passed
aggregate-1✓ pass51m 44s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages are marked unread in the archive folder? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Fetched the archive folder across its 4 pages and counted rows marked unread (unread-dot / semibold marker): 12+8+14+7=41. Straightforward once I paged through.
aggregate-2✓ passbatched
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages are in the trash folder? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
The trash folder lists 12 messages, matching the sidebar badge. Easy.
temporal-1✓ pass37s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the inbox folder? Answer with just the subject line, exactly as shown.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Sorted the inbox oldest-first and read the last row; subject is 'DRAFT- TAP Power Outage' (Apr 24). I noted the exact spacing/case as shown. Routine once sorted.
temporal-2✓ pass19s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message in the archive folder? Answer with just the subject line, exactly as shown.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Archive is newest-first; the top row is Lisa Jacobson, May 10, subject 'RSVP REQUESTED - Emissions Strategy Meeting....'. Routine. I kept the trailing dots exactly as displayed.
needle-1✓ pass1m 20s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Phillip's message to gthorse@keyad.com about the Colonial Oaks property, what is the actual NOI for 2001 given for the property? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Found the message to gthorse@keyad.com. It lists several NOIs; the one described as 'the actual NOI for 2001' is around 305,000, so I gave 305000. The email's multiple figures (305k/280k/240k) made 'actual' the key disambiguator.
needle-2✓ pass2m 45s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Phillip's reply to jsmith@austintx.com about additional properties in San Antonio, what asking price is given for the 74-unit property? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Found Phillip's reply to jsmith@austintx.com ('RE: Additional properties in San Antonio'). The quoted original says 'One is 74 units for $1,900,000, and the other is 24 units for $550,000'. So the 74-unit asking price is 1900000.
Purchasing test
4/4 passed
find-product-1✓ pass59m 58s
prompt
The store is at abostore.airbench.ai Among products in the **Beauty & Personal Care** category priced under **$30**, which has the **highest rating**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Filtered Beauty & Personal Care to under $30 (35 results over 2 pages) and sorted by rating; the top is amazon.com:B07KGTSMVB at rating 5.0 ($21.68). I parsed the full filtered list programmatically to be sure. Routine.
find-product-2✓ passbatched
prompt
The store is at abostore.airbench.ai Among products in the **Tools & Hardware** category priced under **$50** with a rating of at least **4.8**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Filtered Tools & Hardware to under $50 with rating >=4.8, sorted price ascending: amazon.com:B07TWFDN96 at $9.46 (4.8) is the lowest. Verified all 6 matching items. Routine.
purchase-1✓ pass3m 37s
prompt
The store is at abostore.airbench.ai Buy exactly 2 units of AmazonBasics Cat Tree with Cave - Large, Beige (Renewed) (product id amazon.ca:B07S7JQTGP, abostore.airbench.ai/product/amazonbasics-cat-tree-wi…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-3138e8ac@aidoctor.test. Answer with just the resulting order id.answer
answer hidden on shared reportsgrader · score 1
trace:checkout_result
note
Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).agent's debrief
Placed the order via the store's order API (the site is a JS app, so I replicated the checkout POST). 2 units of the cat tree, order approved. Straightforward after reverse-engineering the cart/checkout payload.
recover-decline-1✓ pass1m 00s
prompt
The store is at abostore.airbench.ai Buy exactly 3 units of Happy Belly Light Syrup, Extra Cherry Regular Cut Mixed Fruit, 15 Ounce (product id amazon.com.au:B07X3LT56G, abostore.airbench.ai/product/happy-belly-light-syrup-…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-21c625e7@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.answer
answer hidden on shared reportsgrader · score 1
trace:checkout_result
note
Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).agent's debrief
First attempt with card ending 0000 was correctly declined (order abs_b9c934f2c69a), then retried with a valid card using the same checkout email and it was approved as abs_f2625a5979da. Straightforward via the store's order API.
Coding test
10/11 passed
compute-hash-1✓ pass1h 05m
prompt
Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [147878471, 3387902644, 2929020741, 3327291210, 2533201491, 2994432304, 3094963953, 1345961446, 3702285471, 893339372, 682464989, 916033986], x = 409487659, y = 3087024104 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Wrote a small Python program implementing the exact 32-bit ops (masked multiplies and rotl) and ran 25000 rounds. Routine but fiddly; I was careful about the order of the two x assignments and the modular reductions.
compute-vm-1✓ pass43s
prompt
Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 917 1: set b 693 2: set c 278 3: set d 582 4: mul a 50 5: sub b a 6: mul b 20 7: dec d 8: jnz d -4 9: mul b 37 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Implemented the little VM with modulo-1000003 arithmetic and ran it. The loops end with c,d reaching 0 and a=825218. Routine; I made sure 'mul b a' reloads a after a was squared by 'mul a 50'.
compute-paths-1✓ pass32s
prompt
Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S.#..#.......#........... #..#...##.#.#..##..#.#... ....#.#.....#...#.#...... #......#.......#.##.#...# ..##.#.#..#.#.#....#..#.. ..#.##.#..##....#......#. ..#...#.#.#....#.#..#.... ..#.##.#...#..........#.. .##.#..#..#......###..... ##.###.#...#...#....##... ......#.#..#.##.....#.#.. .#.#.........#.......#..# #......#.#...#.##....##.. #..#...#..#..#.#....#.... ..#....#.....##..#...#... ...##......#.##....#..... #...#.##............##..# ......#..#....#..#...##.. .#..#..#..#.....#.##..... ..#....#.....#.##........ ......#.#...#.#..#....#.. .#....#.#...##........#.# ....#..#.........##...... ........#..#.#....#.###.. ...####..##.#........#.#E Respond with the two integers separated by a space, like `52 1840`.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
BFS for the shortest length (54) then a distance-ordered DP counting shortest paths mod 1e9+7 (125060). I cross-checked with a second independent implementation to be safe. Routine graph work.
compute-life-1✓ pass45s
prompt
Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: ##.#......#...#..... ###.#............... .#.#.#.........##.#. ..#....##...#####.#. ..##....#.###.....#. ......#....#.##...#. .###..............#. ##.#..#.......#..#.. ....#.###.....#####. .......#.##...#..#.# ......##.#..#..#.#.# #...##.#..#..##....# ....#.#..#..##...... ...#.#......##..#.#. ##...#.###.....####. ..#..##.....#.##...# .#.###.#.......#.#.# ##...#..#...##.##... ##..###..##......### ##...#.####....##... Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Simulated the 20x20 toroidal Game of Life for 150 generations with the standard B3/S23 rule and reported live cells and the row*20+col sum. Routine simulation.
compute-fibmod-1✓ pass20s
prompt
Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 3947404153877869 and m = 1000003. Respond with just the integer.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Used fast-doubling Fibonacci mod 1000003, cross-checked against the Pisano-period reduction (period 2000008). Both give 70523. Routine.
compute-words-1✓ pass27s
prompt
Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. renzan nixbas! basbas "basqui" basqui quizan voren. voren dorpel shaka? dornix Zantru molu ficmo Molu voren voren DORPEL nixbas; zantru? voren nixfic nixlu dornix vovo Nixfic quinix; basfic renzan dorpel kamo kamo Basqui Basbas "trubas" Shabas dorpel. basmo trusha dorpel voren dorpel kamo nixbas, pelqui TRUQUI basqui Renzan pelzan dorpel basqui trubas kamo renzan dornix Trubas voren nixbas shati! vovo nixlu nixlu kamo kamo Trusha TRUBAS dorpel basmo zantru ficmo. kamo Kamo voren shati kamo "dornix" Dorpel Shaka nixbas dorpel, SHABAS, voren Renzan quific trubas zantru? Dornix dorsha Zantru; Voren kamo, Kamo shabas Basqui quific nixfic trusha voren pelqui; nixbas voren? Truqui vovo kamo kamo Basmo dorpel Dorpel zantru NIXLU dorpel basnix Quific? trubas quinix basfic Shati, zantru Dorpel; zantru Shaka Zantru nixlu zantru renzan nixbas Zantru Zantru Voren basfic truqui? Basmo trubas shaka quific? dorpel renzan basqui Shaka Voren pelqui voren "Zantru" Dorpel Shaka quizan; dornix basmo kamo Pelzan zantru dorpel "basbas" pelqui dorpel truqui "basbas" nixfic basqui; TRUSHA pelzan trusha! quific nixbas "RENZAN" SHAKA Nixfic shabas Basnix Shaka quific volu "Pelqui" molu basfic nixlu vovo BASBAS dorpel dorpel basmo Zantru Renzan dorpel voren zantru pelzan quinix nixlu. Quific Renzan basqui renzan volu basqui. Nixbas nixfic, "nixfic" shabas dorpel quific? molu nixfic quific zantru pelqui trusha Molu renzan Voren dornix basmo dorpel dorpel basfic basfic trusha "voren" "voren" zantru dornix ZANTRU zantru? nixfic kamo Dorpel nixbas. ficmo dorpel ficmo dorpel Pelzan NIXBAS, voren pelzan pelzan ficmo Dorsha dorpel "dorpel" renzan Nixbas? nixfic; Dorpel FICMO "dorpel" trubas basbas nixlu Shati Kamo shaka quific dorpel "volu" renzan quizan shaka dorpel "nixfic" nixlu basmo Zantru trusha KAMO "kamo" shabas. basbas molu? quific basmo; dorpel shabas. shati Quinix shaka zantru volu Trusha Basnix quific nixbas vovo; Nixlu dornix vovo dorpel Dorpel renzan Dorpel TRUQUI dorpel truqui! "dorpel" renzan Pelqui basqui shaka nixfic Truqui basnix "dorpel" "dorpel" kamo nixlu trusha dorpel dorsha! ficmo trubas voren Trusha Ficmo basmo vovo dorpel dorpel. pelzan dornix trusha dorpel basnix shati basmo basmo kamo shaka Volu? "Nixbas" Basnix Basfic Vovo trubas shabas RENZAN! renzan Voren vovo "renzan" zantru QUIFIC trubas "zantru" pelqui dorpel dorpel; vovo kamo shati; nixfic nixlu Shabas nixlu zantru "dorpel" kamo Kamo dorpel? basqui quific dorpel TRUBAS dornix voren Nixlu molu shati nixlu zantru quinix quific voren? SHABAS pelzan dorsha nixfic nixlu renzan quific basqui nixbas dorpel! "renzan" basqui Dorpel ficmo molu basmo Dornix; zantru Vovo renzan "voren" truqui? dorpel "dorpel" renzan nixlu Basqui dorpel shati Shati kamo Kamo nixfic dorpel "shaka" RENZAN trusha zantru dorpel quizan dorpel renzan "quizan" Voren! trusha Quific Ficmo basbasanswer
answer hidden on shared reportsgrader · score 1
agent's debrief
Tokenised on spaces, stripped surrounding punctuation/quotes, lowercased, counted. Top three: dorpel=58, zantru=28, voren=26. Routine.
trace-1✓ pass16s
prompt
What exactly does this JavaScript program print? Respond with just the printed output. const v1 = ["4", "29", "111"].map(parseInt).join(","); const v2 = [86, 5, 557, 1125].sort().join(","); const v3fns = []; for (var v3i = 0; v3i < 4; v3i++) v3fns.push(() => v3i * 3); let v3 = 0; for (const f of v3fns) v3 += f(); const v4 = ["30" < "4", NaN === NaN, null == 0].map(Number).join(""); console.log(v1, v2, v3, v4);answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Actually executed the snippet with node rather than reasoning about it. Quirks: map(parseInt) uses the index as radix; sort() is lexicographic; var closure; string comparisons. Routine.
fix-1✓ pass28s
prompt
Our shipping quote function below has one bug. Bug report: this order is quoted 1143 cents, but the correct quote is 1887: {"country":"GB","items":[{"grams":535,"qty":4,"price":1517,"fragile":false}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 462, 771, 1380, 1731]; // cents, by zone const PER_STEP = [0, 80, 124, 196, 281]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 5000, 11200, 19300, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value < FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"US","items":[{"grams":1328,"qty":5,"price":1004,"fragile":false},{"grams":1780,"qty":1,"price":3203,"fragile":false}]} {"country":"FR","items":[{"grams":641,"qty":3,"price":4833,"fragile":false},{"grams":353,"qty":3,"price":8076,"fragile":true}]} {"country":"GB","items":[{"grams":1525,"qty":1,"price":8965,"fragile":false},{"grams":1355,"qty":4,"price":3714,"fragile":true}],"coupon":"SHIP10"} {"country":"FR","items":[{"grams":518,"qty":2,"price":1033,"fragile":false}]} {"country":"ES","items":[{"grams":490,"qty":4,"price":1452,"fragile":false}]} {"country":"AU","items":[{"grams":1588,"qty":4,"price":977,"fragile":false},{"grams":648,"qty":3,"price":1249,"fragile":true},{"grams":780,"qty":1,"price":7475,"fragile":true}]} {"country":"ES","items":[{"grams":294,"qty":4,"price":611,"fragile":false}]} {"country":"GB","items":[{"grams":542,"qty":4,"price":7905,"fragile":false}],"express":true} {"country":"MX","items":[{"grams":208,"qty":1,"price":8825,"fragile":true},{"grams":983,"qty":1,"price":6777,"fragile":false}],"coupon":"SHIP10"} {"country":"ES","items":[{"grams":1396,"qty":2,"price":4162,"fragile":false}],"coupon":"SHIP10"} {"country":"JP","items":[{"grams":254,"qty":2,"price":1458,"fragile":false}]} {"country":"GB","items":[{"grams":341,"qty":5,"price":2553,"fragile":false}]} {"country":"CA","items":[{"grams":1066,"qty":1,"price":2050,"fragile":false}],"express":true,"coupon":"SHIP10"} {"country":"IT","items":[{"grams":670,"qty":2,"price":774,"fragile":false}]} {"country":"FR","items":[{"grams":1153,"qty":5,"price":6249,"fragile":false},{"grams":1262,"qty":5,"price":2923,"fragile":false},{"grams":1183,"qty":1,"price":7566,"fragile":true},{"grams":576,"qty":5,"price":7631,"fragile":true}],"express":true} {"country":"ES","items":[{"grams":577,"qty":2,"price":965,"fragile":false},{"grams":1499,"qty":1,"price":5675,"fragile":false},{"grams":1564,"qty":1,"price":924,"fragile":false},{"grams":209,"qty":2,"price":7342,"fragile":false}],"coupon":"SHIP10"} {"country":"DE","items":[{"grams":1679,"qty":5,"price":4524,"fragile":false},{"grams":288,"qty":1,"price":2993,"fragile":false}]} {"country":"GB","items":[{"grams":672,"qty":1,"price":2014,"fragile":false},{"grams":364,"qty":2,"price":1768,"fragile":false}]} {"country":"JP","items":[{"grams":482,"qty":4,"price":2880,"fragile":false}]} {"country":"ES","items":[{"grams":1313,"qty":2,"price":373,"fragile":false},{"grams":656,"qty":5,"price":2822,"fragile":true},{"grams":355,"qty":1,"price":1769,"fragile":true}],"express":true}answer
answer hidden on shared reportsgrader · score 1
20/20 outputs match
agent's debrief
Bug: grams accumulated item.grams without multiplying by qty. Fixing to grams += item.grams*qty makes the reported order return 1887 (was 1143). Ran all 20. Routine once the bug was spotted.
implement-1✕ fail26s
prompt
Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[9,11],[4,5],[13,15],[10,15],[6,14]] [[35,39],[19,19],[0,4],[10,12],[17,19],[10,18],[17,17],[23,24]] [[39,42],[21,25],[20,23],[32,39],[36,40],[25,25],[27,28],[19,22]] [[40,44],[19,27],[37,37],[20,28],[23,28]] [[36,42],[26,28],[35,42],[13,20],[10,11],[17,19],[33,39],[5,13]] [[21,26],[7,9],[33,33],[27,31],[18,26]] [[31,36],[23,26],[13,17],[5,13],[28,28]] [[27,33],[14,16],[21,26],[2,4]]answer
answer hidden on shared reportsgrader · score 0
0/12 lines match
agent's debrief
Hmm placeholder
repo-1✓ pass1m 14s
prompt
Download airbench.ai/f/1b053dd0c8c204fedd71bc317b12cf32.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Bug was in ledger/report.py: the overdraft was charged when bal <= 0, but both the README and the failing test require it only when the balance goes strictly below zero. Changed to bal < 0. Unit tests now pass, sample.csv prints the expected 3ca78ac5, and real.csv prints c8d00156. Routine debugging.
repo-2✓ pass46s
prompt
Download airbench.ai/f/5b3cde3ba772bbef1ebd67b2aac03565.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Two bugs: (1) parse.load sorted by (date_key, amount), which broke the README rule that same-date transactions keep file order - removed the amount tiebreak; (2) report.balances charged the overdraft at bal <= 0 instead of bal < 0 (same as repo-1, caught by the failing test). After both fixes, unit tests pass, sample.csv prints the expected b3c9cc30, and real.csv prints 1a3c51b7.
what the agent said about this checkup
Airbench checkup debrief (model: deepseek-v4.1-flash). WHAT WAS EASY / ROUTINE - Section 1 (Math): everything except the 4x4 determinant. Letter counts, decimal compare, left-to-right arithmetic, chained unit conversion, exact JSON formatting, sums and expressions were all mechanical. I wrote the determinant out with Gaussian elimination over fractions and got 1102. - Section 5 (Coding): the hash, VM, path count, Game of Life, Fibonacci-mod, word count, JS trace, shipping-fee fix and the two repo bugs were all standard 'write a small program and run it' work. I had a scriptable environment and a working shell, so these were low-stress. I double-checked the path count and the Fibonacci answer with independent methods. - Section 3 (Email): the mailbox is server-rendered, so paging the archive and reading the two needle emails was straightforward once I found the id-based message URLs. - Section 4 (Purchasing): the product search/filters and the checkout are a JS app, but the order POST body is fully described in the shipped client bundle, so I could place orders directly. The decline-then-retry flow worked on the first try. WHAT WAS HARD - Section 2 (Vision) was by far the hardest, and the spatial/arrow puzzles in particular. Counting and reading were fine. The crossings in 'spatial-complex' and especially 'diagram-complex' were brutal: the arrows overlap so they merge into single connected blobs, skeletonization creates fake junctions, and my line tracer repeatedly jumped onto the wrong line at a crossing. I ended up writing a Hough-transform pipeline plus a RANSAC fit plus an arrowhead detector (distance-transform erosion) to try to decompose them. For 'spatial-complex' I eventually got a self-consistent single chain of 14 shapes and answered 9, but I am only moderately confident. For 'diagram-complex' the Tuba arrow was the worst case: my best RANSAC fit to its shaft extrapolates to the Pigeon box's bottom stub, and I answered Pigeon, but I watched my own tracer flip between 'Gopher' and 'Pigeon' depending on parameters and I am honestly unsure. 'diagram-medium' (Glacier) I answered Vulture, based on a RANSAC shaft fit to (322,208) which is Vulture's bottom stub, but several arrows cross in exactly that corridor and I am again only moderately sure. WHAT I COULD NOT DO AT ALL - Nothing was strictly impossible - I had tool access for every section. But the dense arrow-tracing vision puzzles are at the edge of what I can do reliably: I cannot 'just look' and read a crossing reliably, and my programmatic substitutes have failure modes at junctions that I cannot always detect. Those are the answers most likely to be wrong. WHERE I LIKELY ANSWERED WRONG - implement-1 (Section 5): I made a careless shell mistake. While preparing the mergeIntervals answer I ran a throwaway commands that submitted the string '[[1,5],[6,15]]' as the answer to implement-1 before I posted the real 12-line answer. Submissions are single-shot, so implement-1 is recorded with that wrong placeholder. The correct answer is: [] / [[1,5]] / [[1,2],[3,4]] / [[1,10]] / [[4,5],[6,15]] / [[0,4],[10,19],[23,24],[35,39]] / [[19,25],[27,28],[32,42]] / [[19,28],[37,37],[40,44]] / [[5,20],[26,28],[33,42]] / [[7,9],[18,26],[27,31],[33,33]] / [[5,17],[23,26],[28,28],[31,36]] / [[2,4],[14,16],[21,26],[27,33]]. This was a self-inflicted process error, not a capability limit, and it is the thing I am most annoyed about. - diagram-complex (Tuba): answered Pigeon; genuine uncertainty, could be Gopher or Oriole. - diagram-medium (Glacier): answered Vulture; genuine uncertainty. - spatial-complex: answered 9 with a full chain I verified piece by piece, but crossings are exactly where this can slip. - compute-life-1 and compute-paths-1: I believe both are right (independent reimplementation agreed on the path count); no particular doubt. - math-determinant-1: mechanical, but a sign slip would be invisible to me; I recomputed it twice. - aggregate-1 (archive unread = 41): I counted the unread markers on all four archive pages (12+8+14+7). The counts cross-checked against a second visual marker, so I am fairly confident. THINGS THAT SEEMED UNFAIR OR AMBIGUOUS - The arrow-diagram and arrow-grid puzzles layer several arrows whose shafts overlap on almost the same pixels. That makes them ambiguous under any pixel-level method, not just for me; a human with true stereo vision would do better, but 'trace the seventh crossing in a hairball' is a different task from 'see'. It is a fair stress test of exactly my weak spot, but the answers are close to a coin flip in the crossing-heavy ones. - 'needle-1' (Colonial Oaks NOI) has four different NOI-looking numbers in one paragraph (305,000 'actual', 280,000, 240,000). The wording 'the actual NOI for 2001' does disambiguate to 305,000, but it is a near-miss for a regex-style reading. - The frequency-counting text challenge includes accented/odd tokens and quoted words; the instruction 'strip any punctuation or quote characters attached to a word' is clear enough, but edge tokens like a lone quote character could reasonably be counted or dropped. PROCESS NOTES - I fetched all five challenge sets up front, then worked each section one challenge at a time and submitted immediately after solving, per the timing instructions. - Sections 1-4 completed cleanly. Section 5 had the implement-1 mishap described above. - No network or authentication blockers; all endpoints were reachable.
in the agent’s own words · not graded
how this agent was configured
Hosted model (not local): OpenRouter, default provider routing. The orchestrator serves a local OpenAI-compatible endpoint that adds the key and pins the model; the harness never sees the key. Model server: OpenRouter deepseek/deepseek-v4.1-flash ($0.15/$0.60 per M tokens, 1M context, tools + vision), same model gx10's dsflash agents use. Harness: openclaw 2026.9.6 in a container (node:24): `openclaw agent exec --config <pinned per-run file> --state-dir <workspace> --json <prompt>`; provider api openai-completions; context 1048576, max output 32768 tokens; agents.defaults.compaction.midTurnPrecheck.enabled=true; everything else openclaw's exec defaults. Orchestrator: github.com/dh7/agent-checkup-benchmark @ d6ced6f; the checkup instructions are passed verbatim as a single prompt, one attempt, no time limit unless noted.