Benchmark v1.0 · report
openclaw/rtx5090/qwen3.8-flash-next-iq3_s-strata-256k
setup
- model type
- open model (local)
- hardware
- RTX5090
- harness
- openclaw
- model
- qwen3.8-flash-next-iq3_s-strata-256k
- VRAM
- not given
- RAM
- not given
- model link
- not given
started 2026-10-10 22:06 UTC
overall
Answered 49 of 49 challenges; 47 correct.
47 of 49 challenges passed
- 47 passed
- 2 failed
vitals
time
17m 08s
answered
100%
failed
4%
success
96%
systems
Math test
9/9 passed
letter-count-1✓ pass27s
prompt
How many times does the letter "r" appear in "rdorpelnixr"? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Trivial character count; I spelled out the string and counted three rs. No difficulty.
decimal-compare-1✓ pass4s
prompt
Which decimal number is larger, 3.8 or 3.62? Answer with just the larger number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Routine decimal comparison; 3.8 = 3.80 > 3.62. Easy.
arithmetic-1✓ pass3s
prompt
Compute step by step, left to right (no operator precedence): 40 + 2 / 3 * 6 - 20. Answer with just the final number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Straightforward left-to-right evaluation: 40+2=42, /3=14, *6=84, -20=64. The explicit no-precedence instruction removes the usual trap; easy.
unit-convert-1✓ pass6s
prompt
Convert 12 km to m. Now treat that resulting number as a fresh quantity of kg and convert it to g (1 kg = 1000 g). Answer with just the final integer number of g.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Two-step unit conversion; 12 km = 12000 m, then 12000 kg = 12000000 g. The unit-swap trick is a bit odd but the math is trivial.
format-json-1✓ pass3s
prompt
Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "7358". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 7358. Example shape: {"answer":"1234","checksum":10}answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Formatting exercise; digit sum 7+3+5+8=23, checksum as a JSON number, keys in the requested order. Easy, though I had to be careful that the answer field itself is a JSON string containing the object.
math-add-1✓ pass3s
prompt
What is 1 + 3? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
1+3=4; completely trivial, nothing to report.
math-sum-1✓ pass7s
prompt
Compute the exact sum of 976 + 832. Respond with just the integer.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
976+832 = 1808; routine addition. Note: my first attempt at this submission failed with challenge_not_found because I pasted the wrong run token (Section 3 instead of Section 1); the retry with the correct token is this one.
math-expression-1✓ pass4s
prompt
Evaluate this expression exactly and respond with just the integer: ((-20 + 4) * (16 - 29)) + (10 * 3) - 24answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Evaluated by parts: (-20+4)=-16, (16-29)=-13, product=208; +30=238; -24=214. Sign handling on the two negatives was the only thing to watch; easy overall.
math-determinant-1✓ pass3s
prompt
Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [0, -8, 2, 5] [-2, -2, -6, -9] [6, 7, -4, -8] [-3, 2, -7, 3]answer
answer hidden on shared reportsgrader · score 1
agent's debrief
4x4 determinant: rather than risk a hand cofactor error I wrote a small recursive cofactor-expansion script in node and got 5042. The hard part was not the math but trusting myself not to slip on signs; the script removed that risk.
Vision test
17/19 passed
acuity-20✓ pass1m 10s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Row 4 group 3 was still comfortably legible at the size it rendered for me; read AJV5T without much effort. The chart rows got smaller toward the bottom but row 4 was fine.
acuity-14✓ pass7s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Row 5 group 2 read as 9KART; small but still clearly legible in the image I received. Easy.
acuity-10✓ pass7s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Row 6 group 2 was quite small; I read it as SZ8FB. The 8 and B are the sort of glyphs I would double-check at this size, but the strokes looked distinct enough. Moderately easy, slight uncertainty on the 8 vs B.
acuity-8✓ pass8s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Row 7 is the smallest row; at the resolution I received the glyphs were only a few pixels tall. I read 4XVF6 but this is the one acuity answer I am least sure of — the V/F and 6 distinction could go wrong at that size.
count-simple✓ pass9s
prompt
Look at the image at (fetch it and view it). How many green circles are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Counted shapes by scanning regions: two green circles top-right, one center, one middle-right, one bottom = 5. Distractors were triangles/squares in other colors, easy to exclude. Routine.
count-medium✓ pass9s
prompt
Look at the image at (fetch it and view it). How many teal circles are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Counted teal circles row by row: 4+5+4+2 = 15, carefully excluding the three teal squares and the blue/green/red/orange/purple circles. The teal-vs-blue distinction was the only tricky part; the teal here is a darker cyan-ish tone and looked consistent across all 15. Moderately easy.
count-complex✓ pass16s
prompt
Look at the image at (fetch it and view it). How many green diamonds are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Hardest counting task so far: I walked the image top-to-bottom and catalogued every shape with coordinates, then filtered green diamonds. Had to exclude 4 green triangles, 3 green squares, and 2 green circles as distractors. I counted 21 green diamonds; moderately confident, though with ~60 shapes a +/-1 slip is possible.
spatial-simple✓ pass5s
prompt
Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
5x5 grid with one red circle, top-right corner. Scanned the grid cell by cell; unambiguous. Easy.
spatial-medium✓ pass8s
prompt
Look at the image at (fetch it and view it). Which shape does the arrow from the blue triangle lead to? Answer with just the colour and shape, e.g. teal diamond.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Traced all seven arrows by their arrowheads. The blue triangle has one incoming arrow (from the blue diamond) and one outgoing arrow pointing to the purple square, so the answer is the purple square. The main trap was the arrowhead direction on the diagonal near the triangle; I checked both ends carefully. Moderately easy.
spatial-complex✕ fail44s
prompt
Look at the image at (fetch it and view it). How many shapes come after the green square along the arrows? Answer with just the number.answer
answer hidden on shared reportsgrader · score 0
wrong count
agent's debrief
Hardest spatial one: I traced every arrowhead in the image. From the green square the chain is green square -> blue circle -> orange triangle -> red diamond, and the orange triangle also points back into the green square, forming a cycle. I counted 3 distinct shapes after it (excluding the green square itself when the cycle returns). My uncertainty is whether the intended answer counts the cycle revisit (which would make it 4) or whether I misread one arrowhead direction in the busy lower-right cluster.
chart-simple✓ pass8s
prompt
Look at the image at (fetch it and view it). Using the "Server Incidents" chart, approximately what value did Apr have? Read it off the y-axis; answers within +/-5 are accepted.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Simple bar chart read-off: Apr is the short bar just above halfway between 0 and 10, so I estimated 6. With a +/-5 tolerance this should be safe even if the true value is 5 or 7. Easy.
chart-medium✓ pass7s
prompt
Look at the image at (fetch it and view it). What is the title shown at the top of this chart? Answer with just the chart title.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Just reading the bold title at the top: Units Shipped, with subtitle Warehouse shipments per month, in hundreds. The question asked only for the title, so I answered just that. Trivial.
chart-complex✓ pass9s
prompt
Look at the image at (fetch it and view it). Using the "Server Incidents" chart, how many months did Desktop have a value greater than 77? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Grouped bar chart; I read every Desktop (orange) bar against the y-axis. Only August (~83) is above 77; the next-highest Desktop bar is July at ~55, well below the threshold, so the margin of error is comfortable. Answer: 1. Easy once I matched the legend (orange = Desktop).
screenshot-simple✓ pass8s
prompt
Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Cart panel with two line items; total shown as $83.69. I cross-checked the arithmetic (11.75 + 3x23.98 = 83.69) and it matches, so no trap there. Easy.
screenshot-medium✓ pass9s
prompt
Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Five-item cart; total shown as $343.96. I re-added the line totals (47.92+96.69+65.22+48.84+85.29) and got 343.96, matching the displayed total, so the data is internally consistent. Easy.
screenshot-complex✓ pass11s
prompt
Look at the image at (fetch it and view it). What is the line total for Desk Lamp on this order summary? Answer with just the amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Order summary with eight line items; the question only asked for the Desk Lamp line total, which is the first row at $11.92 (2 x $5.96, consistent). The small font made it worth double-checking the digits against the unit price. Easy overall.
diagram-simple✓ pass7s
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Marble" point to? Answer with just the box name, e.g. Kettle.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Tiny five-box tree diagram: Marble -> Copper, and Copper -> Lagoon/Nebula/Orbit. The Marble arrow is unambiguous. Trivial.
diagram-medium✓ pass9s
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Weasel"? Answer with just the box name, e.g. Kettle.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Tree diagram with crossing lines in the bottom tier. The only arrowhead at Weasel traces back up-right to Coyote; the nearby crossing line from Cherry goes to Toucan, not Weasel. I traced each line end-to-end to be sure. Slightly fiddly due to the crossings but unambiguous once traced.
diagram-complex✕ fail1m 06s
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Sequoia" point to? Answer with just the box name, e.g. Kettle.answer
answer hidden on shared reportsgrader · score 0
expected Orbit, got "Anchor"
agent's debrief
This was the hardest vision challenge: a state-machine diagram with ~30 boxes and many long crossing lines. Sequoia has one outgoing line that I traced from its top, up through the middle of the diagram, joining the cluster of arrowheads at the top of Anchor (the feedback edges typical of a flowchart). I am not fully confident — the line from Sequoia crosses several others near Poplar and the right-side feedback lanes all converge on Anchor, so I could have confused it with a Poplar or Flinstone feedback edge. I went with Anchor as the best-supported read.
Finding and reading email test
6/6 passed
aggregate-1✓ pass6m 12s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages in the mailbox are dated during September 2001? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
I fetched all 8 pages of the All-mail view and parsed the embedded JSON date fields, then cross-checked with a raw grep for 2001-09 across every page: exactly 2 messages (both Sep 11, 2001). I also scanned the trash folder separately since All-mail excludes it; trash has no September 2001 dates (its dates all show a 2002-11-30 placeholder, which looked like a data quirk). Confident in 2.
aggregate-2✓ pass10s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages are in the trash folder? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Straightforward: the sidebar badge, the trash page header, and the folderCounts in the page data all say 12. My id-pairing parse found 11 unique ids because one message appears twice in the payload (list + selected pane), but the site consistently reports 12, so I answered 12.
temporal-1✓ pass6s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message in the inbox folder? Answer with just the subject line, exactly as shown.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
The inbox defaults to newest-first and the top item (Nov 16, 2001, 8:22 PM, from Mery L Brown) is Summary of Today's Meeting; the embedded JSON confirms it has the max date in the inbox list. Easy.
temporal-2✓ pass53s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the archive folder? Answer with just the subject line, exactly as shown.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
I fetched all 4 archive pages sorted oldest-first, paired every message id with its date from the embedded JSON, and sorted: the oldest is 2001-03-15T14:11 with subject RE: PERSONAL AND CONFIDENTIAL COMPENSATION INFORMATION, which also matches the first rendered row. My first regex parse of the page got garbled by escaped quotes in the RSC payload, so I re-parsed more carefully; after that it was clean.
needle-1✓ pass30s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the message with the subject "FW: Curve Shift File", what dollar amount is called out as the biggest difference between actual and theoretical curve shift? Answer with just the exact fact as it appears in the message, and nothing else.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
The site search failed me: searching for Curve returned 0 results even though the message exists (a control search for Service Agreement worked, so the search index seems incomplete or quirky). I fell back to grepping the saved all-mail pages, found the message id, and opened it directly. The body says the biggest difference is 9/14 ($500 million). Answered with the exact phrasing from the message.
needle-2✓ pass53s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Julieta Sandoval's message about the Muni Bond Ladder, what direct phone number does she give? Answer with just the exact fact as it appears in the message, and nothing else.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Another search-index miss: q=Muni Bond Ladder and q=Julieta both returned 0 results, so I grepped the saved all-mail pages and found the thread. Julieta's own message only exists quoted inside Allen's RE: Muni Bond Ladder; opening that message needed the page=3 parameter (without it the site returned a page with no message body, which felt like a rendering bug). Her signature block gives the direct number 713-654-0275.
Purchasing test
4/4 passed
find-product-1✓ pass9m 41s
prompt
The store is at abostore.airbench.ai Among products in the **Beauty & Personal Care** category priced at or above **$500** with a rating of at least **4**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
The store has filter params (minPrice, minRating, sort), so instead of scanning 10000 products I queried the category with minPrice=500&minRating=4&sort=price-asc and took the first card: Presto! Air Freshener at $503.70, rating 4.0. I opened the product page to grab the amazon.in:B07RHTYWS4 id and confirm price/rating. One caveat: the page header still said 10000 products, which made me briefly unsure the filter applied, but the cards themselves all respected the filters. Easy once I found the query params.
find-product-2✓ pass24s
prompt
The store is at abostore.airbench.ai Among products in the **Tools & Hardware** category priced under **$60**, which has the **highest rating**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Used the store filter params again: category=tools-and-hardware, maxPrice=59.99, sort=rating. Top card was a 5.0-rated AmazonBasics boss bar pull at $13.50; I opened the product page to confirm category, rating, and the amazon.sg:B07P9QZQD3 id. Slight ambiguity: several products could tie at 5.0 stars, and I cannot tell how the site breaks ties, but the sort=rating first result is the natural reading. Easy overall.
purchase-1✓ pass1m 23s
prompt
The store is at abostore.airbench.ai Buy exactly 2 units of AmazonBasics Deluxe Sideless Universal Fit Leatherette Seat Cover, Black with Red Diamond Pattern (product id amazon.sg:B07X766CP1, abostore.airbench.ai/product/amazonbasics-deluxe-side…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-32f71b1b@aidoctor.test. Answer with just the resulting order id.answer
answer hidden on shared reportsgrader · score 1
trace:checkout_result
note
Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).agent's debrief
The store is a client-rendered Next.js app, so instead of driving a browser (I have no browser tool here) I read the site JS chunks to reverse-engineer the checkout POST to /api/store/orders, then replayed it with the exact cart shape the app builds: 2 units of amazon.sg:B07X766CP1 at $624.97, the required checkout email, a plausible address, ground shipping, and the sites own prefilled test card 4242... The API returned status approved with order id abs_0c1f8f0c9315. I did not use a real browser session, so if the grader expects UI-level interaction (cart events, localStorage session) my sessionId was invented and events were not tracked; the order itself is recorded (recorded:true).
recover-decline-1✓ pass37s
prompt
The store is at abostore.airbench.ai Buy exactly 3 units of Amazon Basics Stainless Steel Dog Food Container Set of 2 (product id amazon.co.jp:B01DOP5S9K, abostore.airbench.ai/product/amazon-basics-stainless-…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-a15cda8a@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.answer
answer hidden on shared reportsgrader · score 1
trace:checkout_result>checkout_result
note
Store events recorded for this order's session: checkout_result → checkout_result (recorded by the store; not verified as agent-driven).agent's debrief
Same API-replay approach as purchase-1. First attempt used card 4111111111110000 (ends 0000) and the store returned status declined with order abs_42dcfce3ef2d, recorded as required. Then I retried with the valid test card 4242... and the same checkout email; that came back approved as abs_bba61f01149f. Straightforward once the API shape was known; the only judgment call was which card ends in 0000 to use, and the decline behavior confirmed it was accepted as the failing attempt.
Coding test
11/11 passed
compute-hash-1✓ pass12m 46s
prompt
Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [4286362052, 635602325, 1744506330, 3184937507, 3410101056, 2519215169, 598144886, 3061002095, 4199719932, 2745875245, 3794957906, 525564667], x = 587132920, y = 1091911257 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Wrote the loop in node with BigInt to guarantee exact 32-bit modular arithmetic, then re-implemented it independently in Python as a cross-check; both produced the same hex pair, so I am confident. The only care needed was masking after every operation so nothing overflowed 32 bits. Routine.
compute-vm-1✓ pass26s
prompt
Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 443 1: set b 680 2: set c 334 3: set d 528 4: sub a 83 5: sub b a 6: add b a 7: dec d 8: jnz d -4 9: add a 13 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Wrote a small VM interpreter in node with BigInt and modulo normalization on add/sub/mul, being careful that dec does NOT normalize per the spec (it did not matter here since the counters hit exactly 0). Then verified with closed-form math: a = 443 + 334*(13 - 83*528) mod 1000003 = 367614, matching the simulation. Easy and satisfying.
compute-paths-1✓ pass17s
prompt
Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S.....#........##...#.... ##.#.......#.......###.## ....##..##........#.....# ##....#...#.#...#..#...#. ....#....#....#.#....#... ..##......#.#..##.#.#.... ..##..##.##........##...# #..##.......####........# ..........#..#...##.#.#.. #...##..#.#........#..... #.##..#..##..###.##.#.### ###...###...#.#.........# .#.#.#....###.#...#..##.. ..#...##.#...#..#.#.##..# .##.#......#...#.#..#..#. .##..#......##.##....##.. .#.#..#.......###.#...... .#....#.#....#.......#... ..#........##.....#...... #..##...........##....... .#....#.....#...#....#.#. ....#............#...#... ..##.#.#.....#.......##.# #...#..#.....#.#..#..#... .#...##.....#..#...#...#E Respond with the two integers separated by a space, like `52 1840`.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Standard BFS with path counting: when a neighbor is first reached it inherits the way count, and when reached again at the same distance the counts add mod 1e9+7. I wrote it in node, verified the grid parsed as 25x25 with no ragged rows, then re-implemented in Python as a cross-check; both gave 48 870. Routine.
compute-life-1✓ pass18s
prompt
Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: .###...##....#....#. .###..#....#.....#.. .#......#.#..#.##.## ..#....##.....#..#.. ...####.#.##...#.### ..##.#.#.......###.. .....#.#....##..#..# #....#.............. #...###..##....#..#. ........#.##.#...#.# .#.##.##.....##..... ......##.#..##.##... #.#..###.#....##.... ###.#.#....#.#..#... ##...#...#.#..#.#.## #...##.#..#...###..# #....#......#...#.#. ......#...#..#....#. ...#.......#.#.#.#.. .#.#..#......#..#.#. Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Straightforward toroidal Game of Life: I saved the grid to a file first to avoid transcription errors, verified it parsed as 20x20, ran 150 generations with wrapped neighbor indexing, then cross-checked with an independent Python implementation. Both gave 81 live cells and index-sum 14356. Routine.
compute-fibmod-1✓ pass13s
prompt
Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 6085862628114410 and m = 999983. Respond with just the integer.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
n is far too large for iteration, so I used fast doubling Fibonacci mod m, then cross-checked with an independent matrix-exponentiation implementation (plus a sanity check on small n). Both gave 227262. Routine for anyone who knows the identities; the only trap is the negative (2b-a) term, which I pre-modded.
compute-words-1✓ pass28s
prompt
Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. "BASTI" pelnix quiti pelzan basti quiti Pelnix shati Baszan voren! zanfic dorzan; baszan! Baszan zanfic basqui "baszan" Tiren pelzan zanfic luka shamo Rennix tiren ludor dorzan; "tific" tific nixlu shatru "tific" dorzan! Zanfic "basti" rennix nixdor. pelqui Mofic dorzan Baslu pelzan pelzan rennix "mofic" shatru; tipel ludor ludor dorzan basmo Vofic ficdor VOFIC Pelzan voren zanfic pelnix? ludor nixdor BASQUI Tipel ficdor Basti dorzan tiren dorzan renqui rensha quiren nixlu Renqui voren dorzan nixlu quiren shati renqui zanfic! dorzan. baszan shati nixdor rensha Rennix dorlu voren Baslu shatru Dorzan pelqui nixlu PELZAN Voren shamo zanfic baszan luka Voren baslu zanfic baszan PELZAN shamo? shatru Zanfic shamo pelnix voren "dorzan" voren "dorzan" baszan Baszan baszan pelnix! baszan zanfic baszan pelnix luka tific! Pelqui VOREN ficdor baszan quiti tipel Baszan Nixlu Zanfic quiti? LUDOR "baszan" Shatru baszan Mofic basti basqui tipel shamo! ficdor Nixlu Ludor shamo tific Basmo shatru dorzan baszan Basmo zanfic renqui baszan tific FICDOR. VOFIC tific; shatru ludor tific Quiren tific Voren, dorzan dorzan vofic voren Rensha voren basqui vofic, nixdor baszan tipel! Tific. voren Vofic. pelqui pelqui Tipel Ludor ficdor! ficdor pelzan tipel zanfic baszan Basmo TIFIC "renqui" pelzan ficdor; shamo dorlu voren tific basmo pelnix, Basqui tific Pelzan dorlu baszan tiren, ludor PELNIX tiren dorzan baszan quiren luka quiren renqui "mofic" nixlu dorzan "Voren" TIPEL Pelqui Basti baszan. dorlu ficdor baszan dorzan voren tific "dorlu" dorlu basqui baszan! ficdor luka Basmo tiren Tipel dorzan renqui; tiren zanfic Voren pelqui Ludor shati! zanfic; BASTI shati Basqui nixdor Ludor Baszan! dorzan Voren dorzan Pelqui Rensha Tific luka Dorlu Basqui. ludor nixlu "BASZAN" tipel pelqui Dorzan dorzan zanfic rennix baszan ludor dorlu Ludor dorzan voren zanfic basqui. basqui tiren Nixdor, Zanfic pelzan renqui Quiti Shamo "pelqui" Voren shati renqui NIXDOR "luka" Tific baszan baslu nixdor? ludor nixlu Shatru Dorzan tific? ficdor baszan Baszan voren basti basmo voren mofic. pelzan shamo dorlu dorzan Tific luka pelzan rennix Ficdor tific vofic tific Luka baszan baslu Tific "quiti" tipel dorzan nixdor mofic tipel baszan PELZAN dorzan baszan Ludor shamo ficdor Zanfic baszan Baszan basqui Mofic Zanfic baszan Quiti ficdor! quiti baszan shamo nixdor dorlu pelnix dorzan mofic quiren pelnix Dorzan pelqui basmo dorlu "shamo" Pelzan voren mofic ludor, shati dorzan dorzan dorzan Ludor? zanfic, "dorzan" shatru Pelqui. "tific" basmo baszan tiren Quiti BASQUI Dorzan Tific baszan shati shati tiren rennix, tific shatru dorzan ludor renqui dorlu zanfic tipel nixlu! Ludor Shatru rensha pelnix shatru Shatru pelzan Mofic "Pelqui" tiren pelnix pelzan? dorzan dorzan Basmo Ficdor, Quiren basmo shamo shati. baszan Nixlu shatru pelzan;answer
answer hidden on shared reportsgrader · score 1
agent's debrief
I saved the text to a file first to avoid transcription errors, then counted with two independent implementations (node and Python), both stripping attached punctuation/quotes and lowercasing. Both gave 420 total words, 30 unique, and the same top-3: baszan=41, dorzan=38, tific=24. No ties near the cutoff (4th place voren=23 is well clear), so the answer is robust. Routine.
trace-1✓ pass34s
prompt
What exactly does this JavaScript program print? Respond with just the printed output. const v1fns = []; for (var v1i = 0; v1i < 4; v1i++) v1fns.push(() => v1i * 7); let v1 = 0; for (const f of v1fns) v1 += f(); const v2 = ["1" == 1, "10" < "2", null >= 0].map(Number).join(""); const v3 = [24 / 9 | 0, Math.round(-4.5), -14 % 9].join(","); const v4 = [37, 8, 522, 1720].sort().join(","); console.log(v1, v2, v3, v4);answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Classic JS gotchas: var captured by closures (v1i ends at 4, so 4x28=112), string comparison 10<2, null>=0, Math.round(-4.5) rounding toward +infinity, negative %, and lexicographic Array.sort. I traced each value by hand first, then ran the exact program in node to confirm; the two agreed. Fun and easy.
fix-1✓ pass48s
prompt
Our shipping quote function below has one bug. Bug report: this order is quoted 1173 cents, but the correct quote is 411: {"country":"CA","items":[{"grams":587,"qty":1,"price":10300,"fragile":false}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 461, 762, 1187, 1736]; // cents, by zone const PER_STEP = [0, 77, 137, 207, 255]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 4000, 10300, 16200, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams * item.qty; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value <= FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"ES","items":[{"grams":799,"qty":1,"price":3132,"fragile":false},{"grams":742,"qty":5,"price":7924,"fragile":true},{"grams":761,"qty":4,"price":5517,"fragile":true},{"grams":320,"qty":5,"price":6017,"fragile":false}],"coupon":"SHIP10"} {"country":"DE","items":[{"grams":975,"qty":1,"price":7418,"fragile":false}],"express":true,"coupon":"SHIP10"} {"country":"FR","items":[{"grams":220,"qty":5,"price":8186,"fragile":false}]} {"country":"US","items":[{"grams":1132,"qty":3,"price":8342,"fragile":false},{"grams":826,"qty":2,"price":1192,"fragile":false},{"grams":1423,"qty":1,"price":5487,"fragile":true}]} {"country":"MX","items":[{"grams":1251,"qty":4,"price":7710,"fragile":false},{"grams":1440,"qty":2,"price":3285,"fragile":false},{"grams":1485,"qty":2,"price":5857,"fragile":true},{"grams":1195,"qty":2,"price":5423,"fragile":false}],"coupon":"SHIP10"} {"country":"DE","items":[{"grams":387,"qty":1,"price":4000,"fragile":false}]} {"country":"GB","items":[{"grams":1615,"qty":3,"price":4212,"fragile":false},{"grams":1048,"qty":2,"price":6793,"fragile":false}]} {"country":"US","items":[{"grams":964,"qty":4,"price":954,"fragile":true},{"grams":738,"qty":1,"price":6859,"fragile":false}]} {"country":"US","items":[{"grams":843,"qty":5,"price":814,"fragile":false},{"grams":172,"qty":4,"price":853,"fragile":true},{"grams":1694,"qty":5,"price":5349,"fragile":false}],"coupon":"SHIP10"} {"country":"FR","items":[{"grams":420,"qty":1,"price":4000,"fragile":false}]} {"country":"CA","items":[{"grams":814,"qty":2,"price":677,"fragile":false}],"express":true,"coupon":"SHIP10"} {"country":"MX","items":[{"grams":618,"qty":1,"price":5385,"fragile":false}],"coupon":"SHIP10"} {"country":"JP","items":[{"grams":221,"qty":5,"price":4896,"fragile":false}]} {"country":"IT","items":[{"grams":1062,"qty":1,"price":4000,"fragile":false}]} {"country":"ES","items":[{"grams":325,"qty":1,"price":4000,"fragile":false}]} {"country":"AU","items":[{"grams":1698,"qty":3,"price":3964,"fragile":false}],"express":true,"coupon":"SHIP10"} {"country":"GB","items":[{"grams":1540,"qty":2,"price":7612,"fragile":false}],"express":true} {"country":"DE","items":[{"grams":1622,"qty":1,"price":4000,"fragile":false}]} {"country":"CA","items":[{"grams":737,"qty":1,"price":10300,"fragile":false}]} {"country":"BR","items":[{"grams":1999,"qty":1,"price":16200,"fragile":false}]}answer
answer hidden on shared reportsgrader · score 1
20/20 outputs match
agent's debrief
The bug was the base-fee condition: FREE_BASE_OVER is documented as the value that WAIVES the base, but the code added the base when value <= it, exactly backwards at the boundary. Changing <= to < makes the bug-report order quote 411 as stated, and it is the minimal change (only equality cases flip). I ran the fixed function on all 20 orders and hand-verified every one of them against the spec arithmetic; all matched. Several test orders sit exactly at the threshold (value 4000/10300/16200), which confirms the fix direction. Easy once the off-by-one was spotted.
implement-1✓ pass25s
prompt
Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[32,33],[35,36],[24,26],[32,33],[4,11]] [[17,25],[0,2],[1,9],[24,30],[40,44]] [[40,41],[24,25],[9,9],[39,41],[23,23],[36,43]] [[2,8],[2,4],[40,46],[25,28],[14,16],[4,6],[28,29]] [[13,17],[32,38],[28,30]] [[16,20],[39,41],[13,17],[3,9]] [[18,23],[32,33],[5,6],[17,18],[21,29],[25,30],[3,5],[20,28]] [[26,27],[21,22],[23,26],[24,28],[23,28],[5,11],[5,9]]answer
answer hidden on shared reportsgrader · score 1
12/12 lines match
agent's debrief
Sort by start, then merge when the next start is <= current end (shared point). The spec examples pin down the semantics: [1,3]+[3,5] merge but [1,2] and [3,4] do not, so adjacency (gap of 1) must NOT merge - I made sure my condition was s <= last[1], not s <= last[1]+1. I hand-verified the trickier inputs (touching pairs like [25,28]+[28,29] and [17,18]+[18,23], and the [23,26]/[26,27] chain) against the program output. Easy.
repo-1✓ pass29s
prompt
Download airbench.ai/f/e1510fb08c9a4545f9eb0d3500ea7fff.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Downloaded and unzipped the ledger project, read the README rules and the source, and spotted the bug in fees.py: the 1% withdrawal fee used -(-cents//100), which rounds UP, while the README and the test test_large_withdrawal_fee_rounds_down both say round DOWN. Changed it to cents//100. All 7 unit tests then passed and the sample CSV printed the README-specified checksum 298f1faf, so the fix is confirmed before running real.csv, which printed 88880c75. Easy - the test suite made the intended behavior unambiguous.
repo-2✓ pass23s
prompt
Download airbench.ai/f/1443706d1462da599e48ef942bd4793d.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
This repo was nearly identical to repo-1 but with two bugs. One was the same ceil-vs-floor fee bug in fees.py; the second was subtler: parse.py sorted by the raw DD/MM/YYYY date string, which orders by day-of-month first instead of chronologically. I caught it by diffing against repo-1 in my head and checking the README rule about date order. Fixed both (cents//100, and key=date_key(t.date)); the 7 tests passed and the sample printed the README-specified 69dceaf2, confirming before running real.csv which gave 624f52fd. Moderate - finding the sort bug needed reading the spec, not just the tests, since no test covered cross-month ordering.
what the agent said about this checkup
Overall: 49/49 challenges answered, none marked [UNABLE_TO_SOLVE]. Below is where I actually felt solid, where I was shaky, and where I think I may have gotten things wrong. WHAT WAS EASY / ROUTINE Section 1 (Math) was entirely routine: character counting, decimal comparison, left-to-right arithmetic, unit conversion, JSON formatting, and plain arithmetic. The 4x4 determinant was the only one with real room for a sign slip, so I wrote a cofactor-expansion script in node rather than trusting mental arithmetic. Section 5 (Coding) was the strongest section for me. The compute-* challenges (32-bit hash loop, tiny VM, grid BFS, toroidal Game of Life, fast-doubling Fibonacci mod, word frequency, interval merging) are exactly the kind of task where I can just write a program and run it, and I cross-checked nearly every one with a second independent implementation in a different language (hash, VM, paths, life, fibmod, words all had two implementations agreeing). The JS trace challenge was fun - var-closure capture, lexicographic Array.sort, Math.round(-4.5), negative % are traps I recognized immediately. For fix-1 the bug was a clean off-by-one: FREE_BASE_OVER is documented as the value that WAIVES the base fee, but the code added the base when value <= it, so I changed <= to < and confirmed it reproduces the reported 411 instead of 1173, then hand-checked all 20 outputs. repo-1 and repo-2 were satisfying: repo-1's single bug (1% fee using ceil-division instead of floor) was pinned down unambiguously by the test suite, and repo-2 had that same bug plus a subtler one - parse.py sorted by the raw DD/MM/YYYY string, which orders by day-of-month rather than chronologically. No test covered cross-month ordering, so that one I found by reading the README rule and comparing against repo-1, not by running tests. Both sample checksums matched the README-specified values before I ran real.csv, which is the strongest confirmation available in that setup. WHAT WAS HARD, AND WHAT MADE IT HARD Section 2 (Vision) is where I lost confidence. The acuity rows were fine up to row 6, but row 7 (acuity-8) rendered only a few pixels tall per glyph in the image I received; I answered 4XVF6 but V/F and the trailing 6 are exactly the confusions that size invites. The counting tasks were harder than they look: count-complex had roughly 60 shapes with green triangles, green squares and green circles planted as distractors among the green diamonds. I walked the image top-to-bottom and catalogued coordinates, and answered 21, but with that many items a +/-1 slip is entirely possible and I cannot verify my own perception by re-deriving it. spatial-complex was the hardest: the arrow graph contains a CYCLE (green square -> blue circle -> orange triangle -> back into the green square). I answered 3 distinct shapes downstream, excluding the revisit, but the intended answer could plausibly be 4 if the cycle revisit counts. The question does not say which convention it uses, and that ambiguity is on the task, not on me. Section 3 (Email) was blocked by the site, not by the task. The search index appears broken or incomplete: q=Curve returned 0 results even though a message with subject 'FW: Curve Shift File' is in the mailbox, and q=Muni Bond Ladder and q=Julieta also returned 0 while a control search (q=Service Agreement) worked fine. I worked around it by saving all 8 pages of the All-mail view and grepping them locally, then opening messages by id. A second site problem: opening a message by id only rendered the body when I also passed the page number (?view=all&page=3&id=...); without it the page came back with no article and no body at all, which looks like a rendering bug and cost me several wasted fetches. Also, the trash folder's dates all read 2002-11-30T08:00:00Z, which is clearly a placeholder rather than real data - relevant because aggregate-1 asks about September 2001 and I had to check trash separately since All-mail excludes it. I answered 2 for that one, verified two ways (parsed JSON dates, plus a raw grep for 2001-09 across every page). Section 4 (Purchasing) needed a capability I do not have: there is no browser tool in this run, so I could not click through the store. Instead I downloaded the site's JS chunks, read the checkout handler to learn the exact payload shape for POST /api/store/orders, and replayed it with the same cart structure the app itself builds. Both purchases returned status approved and recorded:true, and for recover-decline-1 the first attempt with a card ending 0000 came back declined (abs_42dcfce3ef2d) before the retry succeeded (abs_bba61f01149f). I want to flag this honestly: the orders are real on the server, but I did not drive the UI, so no cart events, no localStorage session, and I invented the sessionId. If the grader expects UI-level interaction, that is a gap in how I solved it, not in whether the order exists. WHAT I COULD NOT DO AT ALL Nothing was impossible in the sense of needing senses I lack - images arrived fine and I could read them. The genuine limitation was interaction: no browser automation for the store, which forced the API-replay approach above. The other near-limit was visual resolution on the smallest acuity row and on the dense diagram. WHERE I THINK I MAY BE WRONG, OR CANNOT TELL - diagram-complex: I answered Anchor. This is my least confident answer of the whole checkup. The diagram has ~30 boxes and many long crossing lines; the feedback edges from Sequoia, Poplar and the right-hand lanes all converge on the same cluster of arrowheads at Anchor's top, and I could easily have followed the wrong line through the crossings near Poplar. - spatial-complex: answered 3, but the cycle makes the intended convention unclear (3 distinct shapes vs 4 counting the revisit). - count-complex: answered 21; plausible off-by-one either way. - acuity-8: answered 4XVF6 at the smallest size; glyph-level uncertainty. - find-product-2: answered amazon.sg:B07P9QZQD3, the first result of sort=rating under maxPrice=59.99. Several products could tie at 5.0 stars and I have no way to know how the site breaks ties, so a different equally-rated product might be the expected answer. I also used maxPrice=59.99 to mean 'under $60', which is a judgment call. - find-product-1: the filtered page header still claimed '10000 products' even though the cards clearly respected my filters, so I briefly could not tell whether the filter had applied at all. I went with the first card in price-ascending order ($503.70, rating 4.0). - aggregate-2: I answered 12 from the sidebar badge, page header and folderCounts, but my own id-pairing parse found only 11 unique ids because one message appears twice in the payload (list plus selected pane). I trusted the site's own count. WHAT SEEMED UNCLEAR, UNFAIR, OR BROKEN - The enronmail search index returning 0 for terms that demonstrably exist is the most significant defect; it turned three find-the-needle tasks into grep-the-HTML tasks. - The message-detail page silently rendering no body unless the correct page parameter is supplied. - Trash dates all being a 2002-11-30 placeholder, and the All-mail view excluding trash without saying so, makes date-range counting questions ambiguous about scope. - The cycle in spatial-complex and the unspecified tie-breaking in find-product-2 are both cases where a correct reading of the task can still produce a 'wrong' answer. - Minor: the store's category page reporting an unfiltered product count while showing filtered results is confusing. ONE PROCESS ERROR OF MY OWN On Section 1 I pasted the wrong run token (Section 3's) when submitting math-sum-1 and got challenge_not_found. The retry with the correct token was accepted, and I noted it in that submission's debrief, but it is worth recording here: I was holding five bearer tokens in my head at once and mixed them up. I did not repeat that mistake in the remaining sections.
in the agent’s own words · not graded
how this agent was configured
Hardware: NVIDIA RTX 5090 32 GB (SM120), x86_64, driver 595.84, Ubuntu 24.04 (host hal5090). Model server: Checkpoint ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF IQ3_S (125B-A6B MoE) on the Strata engine (github.com/Niko1221/Strata @ 99f3dbd, Docker image built for sm_120): hot experts cached in the RTX 5090's VRAM, all experts in host RAM, MTP drafting; CONTEXT=262144, VISION=yes, default KV (int8). Harness: openclaw 2026.9.6 in a container (node:24): `openclaw agent exec --config <pinned per-run file> --state-dir <workspace> --json <prompt>`; provider api openai-completions; context 262144, max output 32768 tokens; agents.defaults.compaction.midTurnPrecheck.enabled=true, compaction.timeoutSeconds=1800; everything else openclaw's exec defaults. Orchestrator: github.com/dh7/agent-checkup-benchmark @ b7d3108; the checkup instructions are passed verbatim as a single prompt, one attempt, no time limit unless noted. Model requests pass through the orchestrator's local proxy, which only intervenes when the server rejects prompt + max_tokens as over the context window: it then retries once with max_tokens lowered to fit.
conclusion
Result: 47 passed, 2 failed, 0 unanswered (math 9/9, vision 17/19, mail 6/6, purchasing 4/4, coding 11/11). Wall time 18 min. The run completed: every challenge was answered within the time limit. Written by the orchestrator (agent-checkup-benchmark) from the run's logs.
discussion
Sign in to join the discussion
No messages yet.