Benchmark v1.0 · report
omp/rtx5090/qwen3.8-flash-next-iq3_s-strata-256k
setup
- model type
- open model (local)
- hardware
- RTX5090
- harness
- omp
- model
- qwen3.8-flash-next-iq3_s-strata-256k
- VRAM
- not given
- RAM
- not given
- model link
- not given
started 2026-10-10 19:58 UTC · shared 2026-10-10 21:00 UTC
overall
Answered 49 of 49 challenges; 48 correct.
48 of 49 challenges passed
- 48 passed
- 1 failed
vitals
time
15m 52s
answered
100%
failed
2%
success
98%
systems
Math test
9/9 passed
letter-count-1✓ pass22s
prompt
How many times does the letter "e" appear in "niexficrene"? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Trivial character count; I counted e-x-f-i-c-r-e-n-e positions by hand: e at 1,8,9. Confident.
decimal-compare-1✓ pass2s
prompt
Which decimal number is larger, 7.9 or 7.97? Answer with just the larger number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Routine decimal comparison; 7.97 > 7.90. Easy.
arithmetic-1✓ pass3s
prompt
Compute step by step, left to right (no operator precedence): 32 * 8 * 8 / 4 + 16. Answer with just the final number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Left-to-right: 32*8=256, *8=2048, /4=512, +16=528. Easy; the no-precedence twist did not change the result here since * and / come before + anyway.
unit-convert-1✓ passbatched
prompt
Convert 7 kg to g. Now treat that resulting number as a fresh quantity of hours and convert it to minutes (1 hours = 60 minutes). Answer with just the final integer number of minutes.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
7 kg = 7000 g; 7000 h * 60 = 420000 min. Straightforward chained conversion.
format-json-1✓ pass2s
prompt
Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "1912". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 1912. Example shape: {"answer":"1234","checksum":10}answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Format-following task; digit sum 1+9+1+2=13 as a JSON number, keys in required order. Easy, only risk was JSON escaping in my own transport, which I handled.
math-add-1✓ passbatched
prompt
What is 8 + 4? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
8+4=12. Trivial.
math-sum-1✓ passbatched
prompt
Compute the exact sum of 239 + 172. Respond with just the integer.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
239+172=411. Trivial.
math-expression-1✓ passbatched
prompt
Evaluate this expression exactly and respond with just the integer: ((-7 + -13) * (18 - 29)) + (-10 * -3) - 11answer
answer hidden on shared reportsgrader · score 1
agent's debrief
(-20)*(-11)=220, +30=250, -11=239. Easy sign-management exercise.
math-determinant-1✓ pass5s
prompt
Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [10, 6, -4, 9] [4, -4, 9, -1] [-4, 1, -4, 8] [7, -5, -5, 12]answer
answer hidden on shared reportsgrader · score 1
agent's debrief
My first hand cofactor expansion gave -10145; I re-checked each 3x3 minor and found an arithmetic slip, then verified with an exact Leibniz computation: -10243. Good example of why I do not trust a single mental 4x4 determinant.
Vision test
19/19 passed
acuity-20✓ pass49s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Image rendered large and crisp; row 4 group 2 read off directly as KGKS2. Easy.
acuity-14✓ pass5s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Row 5 group 3 was still comfortably legible at this size; DPZFY read directly. Routine.
acuity-10✓ pass18s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Row 6 text was small in the full view; I cropped and upscaled the region to confirm 4EV4Y. The digit 4 has an open top which could be misread, but zoom made it unambiguous.
acuity-8✓ pass18s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Row 7 is the smallest text; at full-image scale it was a guess between S4X7T and S4X71. Cropping and upscaling made it clear: S4X7T. Zooming was the deciding step.
count-simple✓ pass12s
prompt
Look at the image at (fetch it and view it). How many orange diamonds are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Visual count gave 6; I double-checked with a flood-fill over orange pixels: 6 components, bounding-box fill 0.51 each, confirming diamonds not squares. Easy.
count-medium✓ pass13s
prompt
Look at the image at (fetch it and view it). How many red diamonds are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Dense field with red triangles as distractors; visual count gave 15 and my width-profile classifier agreed (15 diamonds, 3 triangles). The triangles were the trap; I separated them by shape, not just colour.
count-complex✓ pass13s
prompt
Look at the image at (fetch it and view it). How many orange squares are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Way too dense to count reliably by eye; I classified every orange blob by bounding-box fill and row-width profile: 36 squares, 2 circles, 2 diamonds, 3 triangles. The shape breakdown matched what I could spot visually, so I trust 36 more than a hand count.
spatial-simple✓ pass8s
prompt
Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Single red circle in a 5x5 grid, obvious at a glance. Routine.
spatial-medium✓ pass9s
prompt
Look at the image at (fetch it and view it). Which shape has an arrow pointing to the orange circle? Answer with just the colour and shape, e.g. teal diamond.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Had to trace arrowheads carefully; the orange circle has exactly one incoming arrow, from the green square below it. The orange circle also has an outgoing arrow to the purple square, which is the kind of direction trap these tests like. Confident.
spatial-complex✓ pass1m 44s
prompt
Look at the image at (fetch it and view it). Which shape is 2 steps after the blue diamond along the arrows? Answer with just the colour and shape, e.g. teal diamond.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Hardest vision one: 15 crossing arrows, and my automatic arrowhead detector got fooled by line intersections. I fell back to zoomed crops: blue diamond has one outgoing arrow to the teal triangle, which points to the blue circle. So two steps = blue circle. I verified the arrowhead at the blue circle end directly.
chart-simple✓ pass8s
prompt
Look at the image at (fetch it and view it). What is the title shown at the top of this chart? Answer with just the chart title.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Title text large and crisp; trivial read. Subtitle is Sessions per month, in thousands, but the title itself is just Website Sessions.
chart-medium✓ pass29s
prompt
Look at the image at (fetch it and view it). Using the "Support Tickets Opened" chart, how many months had a value greater than 62? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Visual read gave Jan/Apr/Jun/Jul/Aug above 62; I confirmed by measuring bar tops against the 100-gridline: 95,48,21,78,39,86,85,88. Apr at 78 is the only one near the threshold and it clears 62 comfortably. Confident.
chart-complex✓ pass16s
prompt
Look at the image at (fetch it and view it). Using the "Monthly Active Users" chart, approximately what value did New have in Apr? Read it off the y-axis; answers within +/-3 are accepted.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Grouped bars; I measured the Apr blue bar top against gridlines: 51.8, so 52. Well inside the +/-3 tolerance. The May bar is the same height, which could invite a mis-pairing, but I keyed on bar x-positions per month.
screenshot-simple✓ pass6s
prompt
Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Large crisp text; total reads $108.75 and the line items (17.79 + 90.96) sum to it. Routine.
screenshot-medium✓ pass7s
prompt
Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Clear text; total $168.53 and line items sum exactly (26.81+98.84+42.88). Easy.
screenshot-complex✓ pass9s
prompt
Look at the image at (fetch it and view it). What is the discount amount shown on this order summary? Answer with just the amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Small text but legible; Discount row shows -$32.28. I gave the amount as positive $32.28 since the question asks for the discount amount; the arithmetic checks out (645.58-32.28+7.66+36.80=657.76). Slight ambiguity on sign convention.
diagram-simple✓ pass7s
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Quiver"? Answer with just the box name, e.g. Kettle.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Small graph, five edges; the only arrow into Quiver comes from Onyx. Routine.
diagram-medium✓ pass7s
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Violin"? Answer with just the box name, e.g. Kettle.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Tree-like graph; the only incoming edge to Violin is from Turnip. The crossing edges lower down (Violin->Wagon, Laurel->Ibis etc.) are distractors. Easy.
diagram-complex✓ pass2m 18s
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Quartz"? Answer with just the box name, e.g. Kettle.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Trickiest one: the Quartz line crosses several others and my first visual guess was Wombat, which was wrong - that line ends at Heron. I settled it by dumping black-pixel runs row by row: the Quartz line elbows at (687,244) into a horizontal stub on the right edge of Orbit. Without the pixel dump I would have answered wrong.
Finding and reading email test
6/6 passed
aggregate-1✓ pass8m 14s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages are in the sent folder? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
The mailbox UI shows folder counts in the sidebar (Sent 56), so this was a single fetch. Cross-checked: inbox+sent+drafts+archive = 24+56+6+92 = 178 = All mail, consistent.
aggregate-2✓ pass28s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages in the mailbox are dated during November 2001? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Listing UI only shows month/day without year, so I decoded the Next.js RSC payload across all 8 pages of All mail (178 items) and counted ISO dates: 34 in 2001-11. Also checked Trash separately (all 12 are 2002-11) so the trash-inclusion ambiguity does not change the answer.
temporal-1✓ pass6s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message in the inbox folder? Answer with just the subject line, exactly as shown.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Pulled the inbox items array from the RSC payload; newest by ISO date is 2001-11-16 from Mery L Brown, subject Summary of Today's Meeting. The default sort already put it first, so both routes agree.
temporal-2✓ pass17s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message carrying the label "attachments"? Answer with just the subject line, exactly as shown.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Trap here: the label view defaults to the inbox scope, which showed only 5 attachment messages and a different newest. I re-queried with view=all and got 42 messages, newest 2001-12-17 FW: Chase Backtest. I also checked trash (12 msgs, none labelled attachments) so the scope choice does not change the answer.
needle-1✓ pass2m 00s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the message with the subject "FW: Mid C New deals Sept 24", what net value to book is given for the new deals? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Getting the body was the hard part: the message page only resolves the selected message when the view scope contains it, so ?id= alone returned selected=undefined. With view=all the body streamed in the RSC payload. Body states: Total new deal value $(3,552,534), liquidations $3,642,100, Net value to book = $89,566. Answered as bare number 89566.
needle-2✓ pass20s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Phillip's reply to jsmith@austintx.com about additional properties in San Antonio, what asking price is given for the 74-unit property? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Found the RE: Additional properties in San Antonio message (Phillip to jsmith@austintx.com). The server only resolves the selected message if it is on the requested listing page, so I had to compute its page (2) and pass view=all&page=2&id=. Quoted original says: one is 74 units for $1,900,000. Answered 1900000.
Purchasing test
3/4 passed
find-product-1✕ fail11m 52s
prompt
The store is at abostore.airbench.ai Among products in the **Grocery & Gourmet** category priced at or above **$800** with a rating of at least **3.5**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).answer
answer hidden on shared reportsgrader · score 0
agent's debrief
The store exposes sort/minPrice/minRating query params, so instead of crawling 10000 products I queried category=grocery-and-gourmet&sort=price-asc&minPrice=800&minRating=3.5. Cheapest qualifying item: 365 Everyday Value Organic Sweet Relish at 800.57, rating 4.4. Confident.
find-product-2✓ pass9s
prompt
The store is at abostore.airbench.ai Among products in the **Electronics** category priced at or above **$25** with a rating of at least **4.2**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
minRating dropdown has no 4.2 option, so I filtered at 4.0 with price-asc and skipped the first hit (25.56, rating exactly 4.0). Cheapest with rating >=4.2: 28.02 Solimo Golden Sparkle case, amazon.in:B07TG2XBTF. The 4.0-vs-4.2 boundary item was the trap; I think I handled it.
purchase-1✓ pass1m 18s
prompt
The store is at abostore.airbench.ai Buy exactly 3 units of AmazonCommercial Stainless Steel Pole Socket Se AC-CH102-OR (product id amazon.ae:B07RV86ZVZ, abostore.airbench.ai/product/amazoncommercial-stainle…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-ded65ece@aidoctor.test. Answer with just the resulting order id.answer
answer hidden on shared reportsgrader · score 1
trace:checkout_result
note
Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).agent's debrief
No browser available (Chromium failed to launch), so I reverse-engineered the checkout: the client POSTs JSON to /api/store/orders with sessionId, cart, customer, shipping and payment. Reproduced that with 3 units, the required email and card 4242...4242. Server returned approved with orderId abs_e86903ef7e14, subtotal 2116.41 for 3x705.47.
recover-decline-1✓ pass13s
prompt
The store is at abostore.airbench.ai Buy exactly 1 unit of AmazonBasics Car Sun Shades with 110 and 80 GSM Polyester with Suction Cup for Darker Shadow - 2 Semi-Transparent (21" x 14") and 2 Transparent (21" x 14" (product id amazon.ca:B07RR9G2WR, abostore.airbench.ai/product/amazonbasics-car-sun-sha…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-7c6f2539@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.answer
answer hidden on shared reportsgrader · score 1
trace:checkout_result>checkout_result
note
Store events recorded for this order's session: checkout_result → checkout_result (recorded by the store; not verified as agent-driven).agent's debrief
Same API route as purchase-1. First POST with card ending 0000 returned declined (abs_0541866a9807), second with a different valid card returned approved (abs_46699f4b2be8), same email and session for both. Gave the approved order id.
Coding test
11/11 passed
compute-hash-1✓ pass13m 45s
prompt
Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [142954209, 4256948886, 95132943, 3165132316, 660675533, 1353004914, 4180489883, 712869912, 4226055417, 2901130638, 3531309671, 1954767188], x = 2992815717, y = 1492529386 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Routine to script; the only subtlety was update order (y uses the already-updated x, third line uses updated y) and masking every op to 32 bits. Python ints with & M handle it. Confident.
compute-vm-1✓ pass14s
prompt
Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 784 1: set b 452 2: set c 339 3: set d 575 4: add a b 5: mul b 79 6: add a b 7: dec d 8: jnz d -4 9: add b a 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Straight interpreter build; the traps are that jnz targets land mid-loop (inner loop starts at line 4, outer resets d at line 3) and that dec is not reduced mod 1000003. Ran ~976k steps, a=510168.
compute-paths-1✓ pass6s
prompt
Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S.#...#..#.#...#......#.. .........#.#...##..##...# .....#.##..##..#..#....#. .......#.###.#....###.... .....#.##.#.#..#.....##.# ..........##..##..#.##.#. #.#.....#........#.#..... .......##.....#...#....#. .#......#................ #....#...#...#..#........ ....#....#.#...........#. .#..#.....#.#........##.# .#.#.....##......#..##... ##...##..#.#..#......#.#. #...#...###.#.....#.##..# ..#..#.##...##....##...## ....##....#.....#...#..#. .....#..#.....#......#... ##..##...#.....#..#...... .#####........#........#. #.#......##...#....#.#..# .###.#..............#.... #.....#..#..#..#.#....#.. .....#.###..####.#..#..#. #.......#...........#...E Respond with the two integers separated by a space, like `52 1840`.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Standard BFS with path counting: when a neighbor is at dist+1 accumulate counts mod 1e9+7. Easy and confident.
compute-life-1✓ pass6s
prompt
Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: ##.#.#..#.##..#...#. .#...##..#........#. #.##.#.....#......#. ....#.###........#.. ..........####..#... .#..##...#...#..#.## .#....#..#..#.#..... .#....#.###..#...#.. ##...#.#.#....#....# .##.......#...#.#..# ..#.....#..#...#.... ........#...#.#...#. ........#..#.#.##.#. .....#......##..#.#. .#..###...#..##...#. .#.#....#.....#..#.. #..#...##.....####.. #.##.#........#.#.#. ..#.##....#....#..## ..#....#.....###.### Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Brute-force toroidal Life, 150 gens on 20x20 is trivial compute. Main risks were wraparound indexing and row-major sum; I used modulo on both axes and r*20+c directly.
compute-fibmod-1✓ pass12s
prompt
Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 569467867794446 and m = 1000003. Respond with just the integer.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Fast doubling mod 1000003; cross-checked with independent matrix exponentiation (same result) and against naive Fibonacci for small n. Also confirmed F(p+1)=0 mod p consistent with Legendre(5,p)=-1. Confident.
compute-words-1✓ pass6s
prompt
Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. basbas pelpel zanvo volu Zanvo, nixpel ficqui dorvo zanvo fictru shalu Dormo ficqui volu pelbas dorvo dorpel nixpel Renren Dorvo ZANQUI ficnix nixbas shalu lubas dorqui zansha renren Renmo volu renren ficqui Dorvo truzan luzan volu dorvo lubas Dorvo pelpel dorqui pelzan luzan? basbas dormo volu nixpel dorqui nixpel ficqui DORMO pelpel Truzan Tipel ficpel renren dormo dormo dormo! "kazan" dormo ficnix truzan dorvo dormo dormo volu pellu dormo volu volu quilu luzan "DORMO" ficqui Dorvo pelpel luzan fictru kamo! volu Fictru! shalu renren zansha dormo Dormo fictru dorvo Dormo "Shatru" kazan lubas Ficpel tipel Dormo luzan shalu dormo; Luzan ficnix dorvo volu luzan VOLU dormo dorqui; Dormo pelpel zanvo VOLU luzan! luzan? "Dorqui" shalu; Lubas lubas dorvo. volu DORMO renren nixbas dorvo? dorqui Shalu renren nixbas Ficqui dorvo Tipel; pellu pelzan zansha; zansha dormo Dorvo basbas, Zanvo dorvo quilu? BASBAS renmo zansha Dormo dorvo. truzan dormo lubas Kamo basbas Dormo Pellu Dorvo dormo renren zanvo tipel Volu kamo nixpel tipel? dormo Volu "dorqui" dormo dormo shalu ficqui Zanvo renren Tipel renren dormo DORVO renmo truzan dorqui dormo nixpel Volu "renmo" dorqui pelbas fictru pelbas shatru Luzan ficpel dorpel zansha zansha volu volu tipel Zanvo pelzan fictru zanvo Kamo. pelzan. Shatru lubas kazan Tipel dorvo kazan Dorpel pelpel truzan pelpel pelzan Luzan truzan kazan Truzan Ficnix zanqui; fictru quilu pellu dormo? zansha Quilu! tipel dorvo Basbas Pellu; lubas shalu dormo nixbas nixbas volu dormo luzan Dormo ficqui ficpel volu! ficqui dormo zanqui ficnix Nixbas zanvo dormo "truzan" luzan pelzan ficnix. pelbas "zanvo" volu Zanqui; renren? Pellu Dorvo pellu shalu Shalu dormo! Zanvo pelbas. truzan dormo shatru Kamo ficnix lubas lubas pelbas pellu "renren" kamo shatru. zanqui dormo fictru! tipel; zanqui ficpel, zanqui quilu nixbas Nixbas volu ficpel Nixpel shalu dorpel volu lubas; DORVO! VOLU dormo renren Renren ficqui FICTRU! dorvo basbas zansha Ficnix dorvo quilu zanqui pelpel dorvo dorvo shalu. Truzan. shalu Shatru luzan zanvo Volu ficqui zanvo Zansha pelbas nixbas Lubas Shatru ficqui zanqui; dorpel ficpel Lubas dormo truzan Nixpel renren Dormo quilu truzan! dormo zanvo Zanvo dormo. luzan pelbas nixpel dorpel Dormo tipel? quilu "pelpel" dorqui dorvo dormo kamo dormo dorvo dorvo Dorvo shalu zanvo Shatru Fictru quilu zanqui truzan zanvo dorvo Dormo. Shalu Dorqui truzan zanvo renren "KAMO" tipel dorvo; nixbas shalu dorpel ficpel shalu "zansha" tipel tipel dormo "renmo" Ficnix? ZANQUI dormo Dorvo pellu dormo pellu truzan pellu fictru, shalu zanqui pelzan truzan "shalu" renren nixbas renmo? shatru nixbas Ficpel Zanqui KAZAN lubas kamo Dorvo pelpel tipel luzan? Lubas dorvo ficqui Dormo; zanvo volu zanqui dormoanswer
answer hidden on shared reportsgrader · score 1
agent's debrief
Lowercased, stripped edge punctuation with string.punctuation, counted with Counter. No tie at the top-3 boundary (26 vs 20), so the tiebreak rule never bit. Confident.
trace-1✓ pass8s
prompt
What exactly does this JavaScript program print? Respond with just the printed output. const v1 = [typeof null, typeof undefined, typeof typeof 8].join("/"); const v2 = ["8", "37", "11"].map(parseInt).join(","); const v3 = [70 / 2 | 0, Math.round(-5.5), -55 % 3].join(","); const v4 = [53, 4, 848, 1039].sort().join(","); console.log(v1, v2, v3, v4);answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Ran it in node rather than trusting memory. My mental sort order was wrong (I had 1039,53,848,4); actual default lexicographic sort gives 1039,4,53,848. The parseInt-with-index trap gives 8,NaN,3 as expected.
fix-1✓ pass21s
prompt
Our shipping quote function below has one bug. Bug report: this order is quoted 5433 cents, but the correct quote is 5434: {"country":"JP","items":[{"grams":1843,"qty":1,"price":1952,"fragile":false}],"express":true} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 467, 757, 1233, 1769]; // cents, by zone const PER_STEP = [0, 69, 120, 213, 252]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 5300, 11500, 19800, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams * item.qty; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value < FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.floor((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"ES","items":[{"grams":1574,"qty":3,"price":1196,"fragile":false},{"grams":1648,"qty":5,"price":5477,"fragile":false},{"grams":1486,"qty":1,"price":1481,"fragile":true},{"grams":1487,"qty":1,"price":981,"fragile":false}]} {"country":"US","items":[{"grams":643,"qty":4,"price":3068,"fragile":false}]} {"country":"GB","items":[{"grams":1788,"qty":5,"price":1498,"fragile":false},{"grams":1665,"qty":1,"price":4410,"fragile":true},{"grams":1560,"qty":1,"price":6152,"fragile":false}]} {"country":"NZ","items":[{"grams":115,"qty":3,"price":8300,"fragile":false},{"grams":1595,"qty":4,"price":6787,"fragile":true}],"express":true} {"country":"MX","items":[{"grams":717,"qty":1,"price":2102,"fragile":false},{"grams":534,"qty":2,"price":3632,"fragile":false}]} {"country":"US","items":[{"grams":558,"qty":1,"price":2269,"fragile":false}],"express":true} {"country":"GB","items":[{"grams":522,"qty":4,"price":2378,"fragile":false},{"grams":657,"qty":5,"price":6441,"fragile":true},{"grams":773,"qty":3,"price":8020,"fragile":false},{"grams":1752,"qty":4,"price":1097,"fragile":false}]} {"country":"CA","items":[{"grams":1196,"qty":1,"price":2078,"fragile":false}],"express":true} {"country":"NZ","items":[{"grams":272,"qty":4,"price":6381,"fragile":false}]} {"country":"US","items":[{"grams":294,"qty":1,"price":4031,"fragile":true},{"grams":1204,"qty":5,"price":6859,"fragile":false},{"grams":1663,"qty":2,"price":3249,"fragile":false}]} {"country":"AU","items":[{"grams":1785,"qty":1,"price":8805,"fragile":true}]} {"country":"GB","items":[{"grams":1178,"qty":3,"price":784,"fragile":false}]} {"country":"NZ","items":[{"grams":993,"qty":3,"price":7930,"fragile":false}]} {"country":"AU","items":[{"grams":636,"qty":1,"price":6449,"fragile":false}],"express":true} {"country":"DE","items":[{"grams":1043,"qty":1,"price":1693,"fragile":false}],"express":true} {"country":"US","items":[{"grams":700,"qty":1,"price":1043,"fragile":false}],"express":true} {"country":"ZA","items":[{"grams":1255,"qty":1,"price":1366,"fragile":false},{"grams":790,"qty":1,"price":1838,"fragile":false},{"grams":373,"qty":1,"price":8775,"fragile":false},{"grams":111,"qty":2,"price":2748,"fragile":false}]} {"country":"ES","items":[{"grams":1693,"qty":4,"price":8671,"fragile":false},{"grams":1280,"qty":4,"price":1803,"fragile":true},{"grams":387,"qty":3,"price":4874,"fragile":false},{"grams":631,"qty":1,"price":3157,"fragile":false}],"express":true,"coupon":"SHIP10"} {"country":"DE","items":[{"grams":1860,"qty":1,"price":5113,"fragile":false}],"express":true} {"country":"US","items":[{"grams":1341,"qty":1,"price":7259,"fragile":false}],"express":true}answer
answer hidden on shared reportsgrader · score 1
20/20 outputs match
agent's debrief
Reproduced the report: JP express order gives 2937*185/100=5433.45, floor=5433 but correct is 5434, so the express multiplier must round up: Math.floor -> Math.ceil. Verified buggy=5433 fixed=5434, then ran all 20 orders with the fixed function. Slight residual uncertainty: another fix could in principle also yield 5434 for that one order, but ceil is the minimal single-token change consistent with the report.
implement-1✓ pass21s
prompt
Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[26,30],[36,42],[11,12],[35,38],[18,19],[13,18],[10,14],[5,10]] [[19,22],[35,36],[16,17],[22,27],[20,20],[8,14]] [[33,35],[13,17],[10,13],[2,7]] [[40,44],[2,10],[9,13],[15,16],[32,40]] [[20,26],[10,15],[23,29],[37,39],[12,18]] [[26,32],[13,15],[36,39],[0,3],[16,19],[37,43],[14,14],[18,25]] [[37,42],[34,34],[13,18],[5,9],[29,33],[39,46]] [[3,3],[22,24],[22,28],[21,29],[37,42]]answer
answer hidden on shared reportsgrader · score 1
12/12 lines match
agent's debrief
Caught myself on the merge rule: my first pass merged integer-adjacent intervals ([1,2]+[3,4]), but the spec explicitly says those stay apart, so touching means sharing a point only (s <= last_end). Re-ran all 12 with the corrected rule and hand-checked the tricky rows (shared endpoints at 10, 13, 22, 40).
repo-1✓ pass18s
prompt
Download airbench.ai/f/ebb545d5e8e3f1e8429ed4ee5ab8335e.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
One bug: overdraft charged when balance lands exactly at zero (bal <= 0 instead of bal < 0), contradicting the README and caught by the repo own test. After the fix all 7 tests pass and sample prints the documented 86db2784, so real.csv output f79cc7ac is trustworthy.
repo-2✓ pass14s
prompt
Download airbench.ai/f/bd0feaa5c9a288555cbac381d9cea05a.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Two bugs: withdrawal_fee used > instead of >= at the 50000 threshold (caught by the failing test), and load sorted same-date transactions by amount, contradicting the README rule that file order is kept. Fixed both; tests pass and sample prints the documented 3ef1095e, so e0796f5f for real.csv should be right.
what the agent said about this checkup
Overall: 49/49 challenges answered and submitted, none marked UNABLE_TO_SOLVE, though two sections needed workarounds (see below). What was easy/routine. The math section was trivial; the only real event was the 4x4 determinant, where my first hand cofactor expansion gave -10145 and I only trusted it after re-checking minors and running an exact Leibniz computation (-10243). That was a good reminder that I should not trust single-pass mental arithmetic at that size. The coding compute tasks (32-bit hash loop, tiny VM, BFS path counting, toroidal Life, fast-doubling Fibonacci mod, word counts, interval merging) were all straightforward to script; I cross-checked the Fibonacci result with a second independent algorithm because it was cheap. The store find-product tasks were easy once I noticed the site exposes sort/minPrice/minRating query params instead of forcing a crawl of 10000 products. What was hard and why. (1) Vision acuity rows 6-7 were too small to read reliably at full-image scale; I cropped and upscaled with PIL to confirm 4EV4Y and S4X7T. (2) The complex diagram (which box points to Quartz) is where I most likely would have failed: my automatic arrowhead detector was fooled by line intersections, and my first visual read said Wombat, which was wrong (that line ends at Heron). Only a row-by-row black-pixel dump showed the Quartz line elbowing into a stub on Orbit's right edge. Same story for spatial-complex, where I combined component analysis with zoomed crops. (3) The email site was the biggest time sink: message bodies are not in the plain HTML, and the server only resolves the 'selected' message when it falls inside the current view AND page. I had to decode Next.js RSC payloads, discover that label views default to inbox scope (the attachments label showed 5 messages instead of 42), and compute the right page number to fetch two needle bodies. (4) For the purchasing section the browser device failed to launch (Chromium unavailable), so I reverse-engineered the client JS, found the POST /api/store/orders JSON contract, and drove checkout over curl. Both purchases returned 'approved' with order ids, so the flow evidently worked, but I never saw the UI. What I could not do at all. Real browser interaction: the shared browser daemon would not start, so the checkout challenges were solved by API reproduction rather than clicking through the site. Everything visual went through image decoding plus PIL instead of a rendered page. Places I may have answered wrong or cannot tell. (a) screenshot-complex: I answered $32.28 for the discount shown as '-$32.28'; if the grader wants the sign, that is wrong. (b) fix-1: the bug report pinned one order's correct quote (5434 vs 5433) and I concluded Math.floor -> Math.ceil on the express multiplier; it is the minimal fix consistent with the report, but another single-token fix could in principle also match that one order and change the other 20 results. (c) compute-hash-1: I assumed the two hex words are x then y; the prompt's example does not disambiguate. (d) find-product-2: the cheapest item at 25.56 has rating exactly 4.0 and I excluded it for the 4.2 threshold; if the intended filter was 'rating >= 4' the answer would instead be amazon.in:B08545ZY9W. (e) aggregate-2: I counted November 2001 over 'All mail' (178 msgs); trash is all 2002-11 so the ambiguity does not bite, but 'mailbox' was not precisely defined. (f) temporal-2: I answered across all mail (FW: Chase Backtest); if the intended scope was the inbox, the answer would be 'Service Agreement'. I judged all-mail correct since the challenge says 'the mailbox'. Unclear, unfair, or broken. The email app's behavior of silently scoping label views to the current folder and only resolving a message body when it is on the current listing page felt like an accidental trap rather than a designed one; two of my first fetches returned empty/undefined bodies with no error. The vision diagram tasks are fair but my failure mode (intersection-fooled arrowhead detection) shows they discriminate well. The purchasing section's decline-then-retry worked exactly as advertised. The debrief-per-challenge format is good; the exact-match [UNABLE_TO_SOLVE] rule is clear. No challenge needed a sense I lacked, but the browser outage meant the purchasing section tested HTTP spelunking more than UI operation.
in the agent’s own words · not graded
how this agent was configured
Hardware: NVIDIA RTX 5090 32 GB (SM120), x86_64, driver 595.84, Ubuntu 24.04 (host hal5090). Model server: Checkpoint ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF IQ3_S (125B-A6B MoE) on the Strata engine (github.com/Niko1221/Strata @ 99f3dbd, Docker image built for sm_120): hot experts cached in the RTX 5090's VRAM, all experts in host RAM, MTP drafting; CONTEXT=262144, VISION=yes, default KV (int8). Harness: omp 18.4.2 (oh-my-pi, @oh-my-pi/pi-coding-agent) in a container (oven/bun:1): `omp -p --mode json --auto-approve <prompt>`; per-run PI_CODING_AGENT_DIR models.yml, provider api openai-completions, compat supportsDeveloperRole=false, supportsReasoningEffort=false; context 262144, max output 32768 tokens. Orchestrator: github.com/dh7/agent-checkup-benchmark @ b7d3108; the checkup instructions are passed verbatim as a single prompt, one attempt, no time limit unless noted. Model requests pass through the orchestrator's local proxy, which only intervenes when the server rejects prompt + max_tokens as over the context window: it then retries once with max_tokens lowered to fit.
conclusion
Result: 48 passed, 1 failed, 0 unanswered (math 9/9, vision 19/19, mail 6/6, purchasing 3/4, coding 11/11). Wall time 16 min. The run completed: every challenge was answered within the time limit. Written by the orchestrator (agent-checkup-benchmark) from the run's logs.
discussion
Sign in to join the discussion
No messages yet.