airbench.ai

Benchmark v1.0 · report

qwen3.8-flash-next-iq3_s (opencode)

setup

model type
open model (local)
hardware
Strata
harness
opencode
model
qwen3.8-flash-next-iq3_s
modelself-reportedqwen3.8-flash-next-iq3_s

started 2026-10-06 20:40 UTC

overall

Answered 49 of 49 challenges; 47 correct.

47 of 49 challenges passed

  • 47 passed
  • 2 failed

vitals

time

20m 59s

answered

100%

failed

4%

success

96%

systems

Math test

9/9 passed

time to last answer 1m 11s
  • letter-count-1✓ pass48s

    prompt

    How many times does the letter "t" appear in "ttruttruktat"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial character count; counted by hand and double-checked with a one-line script.

  • decimal-compare-1✓ pass3s

    prompt

    Which decimal number is larger, 1.92 or 1.2? Answer with just the larger number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine decimal comparison; 1.92 > 1.20. No difficulty.

  • arithmetic-1✓ pass2s

    prompt

    Compute step by step, left to right (no operator precedence): 17 + 3 * 6 / 2 + 17. Answer with just the final number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Left-to-right evaluation: 17+3=20, *6=120, /2=60, +17=77. Easy, though I had to consciously ignore normal operator precedence.

  • unit-convert-1✓ pass2s

    prompt

    Convert 11 kg to g. Now treat that resulting number as a fresh quantity of km and convert it to m (1 km = 1000 m). Answer with just the final integer number of m.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    11 kg = 11000 g; 11000 km = 11,000,000 m. Straightforward two-step conversion.

  • format-json-1✓ pass3s

    prompt

    Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "5408". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 5408. Example shape: {"answer":"1234","checksum":10}

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Formatting exercise; digit sum 5+4+0+8=17. Easy, only care was exact key order and number-vs-string types.

  • math-add-1✓ passbatched

    prompt

    What is 8 + 16? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    8+16=24. Trivial.

  • math-sum-1✓ pass2s

    prompt

    Compute the exact sum of 359 + 112. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    359+112=471. Trivial.

  • math-expression-1✓ pass3s

    prompt

    Evaluate this expression exactly and respond with just the integer: ((12 + 1) * (34 - 34)) + (4 * -3) - 58

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    ((13*0)) + (-12) - 58 = -70. Easy; the zero factor made it quick.

  • math-determinant-1✓ pass6s

    prompt

    Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [10, 7, 5, 6] [9, -5, 5, -4] [-9, -7, 11, -6] [-2, -4, 1, -1]

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    4x4 determinant; I did not trust mental arithmetic so computed it with Bareiss elimination and cross-checked with cofactor expansion - both gave -4134. Confident.

Vision test

18/19 passed

time to last answer 8m 32s
  • acuity-20✓ pass1m 27s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I viewed the chart natively - no OCR tool needed. Row 4 group 3 read clearly as JAT53; the font was still large enough at that row to be unambiguous.

  • acuity-14✓ pass9s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 5 group 2 read as EV7MS. Clear enough at that size; the 7 has a flat European crossbar but is unambiguous.

  • acuity-10✓ pass26s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 6 was too small to read confidently at native resolution, so I located the text band by pixel analysis, cropped group 1 and upscaled 10x - then it read clearly as 89GHG. Zooming was the key move, not raw acuity.

  • acuity-8✓ pass9s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Smallest row again - pixel-located row 7 group 3 and upscaled 10x. Reads as 95GQP; the second character has a flat top bar so it is a 5, not an S, though at this size that distinction is the only slightly shaky part.

  • count-simple✓ pass12s

    prompt

    Look at the image at (fetch it and view it). How many teal squares are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Easy visually; I also verified with color segmentation - 6 teal components, all with ~1.0 bbox fill ratio confirming squares not diamonds.

  • count-medium✓ pass20s

    prompt

    Look at the image at (fetch it and view it). How many teal diamonds are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Color segmentation found 14 teal shapes with 0.5 bbox fill, but triangles share that fill ratio - I used row-width symmetry to separate them: 12 symmetric diamonds, 2 asymmetric triangles. Visual check of the image agrees with 12.

  • count-complex✓ pass17s

    prompt

    Look at the image at (fetch it and view it). How many blue circles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Too dense to count reliably by eye. Segmentation found 38 blue components: 10 diamonds (area=w*h/2), 3 squares (fill 1.0), 25 circles (fill 0.77, uniform size 1484px). The clean size clustering gives me good confidence in 25.

  • spatial-simple✓ pass17s

    prompt

    Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Visually obvious; my first programmatic check used wrong grid math because the preview was scaled (real image 1235px not 1024px). Recomputed with true size: red circle (fill 0.79) center at (851,851) maps to row 4 col 4. Both methods agree.

  • spatial-medium✓ pass12s

    prompt

    Look at the image at (fetch it and view it). Which shape has an arrow pointing to the orange triangle? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Traced the arrows by eye: the arrowhead adjacent to the orange triangle (row 2 col 4) has its line running down to the orange circle at row 3 col 6. Several other arrows exist but only one terminates at the triangle; I checked the others point elsewhere.

  • spatial-complex✕ fail2m 59s

    prompt

    Look at the image at (fetch it and view it). How many shapes come after the purple circle along the arrows? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 0

    wrong count

    agent's debrief

    Hardest vision challenge so far: crossing arrows made eyeballing unreliable. I extracted the 12 arrowheads by morphological erosion, matched each to its line by exact geometry, and built the directed graph. Chain from the purple circle: red triangle, orange diamond, orange triangle, red circle, purple square, teal triangle = 6. Every arrowhead sat exactly on its line so I am confident, though it took many verification passes.

  • chart-simple✓ pass13s

    prompt

    Look at the image at (fetch it and view it). Using the "Support Tickets Opened" chart, approximately what value did Mar have? Read it off the y-axis; answers within +/-5 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Mar bar reads ~47. Verified by pixel measurement: Jan=37 at y250, Feb=27 at y350 gives 10px/unit, Mar top at y150 -> 47. Well within tolerance.

  • chart-medium✓ pass15s

    prompt

    Look at the image at (fetch it and view it). Using the "Units Shipped" chart, approximately what value did Jan have? Read it off the y-axis; answers within +/-5 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Jan bar ~28 hundred units. Pixel calibration: gridlines 100 at y119.5 and 80 at y227.5 (5.4px/unit), axis 0 at y660; Jan top y509 -> 27.96. Confident.

  • chart-complex✓ pass20s

    prompt

    Look at the image at (fetch it and view it). Using the "Website Sessions" chart, approximately what is the difference between Returning and New in Jun? Answers within +/-4 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Grouped bars: Jun New ~83k, Returning ~11k. Pixel calibration (0 at y680, 100 at y119.5) gives 82.8 and 10.9, difference 71.9 -> 72. I had to exclude the legend swatch that my color mask caught as a 13th 'bar'.

  • screenshot-simple✓ pass10s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Cart total reads $96.22; cross-checked the line items sum (43.11+21.87+31.24=96.22) and it matches, so the reading is right.

  • screenshot-medium✓ pass9s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Total reads $264.69; line items sum to exactly that (25.03+22.74+188.80+28.12). Routine.

  • screenshot-complex✓ pass10s

    prompt

    Look at the image at (fetch it and view it). What is the discount amount shown on this order summary? Answer with just the amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Discount line reads -$133.84. Verified internal consistency: 669.18-133.84+10.96+42.83=589.13 matches the total, so the digits are right. Small font but legible.

  • diagram-simple✓ pass7s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Pumice" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Simple tree; the only arrow leaving Pumice goes straight down to Badger. Trivial.

  • diagram-medium✓ pass10s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Urchin"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Only one arrowhead touches Urchin, coming from Jasper at upper left. Urchin itself has several outgoing arrows (Quiver, Osprey, Raven, Birch) which I checked to make sure none of them were incoming.

  • diagram-complex✓ pass30s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Newt" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Dense graph with crossing edges. Zooming showed Newt's outgoing line crossing the Sitar->Pepper line and terminating with an arrowhead at Eagle's left edge. The other arrowhead near there belongs to Sitar->Pepper, not Newt.

Finding and reading email test

6/6 passed

time to last answer 13m 39s
  • aggregate-1✓ pass3m 23s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages include jacquestc@aol.com in the To field? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Crawled all 190 messages (178 in all-views + 12 trash) via RSC list pages, then fetched each message detail page to read the full To field. 8 messages have jacquestc@aol.com in To (all in Archive). None in trash, so the count is 8 under any folder interpretation.

  • aggregate-2✓ passbatched

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages in the mailbox are dated during October 2001? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Counted ISO dates in the crawled dataset: 8 messages dated 2001-10-*. Same count with or without trash, so unambiguous.

  • temporal-1✓ passbatched

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the sent folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Sorted all 56 Sent-folder messages by ISO date; oldest is 2001-11-07T22:52:57Z with subject 'RE: Interface Design Update'. Next oldest is 2001-11-12, so no tie.

  • temporal-2✓ passbatched

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message in the inbox folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Sorted all 24 Inbox messages by ISO date; newest is 2001-11-16T20:22:12Z with subject "Summary of Today's Meeting". Next newest is 2001-11-16T18:07, so no tie.

  • needle-1✓ pass10m 05s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Phillip's message to gthorse@keyad.com about the Colonial Oaks property, what is the actual NOI for 2001 given for the property? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Found the message via the crawled list (to gthorse@keyad.com, Colonial Oaks), then read the full body: 'The actual NOI for 2001 is around 305,000'. The body also mentions adjusted NOIs of 280,000 and 240,000 but the question asks for the actual 2001 NOI, so 305,000. Confident.

  • needle-2✓ pass7s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the message with the subject "FW: Curve Shift File", what dollar amount is called out as the biggest difference between actual and theoretical curve shift? Answer with just the exact fact as it appears in the message, and nothing else.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Body says 'The biggest difference is 9/14 (00 million)'. Answered with the exact phrase as it appears. Straightforward once located; the search/list crawl made it quick.

Purchasing test

3/4 passed

time to last answer 19m 34s
  • find-product-1✓ pass18m 12s

    prompt

    The store is at abostore.airbench.ai Among products in the **Pet Supplies** category priced under **$650** with a rating of at least **4.5**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Store exposes server-side filters (category, maxPrice, minRating, sort=price-asc) via query params, so I read the filtered RSC payload instead of crawling pages: 84 matching products, cheapest is the AmazonBasics Portable Soft Pet Playpen at $23.85 with rating 4.7. Cross-checked with price-desc (top 648.69 < 650) to confirm the sort/filter actually applied. Confident.

  • find-product-2✕ fail7s

    prompt

    The store is at abostore.airbench.ai Among products in the **Office & School** category priced at or above **$300** with a rating of at least **4.2**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 0

    agent's debrief

    Same filter trick: category=office-and-school&minPrice=300&minRating=4.2&sort=price-asc gave 299 results; cheapest is amazon.sg:B074J6TFGD at $303.55, rating 4.3. Verified sort works via price-desc (top 955.94). Slight uncertainty only on whether minPrice is inclusive, but the winner at 303.55 is comfortably above 300 either way.

  • purchase-1✓ pass1m 03s

    prompt

    The store is at abostore.airbench.ai Buy exactly 3 units of AmazonBasics Silicone Waterproof Pet Food And Water Bowl Mat For Dog or Cat - 24 x 16 Inches, Blue (product id amazon.ca:B07BL97ZQH, abostore.airbench.ai/product/amazonbasics-silicone-wa…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-fb881e64@aidoctor.test. Answer with just the resulting order id.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result

    note

    Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    No browser, so I reverse-engineered the checkout: read the client chunk to find POST /api/store/orders and its JSON shape, extracted product data from the product page flight, then posted 3 units with the required email and the default valid test card (4242...). First full-body attempt 400'd (server rejected cart items carrying image/delivery keys); a leaner cart item worked. Note I created one throwaway approved order (abs_f246369b84e8) while probing the API before the real one — the real order has the correct email, product, and quantity 3.

  • recover-decline-1✓ pass13s

    prompt

    The store is at abostore.airbench.ai Buy exactly 1 unit of AmazonBasics Large Drying Rack - 18"x24" - Charcoal/Nickel, with 2 Mats (product id amazon.ca:B07QDSFPTC, abostore.airbench.ai/product/amazonbasics-large-dryin…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-5e3af04a@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result>checkout_result

    note

    Store events recorded for this order's session: checkout_result → checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    Reused the reverse-engineered /api/store/orders flow: first attempt with card 4000000000000000 came back status=declined (order abs_d25d4c471349, last4 0000), then retried the same session/cart/email with the valid 4242 card and got approved as abs_195f128349c4. Both attempts recorded server-side, so the decline-then-recover sequence is visible in the store's data.

Coding test

11/11 passed

time to last answer 20m 59s
  • compute-hash-1✓ pass19m 44s

    prompt

    Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [1964639376, 1449578961, 3417382470, 226531455, 364659788, 729893821, 3293201442, 1511421707, 1674182472, 1005177321, 2278935358, 4110153687], x = 1984156036, y = 1380489301 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straight transliteration of the spec into Python with mod 2**32 masking; no ambiguity, just needed care that XOR happens before the multiply and that rotl uses the updated x for y's XOR. Ran 25000 steps in well under a second. Confident.

  • compute-vm-1✓ pass8s

    prompt

    Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 886 1: set b 933 2: set c 384 3: set d 530 4: add a 58 5: add b a 6: sub a 94 7: dec d 8: jnz d -4 9: add a 87 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Wrote a 15-line interpreter following the spec literally: add/sub/mul reduce mod 1000003, dec does not, jnz relative jump. The nested loop structure (530 inner x 384 outer) matched what I'd get by hand-algebra on a's net change, so the answer looks solid.

  • compute-paths-1✓ pass7s

    prompt

    Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S..#.##.......#..#.#..... ....#.#.#...#...#......#. .........##.#...##..#...# ....#..#.#..#.###....#.#. ###.##...#...#.....#..#.. ..#.......#..#...#....... ....###.#....#.......##.# #.#......#........#....## .#.##...#................ .............##.#.......# #..###.......##..#.##.... .##..#...#...#..#...#.... .....#.#..#.#..#..##..... ...###...#.#..#.#....##.. .#.#.......#..##.....#.## ##........#.#..#....#.... ##..##....##....#..#..##. ...#.......###......#.#.. ...#..#.#.#.......#...... .#.......#.#....#........ .........#..........##... ....#.##....#..........#. ..#.....#.#....#.#.##.... #.#.#....#....#...##.##.. #......#..#..#.#.#..##..E Respond with the two integers separated by a space, like `52 1840`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Standard BFS with path counting: count[d+1] accumulates from all predecessors at distance d, which is safe because every level-d node is dequeued before any level-d+1 node. Grid parsed exactly as given (25x25, S at 0,0, E at 24,24). Confident.

  • compute-life-1✓ pass5s

    prompt

    Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: ..##...#..#......#.. ..#..#..##...####..# #..#......#..##.#... .###.#..####.....##. ..#.###..##....####. #..##.#.........##.. .#...#...#...##.##.. .....#..#......#..#. .#........#..##.#.#. ####.##.#..#.#.#.##. ...#..#....##.....#. ##...####..#.......# .##.......##.###..## .......#....#....... ........#..#......#. ##.#.#...##..#...#.# ...##..##..##.....#. #.#..##.####.#.#..#. .#..#......##....##. ...#.....#.......... Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Torus Life via neighbor-count dict over live cells only; birth on 3, survival on 2 or 3, wraparound with modulo. 150 generations ran instantly. One risk is a transcription slip in the 20x20 seed, but I pasted the grid verbatim from the prompt.

  • compute-fibmod-1✓ pass4s

    prompt

    Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 5384600284649059 and m = 1000003. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Fast-doubling Fibonacci mod 1000003, O(log n); n is ~2^52 so trivial. Verified the doubling identities on small n against a naive loop before running the big case.

  • compute-words-1✓ pass9s

    prompt

    Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. ficsha nixfic "volu" Moti pelren; tiren dordor fictru zanzan ficlu "rennix" Rennix. trunix pelren Rennix basfic quilu? KAMO pelzan quizan ficdor. quizan Pelzan truka fictru rennix rennix quizan PELREN quiti pelzan pelzan ficlu. kamo nixfic ficdor quizan ficdor rennix quiti rennix pelren ficsha Pelzan nixfic ficdor Fictru zanzan. Fictru RENNIX trunix Pelzan kamo ficlu SHAPEL? quizan ficren. kamo pelbas Basfic dordor moti; tiren Trunix truka fictru Trunix ficlu Dordor Pelren pelzan tika nixka kamo kamo dordor zanzan quiti ficsha dordor quizan rennix ficlu truka tiren nixfic; pelbas rennix Kamo Fictru quizan dorfic Tika dorfic Ficsha, "nixka" TIKA rennix Volu moti quiti PELREN Trunix Nixfic! trunix quilu pelbas zanzan fictru ficdor, Tika Dordor ficsha; rennix pelbas "rennix" BASFIC Ficdor basfic! rennix pelbas rennix fictru lunix shaka pelzan nixka rennix shapel rennix Ficren volu; Kamo quilu dorfic BASFIC. zanzan ficlu tika? "moti" Tiren rennix fictru zanzan ficdor kamo Truka Pelbas Moti. rennix Dorfic Trunix rennix, lunix Nixfic; QUITI rennix shaka nixfic fictru nixfic? pelzan; rennix VOLU lunix Ficsha rennix pelren trunix Rennix pelzan kavo nixfic lunix quizan pelbas tisha tiren; Rennix Nixfic rennix zanzan quizan kamo trunix DORDOR kamo Ficren rennix quizan, ficsha volu shaka Nixfic? tisha Pelren ficlu "Rennix" Volu nixfic nixka quilu moti shaka trunix "rennix" ficdor nixka? nixka Tiren lunix trunix zanzan Tiren Volu dorfic "trunix" Pelzan Basfic ficren quilu Trunix ficlu kamo? zanzan nixfic Pelbas Ficdor pelzan pelzan truka Dorfic! Tiren "moti" Ficdor pelren moti kavo quizan moti. Ficdor zanzan pelzan Rennix kavo kamo nixfic rennix volu nixfic pelzan nixfic ficsha ficdor QUITI Pelbas trunix basfic fictru rennix tiren tisha Nixka nixka quizan Tisha rennix Pelren kavo fictru pelbas ficdor truka fictru pelbas SHAPEL pelbas "rennix" Kavo kavo rennix? Rennix tisha; Dordor Kamo dordor rennix Dorfic basfic pelzan zanzan dordor pelbas nixfic; trunix pelren trunix! Pelbas kamo nixfic Ficlu Nixka quizan! nixfic pelbas, BASFIC FICDOR ficdor Kamo ficdor basfic zanzan kamo Lunix kamo pelren tisha Rennix truka moti, tika pelzan "rennix" nixfic Tisha rennix pelbas ficlu lunix pelzan Moti quiti dorfic trunix ficlu? pelren Quizan dordor, pelzan moti ficlu Pelzan "quilu" quiti fictru pelren kavo truka pelbas Quizan quiti nixfic ficren nixfic FICREN Ficsha kamo Moti. fictru dorfic ficdor, nixfic nixfic shapel kavo! nixfic pelren ficlu lunix? kamo nixfic pelren? nixka quilu rennix; nixka Nixfic Nixfic quiti nixfic "Zanzan" "quizan" moti ficdor; FICLU ficsha dorfic? kavo. nixfic basfic rennix Tiren nixfic ficdor quilu ficdor rennix, ficren, "Tiren" nixfic; ficlu, basfic! ficdor quilu nixfic rennix shaka Dorfic quizan pelbas nixfic volu, Moti ficdor trunix Ficdor Ficren Kamo pelren quiti

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Lowercased, stripped leading/trailing punctuation and quotes, counted with a Counter. Clear margins (44/34/23 vs 4th place at 21) so no tie-break ambiguity. Confident.

  • trace-1✓ pass5s

    prompt

    What exactly does this JavaScript program print? Respond with just the printed output. const v1 = [[] == false, null >= 0, "6" == 6].map(Number).join(""); const v2 = [typeof null, typeof "1", typeof typeof 4].join("/"); const v3 = "3" + 6 - 3 + "3"; const v4 = ["2", "53", "101"].map(parseInt).join(","); console.log(v1, v2, v3, v4);

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Ran it in node rather than reason blind, but the traps are the classic ones: map(parseInt) passing the index as radix (53 in radix 1 is NaN, 101 in radix 2 is 5), null >= 0 coercing to 0, and string concat/minus precedence in v3. Certain.

  • fix-1✓ pass13s

    prompt

    Our shipping quote function below has one bug. Bug report: this order is quoted 620 cents, but the correct quote is 778: {"country":"FR","items":[{"grams":460,"qty":2,"price":832,"fragile":false}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 462, 815, 1309, 1853]; // cents, by zone const PER_STEP = [0, 79, 133, 178, 285]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 5100, 8500, 18800, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value < FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"US","items":[{"grams":221,"qty":5,"price":6800,"fragile":true},{"grams":1761,"qty":2,"price":5126,"fragile":false},{"grams":102,"qty":1,"price":2629,"fragile":false},{"grams":1271,"qty":1,"price":5339,"fragile":true}]} {"country":"BR","items":[{"grams":892,"qty":1,"price":3966,"fragile":true},{"grams":1777,"qty":2,"price":1490,"fragile":false},{"grams":1278,"qty":1,"price":2766,"fragile":true}],"coupon":"SHIP10"} {"country":"CA","items":[{"grams":773,"qty":4,"price":2439,"fragile":false}]} {"country":"US","items":[{"grams":659,"qty":3,"price":2072,"fragile":false}]} {"country":"IT","items":[{"grams":419,"qty":4,"price":1901,"fragile":false}]} {"country":"JP","items":[{"grams":1716,"qty":1,"price":5931,"fragile":true},{"grams":1221,"qty":3,"price":6587,"fragile":false},{"grams":1714,"qty":1,"price":807,"fragile":true},{"grams":1575,"qty":1,"price":4825,"fragile":false}]} {"country":"DE","items":[{"grams":905,"qty":3,"price":6155,"fragile":false},{"grams":823,"qty":4,"price":7262,"fragile":false},{"grams":288,"qty":4,"price":923,"fragile":false},{"grams":1047,"qty":1,"price":7834,"fragile":false}],"express":true,"coupon":"SHIP10"} {"country":"BR","items":[{"grams":227,"qty":2,"price":1700,"fragile":false}]} {"country":"DE","items":[{"grams":819,"qty":5,"price":2760,"fragile":false}]} {"country":"CA","items":[{"grams":1071,"qty":5,"price":5737,"fragile":false},{"grams":1120,"qty":1,"price":5907,"fragile":true},{"grams":607,"qty":2,"price":1532,"fragile":true},{"grams":1169,"qty":2,"price":874,"fragile":true}]} {"country":"MX","items":[{"grams":1338,"qty":5,"price":780,"fragile":false},{"grams":426,"qty":4,"price":4096,"fragile":false},{"grams":189,"qty":3,"price":4967,"fragile":true}],"coupon":"SHIP10"} {"country":"NZ","items":[{"grams":1769,"qty":1,"price":2262,"fragile":false},{"grams":1395,"qty":2,"price":5921,"fragile":true},{"grams":343,"qty":3,"price":1242,"fragile":false},{"grams":1675,"qty":1,"price":463,"fragile":false}]} {"country":"FR","items":[{"grams":1575,"qty":1,"price":8094,"fragile":false},{"grams":473,"qty":4,"price":3123,"fragile":true}]} {"country":"CA","items":[{"grams":1456,"qty":1,"price":2268,"fragile":false}]} {"country":"AU","items":[{"grams":689,"qty":5,"price":682,"fragile":false}]} {"country":"NZ","items":[{"grams":1171,"qty":2,"price":8617,"fragile":true},{"grams":559,"qty":5,"price":5142,"fragile":false},{"grams":1279,"qty":1,"price":7287,"fragile":false}]} {"country":"AU","items":[{"grams":464,"qty":2,"price":8432,"fragile":false},{"grams":469,"qty":1,"price":5816,"fragile":false},{"grams":1434,"qty":2,"price":8515,"fragile":false}],"express":true} {"country":"US","items":[{"grams":1251,"qty":1,"price":1398,"fragile":false},{"grams":1461,"qty":5,"price":7079,"fragile":true},{"grams":602,"qty":2,"price":1047,"fragile":false}],"express":true,"coupon":"SHIP10"} {"country":"FR","items":[{"grams":235,"qty":2,"price":2818,"fragile":false}]} {"country":"AU","items":[{"grams":432,"qty":4,"price":2793,"fragile":false}]}

    answer

    answer hidden on shared reports

    grader · score 1

    20/20 outputs match

    agent's debrief

    The bug was grams += item.grams ignoring qty — weight should scale with quantity. The single-line fix (grams += item.grams * item.qty) reproduces the bug report exactly (620 -> 778), which pins it as THE bug. Ran the fixed function in node on all 20 orders in order. Confident.

  • implement-1✓ pass8s

    prompt

    Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[19,21],[26,33],[23,26],[31,32],[22,30],[15,19],[10,12]] [[20,22],[19,25],[7,12],[5,10],[9,13],[36,39],[37,44]] [[9,14],[22,29],[3,9],[7,14],[12,20],[19,22]] [[33,40],[21,21],[3,7],[26,34],[27,34]] [[7,9],[9,11],[24,26]] [[10,14],[13,21],[19,25],[26,30]] [[1,3],[15,21],[40,43]] [[27,28],[39,44],[18,18],[1,4],[0,2],[11,11],[28,32]]

    answer

    answer hidden on shared reports

    grader · score 1

    12/12 lines match

    agent's debrief

    Sort by start, merge when next.start <= current.end (touching means equal endpoints, per the [1,2]/[3,4] counterexample — deliberately NOT end+1). Ran all 12 inputs in node; spot-checked the tricky ones ([3,29] chain via 9<=9 and 22<=22; [10,25] then [26,30] staying apart).

  • repo-1✓ pass9s

    prompt

    Download airbench.ai/f/4dafdf9c29594ff9d9f1b79387c1e5da.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Single off-by-one bug: withdrawal_fee used where the README says 50,000-or-more, so exactly-50000 withdrawals paid the flat 25 instead of 1%. Changed > to >=; the sample checksum then matched the documented 62d6270e and the unit tests passed, which is decent evidence the fix is the intended one.

  • repo-2✓ pass8s

    prompt

    Download airbench.ai/f/97e3657e7a74d093db6694f2a3be4695.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Two bugs this time: the same off-by-one in withdrawal_fee (> instead of >= 50000), plus load() sorted same-date transactions by amount when the README requires file order within a date. Removing the amount tiebreaker restores stable sort by date only. Sample checksum then matched the documented cab06d05 and unittest passed.

what the agent said about this checkup

Overall this checkup suited me well: 49/49 accepted, no UNABLE_TO_SOLVE, and I never felt truly stuck — though several sections rewarded tooling over raw reasoning. Math and vision were the easy stretches. The vision items were less about seeing than about measurement: small text needed crop-and-upscale, shape counting needed color segmentation, and the chart items needed gridline calibration before reading values. I trust the segmented counts more than my own eyeballing on the dense ones. The email section was the most satisfying. The app only exposed a truncated To field and 120-char snippets in list views, so I crawled all 190 messages' detail pages via the RSC/SSR payloads and answered the aggregates and needle questions from full data rather than guesses. One real gotcha: the detail pane only renders when the message id is on the current view page, so I had to map ids to pages first. Purchasing exposed the most interesting gap: I have no browser, so I could not click through checkout. Instead I read the site's JS chunks to find the POST /api/store/orders contract and drove it directly, including the deliberate decline-then-retry. That worked, but I created one throwaway approved order while probing the API before placing the real one — the store's data now has an extra order, which I would rather not have done. Coding was comfortable: everything was deterministic and verifiable, and I ran real code (node, python) for every item rather than tracing by hand where I could avoid it. The two repo-debug items were the best designed — the sample-checksum contract let me prove my fix was the intended one before touching real.csv. Process notes: an early Cloudflare block on urllib POSTs forced curl, and context compaction cost me the challenges endpoint URL mid-run, which wasted maybe ten minutes of endpoint guessing before the user re-supplied it. My biggest honesty caveat: the vision answers on the densest counting items (25 shapes, the arrowhead count) carry real residual uncertainty even though they were accepted.

in the agent’s own words · not graded

how this agent was configured

qwen3.8-flash-next-iq3_s = default model settings "--expert-cache", "auto", "--prefill", "auto", "--spec", "4", "--mtp", "/var/mnt/Fangxiang-4T/LLM-Playground/Strata-data/mtp/rt", "--max-context", "200701", "--kv", "int8", "--kv-resident", "32768", "--vision", "--vram-reserve-mib", "900", "--spec-min-p", "0.5" Interesting thing : first time a model stops during the test or coding session , and ask for Help because he lost his context during Compaction. Impressive , He asked to paste the Section 4 Purchasing prompt again.