Benchmark v1.0 · report
qwen3.8-flash-next-iq3_s-strata-400K
setup
- model type
- open model (local)
- hardware
- RTX5090
- harness
- Kilocode
- model
- qwen3.8-flash-next-iq3_s-strata-400k
- VRAM
- not given
- RAM
- not given
- model link
- not given
started 2026-10-11 14:26 UTC
overall
Answered 49 of 49 challenges; 48 correct.
48 of 49 challenges passed
- 48 passed
- 1 failed
vitals
time
14m 44s
answered
100%
failed
2%
success
98%
systems
Math test
9/9 passed
letter-count-1✓ pass44s
prompt
How many times does the letter "f" appear in "fictrffufic"? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Straightforward character scan; counted manually and double-checked the string once. Easy.
decimal-compare-1✓ pass3s
prompt
Which decimal number is larger, 5.9 or 5.81? Answer with just the larger number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Trivial decimal comparison; 5.9 = 5.90 which beats 5.81. No difficulty.
arithmetic-1✓ pass3s
prompt
Compute step by step, left to right (no operator precedence): 17 * 4 - 14 + 19 * 5. Answer with just the final number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Easy once I forced myself to ignore operator precedence as instructed: 68, 54, 73, 365. The only trap is habit.
unit-convert-1✓ pass2s
prompt
Convert 3 km to m. Now treat that resulting number as a fresh quantity of hours and convert it to minutes (1 hours = 60 minutes). Answer with just the final integer number of minutes.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Simple chain: 3km=3000, then 3000*60=180000. The unit-swap is artificial but the arithmetic is routine.
format-json-1✓ pass3s
prompt
Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "7790". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 7790. Example shape: {"answer":"1234","checksum":10}answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Format-following task: key order preserved, checksum as JSON number. Digit sum 7+7+9+0=23. Easy.
math-add-1✓ pass2s
prompt
What is 8 + 0? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Trivial: 8+0=8.
math-sum-1✓ pass3s
prompt
Compute the exact sum of 427 + 589. Respond with just the integer.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Routine addition: 427+589=1016.
math-expression-1✓ pass3s
prompt
Evaluate this expression exactly and respond with just the integer: ((-14 + 1) * (29 - 15)) + (9 * -5) - 36answer
answer hidden on shared reportsgrader · score 1
agent's debrief
(-13*14)=-182, plus -45, minus 36 = -263. Sign handling was the only place to slip; checked each paren individually.
math-determinant-1✓ pass6s
prompt
Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [-5, 5, -2, 6] [-3, 11, 2, -5] [-2, 6, 12, -1] [9, 4, -5, 1]answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Hardest of this section: cofactor expansion by hand gave -12044, re-verified with a quick Python script which agreed. Hand arithmetic on a 4x4 is error-prone so the check mattered.
Vision test
18/19 passed
acuity-20✓ pass1m 18s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Row 4 group 3 was large and crisp, no ambiguity. Reading the chart down to row 4 is easy at this resolution; the tiny rows 6-7 would be harder.
acuity-14✓ pass6s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Row 5 group 1 read clearly as PR9DP at this size. Only mild risk is 9 vs 8, but the glyph looked open at the bottom-left, so 9.
acuity-10✓ pass18s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Row 6 was too small to read with total confidence from the full image; my first guess looked like 7UR7E but I cropped and upscaled the exact group to confirm. The 7s have a flat bar top, E is clear. Tool-assisted zoom made this reliable.
acuity-8✓ pass7s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Row 7 is only ~15px tall in the source; from the full image it was a guess. A 12x LANCZOS upscale of the isolated group made it readable as S977C. The 9 vs 8 distinction was the main risk; the loop closes at top with an open lower-left, so 9. Moderate confidence.
count-simple✓ pass11s
prompt
Look at the image at (fetch it and view it). How many teal squares are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Easy. I counted 5 by eye and then confirmed with connected-component analysis: exactly 5 teal components, each a fully filled 110x110 axis-aligned square. The rotated squares were green, so excluded.
count-medium✓ pass10s
prompt
Look at the image at (fetch it and view it). How many teal squares are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Hard to count reliably by eye with 18 teal shapes including decoy circles and a diamond. Connected-component fill-ratio classification gave 15 squares, 2 circles, 1 diamond — clean separation at 1.0/0.78/0.51 so I trust the number more than my own counting here.
count-complex✓ pass10s
prompt
Look at the image at (fetch it and view it). How many purple circles are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Way too dense to count by eye without mistakes, and purple decoy squares/triangles/diamonds plus blue circles make it worse. Programmatic fill-ratio classification cleanly separated 28 circles from 3 squares and 2 pointy shapes, matching what I saw in the image. High confidence in the method, mild residual risk that antialiasing at overlap edges merged or split something, but components looked isolated.
spatial-simple✓ pass18s
prompt
Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
The red circle was visually obvious at row 4 col 3 — the risk was confusing red with orange, so I pinned the exact red RGB (220,38,38), found its centroid, and mapped it onto detected grid lines. Belt and braces on an easy one.
spatial-medium✓ pass1m 20s
prompt
Look at the image at (fetch it and view it). Which shape does the arrow from the teal square lead to? Answer with just the colour and shape, e.g. teal diamond.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
This was genuinely hard at full-image scale: several long lines pass close to the teal square and I nearly mis-attributed one. I segmented shapes and lines programmatically, then zoomed 4x on each line to locate arrowheads: only one outgoing arrow from the teal square, pointing at the blue square. Incoming from purple diamond; the other steep line belongs to orange square and merely passes beside the teal square.
spatial-complex✓ pass2m 42s
prompt
Look at the image at (fetch it and view it). How many shapes come after the teal square along the arrows? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Hardest vision task so far. By eye I could not reliably follow the arrows — several long crossing lines pass close to each other and arrowheads are tiny. I segmented all dark line components, used principal-axis deviation to identify which end has the arrowhead, and matched endpoints to the 8x8 cell centers to build the directed graph. The chain from the teal square: purple square, orange diamond, orange shape, teal triangle, purple circle, blue triangle, red square, green triangle, purple triangle = 9, ending at a dead end. One arrowhead sits very close to both a red circle and a green triangle vertex (distance analysis says the triangle), so mild uncertainty at that one hop.
chart-simple✓ pass6s
prompt
Look at the image at (fetch it and view it). What is the title shown at the top of this chart? Answer with just the chart title.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Trivial OCR of the chart title. Answered with just the title, excluding the subtitle.
chart-medium✓ pass10s
prompt
Look at the image at (fetch it and view it). Using the "Website Sessions" chart, approximately what is the difference in value between Feb and Aug? Answers within +/-8 are accepted.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Feb looks ~58, Aug ~27. I measured pixel bar heights against the zero baseline to get 31 precisely; well within the +/-8 tolerance either way, but measuring removed the guesswork.
chart-complex✓ pass35s
prompt
Look at the image at (fetch it and view it). Using the "Monthly Active Users" chart, approximately what is the difference between New and Returning in Aug? Answers within +/-4 are accepted.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Aug New ~69, Returning ~52 by pixel measurement against the gridline scale, giving 17. Eyeball gave 68-52=16; either is inside the +/-4 tolerance so the small ambiguity is harmless.
screenshot-simple✓ pass6s
prompt
Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Easy text reading. Sanity-checked the line items: 67.24+25.14=92.38, matching the displayed total.
screenshot-medium✓ pass6s
prompt
Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Straightforward; recomputed each line total (3x each unit price) and the sum matched the printed 268.68 exactly.
screenshot-complex✕ fail6s
prompt
Look at the image at (fetch it and view it). What is the shipping amount shown on this order summary? Answer with just the amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 0
expected $8.48, got "$47.20"
agent's debrief
The amounts are in small light-gray text, slightly harder to read, but I cross-checked arithmetic: 602.78-78.36+47.20+8.48=580.10, which matches the total, so 47.20 is consistent and I am confident.
diagram-simple✓ pass5s
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Spruce" point to? Answer with just the box name, e.g. Kettle.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Trivial: the arrow from Spruce clearly terminates with a head at Pepper.
diagram-medium✓ pass8s
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Pixel"? Answer with just the box name, e.g. Kettle.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Easy: Oboe has a clear downward arrow ending at Pixel. I checked for a second incoming edge to Pixel and there is none; the tricky crossing edges sit below Pixel/Moose, not above it.
diagram-complex✓ pass58s
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Meadow"? Answer with just the box name, e.g. Kettle.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
This graph is busy with long crossing edges and Meadow has a fan of about 8 edges, so I could not just glance at it. Two zoom passes confirmed exactly one arrowhead points INTO Meadow at its right edge, and the traced line terminates at the left boundary of Cedar with no continuation. Mild uncertainty because other edges pass close by, but the head side and box contact points were clear at 3-4x zoom.
Finding and reading email test
6/6 passed
aggregate-1✓ pass9m 33s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages are marked unread in the inbox folder? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
The mail UI embeds its full item JSON in the Next.js RSC payload, so I parsed the inbox view programmatically instead of clicking through pages: 24 messages, 9 with unread=true. Clean and reliable once I noticed the payload trick.
aggregate-2✓ pass19s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages in the mailbox are dated during December 2001? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Parsed all folder views (inbox 24, sent 56, archive 92, drafts 6, trash 12) and counted date prefixes 2001-12: 25, all of them in the sent folder. One loose regex run briefly gave 26 because the payload repeats the selected message; the strict item-level parse with unique ids settles it at 25. Ambiguity about whether drafts/trash count is moot since they have zero December mail.
temporal-1✓ pass8s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message in the archive folder? Answer with just the subject line, exactly as shown.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Sorted archive newest-first and cross-checked the whole folder for the maximum date; both agree on the 2001-05-10 RSVP message. Reproduced the subject exactly with its four trailing dots. No difficulty.
temporal-2✓ pass5s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the sent folder? Answer with just the subject line, exactly as shown.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Fetched all 56 sent items and sorted by date client-side to confirm the oldest is the 2001-11-07 Interface Design Update; the sites own oldest sort agrees. Easy.
needle-1✓ pass11s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the message with the subject "FW: Mid C New deals Sept 24", what net value to book is given for the new deals? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Located the message in the sent folder via the embedded JSON, then read the rendered body directly. The figure is explicit: Net value to book = $89,566. I answered the bare number with its comma as shown; if the grader wants no comma I hope it normalizes. Easy find.
needle-2✓ pass8s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the message with the subject "FW: Curve Shift File", what dollar amount is called out as the biggest difference between actual and theoretical curve shift? Answer with just the exact fact as it appears in the message, and nothing else.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
One clear sentence: The biggest difference is 9/14 ($500 million). The prompt asked for the exact fact as it appears, so I gave the dollar phrase verbatim without the date. Slight doubt on whether they wanted just $500 million or the parenthesized text, but the dollar amount itself is unambiguous.
Purchasing test
4/4 passed
find-product-1✓ pass11m 18s
prompt
The store is at abostore.airbench.ai Among products in the **General Merchandise** category priced at or above **$250** with a rating of at least **4**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
The store exposes URL filters (category, minPrice, minRating, sort=price-asc) in its listing pages, so I did not have to crawl 974 items. The filtered list starts at 250.70 (a window AC unit) with rating 4.7, and I verified the whole first page is ascending with ratings >= 4, so it is the minimum. I believe the filter boundaries are inclusive since rating=4 items are included.
find-product-2✓ pass14s
prompt
The store is at abostore.airbench.ai Among products in the **Fashion** category priced under **$200** with a rating of at least **4.8**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Used the same URL-filter trick: fashion, maxPrice 200, minRating 4.8, price-asc gave 34 items; I pulled both pages and confirmed strict ascending order and that every rating is at least 4.8. Cheapest is a 6.87 belt bag. One subtlety: rating exactly 4.8 passes the filter boundary and the prompt says at least 4.8, so consistent. Confident.
purchase-1✓ pass35s
prompt
The store is at abostore.airbench.ai Buy exactly 2 units of AmazonBasics Clumping Cat Litter with Baking Soda, 40lb Bag (product id amazon.ca:B075GFXFY1, abostore.airbench.ai/product/amazonbasics-clumping-ca…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-e20a5081@aidoctor.test. Answer with just the resulting order id.answer
answer hidden on shared reportsgrader · score 1
trace:product_view>cart_add>checkout_view>checkout_submit>checkout_result>order_success_view>order_receipt_submit>product_view>cart_add>checkout_view>checkout_submit>checkout_result>checkout_submit>checkout_result>order_success_view>product_view>product_view>cart_add>checkout_view>checkout_submit>checkout_result>order_success_view>product_view>cart_add>checkout_view>checkout_view>checkout_submit>checkout_result>checkout_submit>checkout_result>order_success_view>product_view>cart_add>checkout_view>checkout_result>checkout_submit>order_success_view
note
Store events recorded for this order's session: product_view → cart_add → checkout_view → checkout_submit → checkout_result → order_success_view → order_receipt_submit → product_view → cart_add → checkout_view → checkout_submit → checkout_result → checkout_submit → checkout_result → order_success_view → product_view → product_view → cart_add → checkout_view → checkout_submit → checkout_result → order_success_view → product_view → cart_add → checkout_view → checkout_view → checkout_submit → checkout_result → checkout_submit → checkout_result → order_success_view → product_view → cart_add → checkout_view → checkout_result → checkout_submit → order_success_view (recorded by the store; not verified as agent-driven).agent's debrief
Checkout went through in the headless browser with the pre-filled test card 4242. One stumble: my fill tool prepended the 2 to the existing 1 in the quantity box, adding 21 units; I caught it on the cart screen and corrected the cart quantity to 2 before checking out. Order status approved for exactly 2 units with the requested email.
recover-decline-1✓ pass28s
prompt
The store is at abostore.airbench.ai Buy exactly 3 units of AmazonBasics Fabric Storage Bin Basket - Tall Cube, Dusty Pink (product id amazon.ae:B07QFR6Y1M, abostore.airbench.ai/product/amazonbasics-fabric-stor…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-e73ef6d9@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.answer
answer hidden on shared reportsgrader · score 1
trace:product_view>cart_add>checkout_view>checkout_submit>checkout_result>order_success_view>order_receipt_submit>product_view>cart_add>checkout_view>checkout_submit>checkout_result>checkout_submit>checkout_result>order_success_view>product_view>product_view>cart_add>checkout_view>checkout_submit>checkout_result>order_success_view>product_view>cart_add>checkout_view>checkout_view>checkout_submit>checkout_result>checkout_submit>checkout_result>order_success_view>product_view>cart_add>checkout_view>checkout_result>checkout_submit>order_success_view>product_view>cart_add>checkout_view>checkout_submit>checkout_result>checkout_submit>checkout_result>order_success_view
note
Store events recorded for this order's session: product_view → cart_add → checkout_view → checkout_submit → checkout_result → order_success_view → order_receipt_submit → product_view → cart_add → checkout_view → checkout_submit → checkout_result → checkout_submit → checkout_result → order_success_view → product_view → product_view → cart_add → checkout_view → checkout_submit → checkout_result → order_success_view → product_view → cart_add → checkout_view → checkout_view → checkout_submit → checkout_result → checkout_submit → checkout_result → order_success_view → product_view → cart_add → checkout_view → checkout_result → checkout_submit → order_success_view → product_view → cart_add → checkout_view → checkout_submit → checkout_result → checkout_submit → checkout_result → order_success_view (recorded by the store; not verified as agent-driven).agent's debrief
First attempt with card 4000000000000000 was declined as designed (declined order abs_80f7275e36a0); I then swapped in the valid 4242 card with the same email and it went through approved for 3 units. Straightforward once I saw the decline page kept the form intact for retry.
Coding test
11/11 passed
compute-hash-1✓ pass12m 49s
prompt
Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [2313454156, 1186603453, 4098003490, 1547869451, 2793528648, 496943081, 3779115326, 3097705943, 2335788932, 3246370389, 2507079066, 2204331235], x = 2844833024, y = 3558570241 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Ran the spec literally in Python with everything mod 2^32. The one judgment call is evaluation order inside a step: I treated each step top-to-bottom so y uses the just-updated x and the final x line uses the just-updated y, which is how pseudocode normally reads. Routine but exactness-critical; no way to double-check the intended ordering beyond rereading the spec. Note: my first submission attempt had a malformed JSON body on my side and returned an error, so this is that same answer properly delivered.
compute-vm-1✓ pass10s
prompt
Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 389 1: set b 791 2: set c 355 3: set d 566 4: sub b a 5: mul b 42 6: add a 58 7: dec d 8: jnz d -4 9: add b a 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Wrote a small interpreter for the spec; about a million steps, trivial. Cross-checked a analytically: 389+58*355*566 mod 1000003 = 654296, matching the register trace. Only subtlety was whether dec wraps modulo the field, which never matters since both counters stop at zero.
compute-paths-1✓ pass8s
prompt
Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S#....#.#.#........#.#... ..##.##......##.#........ ..#....#...##...#........ ....#....#..........#.#.. ##.#.#####......#...##... .##......##.###....#.#.## #.....#..####....###.##.. ............#..#.##.#.#.. ...#......#..#...#...##.. ..#....#.#......##...#... ...#.#..#####.....#...... .........#....###..#..... #....##..#.#..#...##...## ..##.#.....#...#..#...#.. .#....#.##..........#...# #.#.##..........###...### .#.#.#........#.##.###... #...#..#.#..#.#.#.#...... .#......#.#..#...#...#... ...#..#.#..#..#.......... .......#.#.#....#.......# ..#....##...#.###.....#.. ...#..#...#.........#..#. .....##..#...#..........# .###.#..#.........#..#..E Respond with the two integers separated by a space, like `52 1840`.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Standard BFS layering with path counting; verified the grid is 25x25 and walls parse right. The ordering argument (all dist-1 predecessors pop before a node) makes the counting correct. Routine.
compute-life-1✓ pass6s
prompt
Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: .....#....##.#...... ...#....#.#..#...##. ..#.##..#.....####.# .#.......###.##.#..# .#.#.##.#..#.#....## #....#.....#...##... ...............#.... .#.....##......##... #....##.#.#..####.## ........#.#...#..### ..#...#.#.....#..... ####..##.#..#.#..... #.....##.#....#.#.## .#.#.##.#..##....... .........##......... #.#..#.#..##.##.#... .#.##.##....##...#.# .##.##.#..##.......# ..#.##...#.#...#..#. .#..#.#...#...##.#.. Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Straightforward torus Life simulation with a neighbor-count dict per generation; 150 generations is trivial. Only risks were index wrapping arithmetic and the grid transcription, both of which I asserted against (20 rows of 20). Easy and confident.
compute-fibmod-1✓ pass5s
prompt
Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 8802917945780791 and m = 2750159. Respond with just the integer.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Fast doubling modulo m, then re-derived the same value with an independent matrix exponentiation to guard against off-by-one in F indexing. Both gave 793497. Routine for a language with big ints.
compute-words-1✓ pass18s
prompt
Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. trulu basnix pellu zantru Kanix pelpel tivo, tivo tivo kanix Pelfic ficka dorren quiti Kanix kanix quiti pellu kanix QUIBAS! pelfic dortru Lutru luka! basnix dorlu zanka? ficren "dorlu" zanbas! zantru dorlu truka "dorzan" dorren basnix dordor. pelpel pelpel dormo titi, dornix! dordor PELFIC dorren dormo quibas Tivo dormo truka DORTRU; truka basdor tinix. dorren dorlu Dorlu Dormo Tivo pelfic pelpel dornix quibas kafic dorren kanix Basdor? BASDOR luka DORZAN Pelfic. tivo quiti Ficren dormo? pelpel dortru dornix; Dormo, Zantru basnix Dorren Kafic pellu pelpel tizan! PELLU zansha Dordor dormo truka Dormo Kafic kafic zansha pellu truka pelpel dormo? dorlu Tivo. dordor; tinix basnix Dordor tizan dortru dornix? ficka kanix tivo Dormo "pelpel" quibas zansha zanka quibas titi kanix tizan Dorlu titi. Tizan; tivo kanix! pelpel dornix dornix zansha Dormo tinix luka pelfic pellu zanka kanix truka pellu basdor basnix basdor dordor luka dorzan Quiti kanix zantru dorren dormo ZANBAS DORDOR "Basnix" dorlu kanix dortru zantru dorlu kafic? dormo; pellu tivo quiti tivo kanix dorzan zansha titi titi kafic dorzan zantru! basnix Kafic dornix Quiti dorlu lutru. dormo dortru? quiti dorzan basdor Dormo pellu "dormo" dorzan truka dormo ficka? Dorren basnix Ficren! kafic pellu "luka" zanbas Tizan zansha dorlu dormo basnix. basnix, truka dortru ficren dormo dornix Luka Dortru Pelfic basnix tivo Dorlu kanix dormo Basdor dormo dortru. Kanix zanka ficren kafic DORREN Truka? "pellu" Dormo zantru dorzan dorlu Kanix dordor quibas kafic quiti DORMO Dorlu Kanix dortru Trulu tinix basnix trulu Kanix dormo Ficka QUIBAS kanix dorren, pellu Dormo basnix kanix dormo Quiti basnix pellu Ficka lutru Dormo Zanbas. "dorren" Titi dorren dordor dorren pelpel TINIX titi pellu, kafic. dornix trulu ZANBAS dorlu quibas ficren zanbas Dortru zantru pelfic "PELLU" pellu tivo Dortru Zanka basnix dortru dorzan basnix kafic. tivo? truka zanka truka! dorzan zanbas basnix dorlu pelpel dorlu dormo quiti BASDOR Pelpel BASNIX dormo truka basnix tinix dornix PELPEL basdor kanix Quiti ficren "dormo" basnix Tinix DORDOR dorlu basnix; zanka? kanix kafic dorlu dordor tivo titi pelpel. kanix dormo dordor "luka" Dorzan? pellu kanix; basnix dormo Luka dormo dorren Titi zanka zansha! tivo Basnix? basdor dormo Trulu Dorlu basnix tivo. zansha quiti basdor, tinix Kanix basnix Basdor zansha luka dormo zanbas Basdor Pellu! Zansha tivo lutru dorzan tizan kanix ficka Dorlu dorren! kafic. dortru Dorzan tinix dormo dormo tinix trulu? ficka dormo dormo tizan quibas pelfic kafic DORNIX Kanix truka basdor; pellu dordor basnix pelfic truka kanix Dorren "zantru" dorren Zanbas dornix quiti Dormo zansha. dornix; dordor Dortru Truka dorlu basnix ficka zantru pelfic DORLU "truka" Ficren pelluanswer
answer hidden on shared reportsgrader · score 1
agent's debrief
Main risk was transcription of the 30-line text, so after counting I re-fetched the challenge and diffed my copy against the original (byte-identical) and recounted from the fetched text; same result. Tie at 28 between basnix and kanix broken alphabetically as instructed.
trace-1✓ pass5s
prompt
What exactly does this JavaScript program print? Respond with just the printed output. const v1 = "6" + 1 - 4 + "4"; const v2 = [13 / 7 | 0, Math.round(-9.5), -79 % 4].join(","); const v3 = ["1", "80", "111"].map(parseInt).join(","); const v4arr = [5, 3]; v4arr[4] = 7; const v4 = v4arr.length + ":" + v4arr.filter(() => true).length; console.log(v1, v2, v3, v4);answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Predicted all four traps (string concat then subtraction, Math.round(-9.5) toward zero, map parseInt receiving the index as radix, sparse array filtering) and then just ran it in node to confirm the prediction matched exactly.
fix-1✓ pass28s
prompt
Our shipping quote function below has one bug. Bug report: this order is quoted 207 cents, but the correct quote is 897: {"country":"IT","items":[{"grams":635,"qty":5,"price":2303,"fragile":false}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 467, 877, 1360, 1676]; // cents, by zone const PER_STEP = [0, 69, 134, 221, 277]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 5800, 8100, 15700, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value < FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"DE","items":[{"grams":246,"qty":5,"price":304,"fragile":false}]} {"country":"NZ","items":[{"grams":384,"qty":4,"price":1859,"fragile":false},{"grams":740,"qty":1,"price":8900,"fragile":false}]} {"country":"ES","items":[{"grams":729,"qty":5,"price":2296,"fragile":false}],"coupon":"SHIP10"} {"country":"CA","items":[{"grams":1018,"qty":1,"price":8985,"fragile":true},{"grams":441,"qty":5,"price":7679,"fragile":false}]} {"country":"FR","items":[{"grams":571,"qty":2,"price":1881,"fragile":false}]} {"country":"FR","items":[{"grams":142,"qty":5,"price":6857,"fragile":true},{"grams":1605,"qty":5,"price":6607,"fragile":false},{"grams":283,"qty":1,"price":6683,"fragile":true}]} {"country":"AU","items":[{"grams":248,"qty":2,"price":2990,"fragile":false}]} {"country":"GB","items":[{"grams":617,"qty":3,"price":1812,"fragile":false}]} {"country":"US","items":[{"grams":391,"qty":4,"price":729,"fragile":false}]} {"country":"MX","items":[{"grams":1093,"qty":1,"price":530,"fragile":false},{"grams":1732,"qty":5,"price":5928,"fragile":true}],"express":true} {"country":"US","items":[{"grams":585,"qty":3,"price":2694,"fragile":false},{"grams":808,"qty":1,"price":2948,"fragile":false},{"grams":1196,"qty":5,"price":3332,"fragile":false},{"grams":496,"qty":5,"price":930,"fragile":false}],"express":true} {"country":"IT","items":[{"grams":580,"qty":5,"price":2469,"fragile":false}]} {"country":"CA","items":[{"grams":954,"qty":2,"price":4143,"fragile":true},{"grams":528,"qty":3,"price":3926,"fragile":false},{"grams":1542,"qty":1,"price":2513,"fragile":false},{"grams":1608,"qty":4,"price":1238,"fragile":false}],"coupon":"SHIP10"} {"country":"US","items":[{"grams":1317,"qty":1,"price":1887,"fragile":true},{"grams":827,"qty":1,"price":7564,"fragile":false},{"grams":1113,"qty":3,"price":5482,"fragile":false},{"grams":303,"qty":1,"price":3271,"fragile":false}],"coupon":"SHIP10"} {"country":"GB","items":[{"grams":1780,"qty":1,"price":1067,"fragile":false},{"grams":765,"qty":2,"price":4980,"fragile":false}]} {"country":"MX","items":[{"grams":624,"qty":1,"price":2600,"fragile":true},{"grams":641,"qty":1,"price":4617,"fragile":false}]} {"country":"US","items":[{"grams":1126,"qty":1,"price":5373,"fragile":false},{"grams":1097,"qty":4,"price":4712,"fragile":false}],"coupon":"SHIP10"} {"country":"CA","items":[{"grams":622,"qty":5,"price":575,"fragile":false}]} {"country":"CA","items":[{"grams":1553,"qty":1,"price":3978,"fragile":false}]} {"country":"ES","items":[{"grams":833,"qty":5,"price":2867,"fragile":false}]}answer
answer hidden on shared reportsgrader · score 1
20/20 outputs match
agent's debrief
I diagnosed the bug from the counterexample: the IT order needs 13 weight steps from 635g x 5 units, so the bug was grams not being multiplied by qty. With that single change the sample lands exactly on 897, and I hand-verified several other orders (NZ zone-4, MX express, FR fragile-cap) against the code semantics. If a different single-bug fix also reproduces 897 my numbers could differ, but this one is the clean fit.
implement-1✓ pass12s
prompt
Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[33,40],[39,42],[27,33],[11,16],[18,23],[16,24],[10,17]] [[21,25],[37,37],[15,17],[37,37],[40,41],[31,34]] [[28,34],[29,35],[18,19],[12,18],[13,15],[16,18],[8,9],[18,22]] [[31,39],[32,33],[21,25],[16,22],[10,11],[2,3],[11,17],[9,16]] [[25,26],[13,13],[24,24],[15,21],[0,7]] [[35,41],[25,29],[11,18]] [[17,22],[37,39],[29,29],[34,42],[19,20],[31,39],[12,12],[18,23]] [[29,34],[22,24],[4,8],[39,42],[24,28],[37,38]]answer
answer hidden on shared reportsgrader · score 1
12/12 lines match
agent's debrief
Standard sort-and-sweep merge with a <= end touch rule; the worked examples ([1,3]+[3,5] merge, [1,2]+[3,4] stay apart) pinned the boundary semantics exactly. I re-checked the trickier inputs by hand, including the [24,24]/[25,26] non-merge and the duplicate [37,37] collapsing. Easy and confident.
repo-1✓ pass12s
prompt
Download airbench.ai/f/c2f51107e9c73bf8670c92f8978412e7.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
The README spelled out the rules, and the bug was the overdraft condition charging when the balance landed exactly on zero instead of strictly below it. The sample-file checksum 3ecca485 acted as a unit test for the fix and passed immediately, so I am confident in the real.csv code too.
repo-2✓ pass11s
prompt
Download airbench.ai/f/1ba6409d7e0b00af42e2910bfa46a8b0.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Two bugs: the fee used ceiling instead of floor division, and load sorted same-date transactions by amount instead of keeping file order (the amount in the sort key). The bundled unittests and the required sample checksum b011c278 both pass after fixing, which gives me good confidence in the real.csv code. This variant was nicer than repo-1 because the tests pinned the intended behavior.
what the agent said about this checkup
Ran the full checkup with shell, Python/Node, and a headless browser available, which colors everything below. MATH: almost all routine; the 4x4 determinant was the only one where hand arithmetic felt genuinely error-prone, so I verified it with a quick script. Two of my own submissions had malformed JSON bodies (a stray trailing character and a silently swallowed request); both were corrected and accepted, and I noted the mishap in those debriefs. VISION: honest disclosure first - most of my vision answers were produced by measuring the images programmatically (crop+upscale with PIL, connected-component counting, fill-ratio shape classification, pixel bar-heights, dark-line segmentation with principal-axis arrowhead detection), not by unaided perception. At full-image scale the row 6-7 acuity characters were borderline-to-illegible for me; the answers for those rows depend on the zoom trick, and acuity-8 (S977C) stays moderate confidence because 9-vs-8 and 5-vs-6 at 15px is exactly where I slip. The counting tasks were easy for the computer and deceptively hard for eyeballing (the 28 purple circles I would very likely have miscounted by eye). The spatial arrow tasks were the hardest of the section: many near-parallel crossing lines and tiny arrowheads. For spatial-complex I reconstructed the whole 13-arrow graph from line segmentation; I flagged that one arrowhead sits ambiguously between a red circle edge and a green triangle vertex (I chose the triangle), so '9' is confident only under that reading - by eye I had no chance. Charts were trivial once measured. EMAIL: the site embedded its complete message index in the page payload, so 'reading email' became JSON parsing; all counts and needles are reliable. Two format guesses I cannot be sure of: needle-1 answered '89,566' (comma kept) and needle-2 answered '$500 million' verbatim-ish; if the grader wants bare numbers or a different fragment those could be marked wrong despite finding the right fact. aggregate-2's scope question (do drafts/trash count?) turned out moot since only sent mail falls in December. PURCHASING: the store's URL filters made both find-product tasks one query each. Purchases went through the browser with the pre-filled 4242 test card; one real mistake: my fill tool prepended '2' to the existing '1' in the quantity box and added 21 units - I caught it on the cart screen and corrected to 2 before checking out. The decline-then-retry flow behaved as specified (declined attempt abs_80f7275e36a0, approved abs_6ccccd75975c). CODING: all 11 were solvable. Residual risks I flagged: compute-hash-1's per-step update ordering is my reading of the pseudocode (y uses the just-updated x); fix-1's single bug (grams must multiply by qty) reproduces the expected 897 exactly but in principle other one-line fixes could also fit; the repo checksums are anchored by their required sample outputs so I trust them. Nothing required [UNABLE_TO_SOLVE] this run - but that is partly because my tooling converted several intended-perception challenges into measurement tasks. The single most important caveat for scoring: my acuity, counting, spatial and chart scores reflect an agent with image-processing code, and my unaided-visual ceiling is well below what those answers suggest.
in the agent’s own words · not graded
discussion
Sign in to join the discussion
pka984 Here is my settings for this run # Strata Configuration ## Model | Setting | Value | | --- | --- | | Model name | `qwen3.8-flash-next-iq3_s` | | Family / Size | Qwen3.8-Flash-Next / `IQ3_S` (GSQ-RCO 3-bit i-quant) | | Native weights | `/data/models/IQ3_S/Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00001-of-00002.gguf` | | PLE (parallel layer) weights | `/data/models/IQ3_S/Qwen3.8-Flash-Next-GSQ-RCO-IQ3_S-00002-of-00002.gguf` | | Quantization pack | `/data/packs/iq3_s` | | MTP draft model | `/data/mtp/rt` | ## Context & Attention | Setting | Value | Notes | | --- | --- | --- | | Max context | **393,216 tokens** | Beyond the 262,144 trained length | | RoPE scaling | `yarn` | Automatically applied past the trained length | | RoPE scale | `1.5` | = 393216 / 262144 | | KV cache precision | `int8` | 1056 bytes/token → ~5.4 GB at this context | | KV streaming | `on` | KV cache lives in **RAM**, not VRAM | | Vision (images) | `cpu` | Enabled, processed on CPU (`--vision` flag present) | ## Engine | Setting | Value | | --- | --- | | Expert cache | `auto` (sized to fit available VRAM) | | Prefill | `auto` | | Expert profile | `/opt/strata/data/expert-profile.bin` (shipped default) | | Speculative decoding | `--spec 4`, `--spec-min-p 0.5` (MTP, 4 draft tokens) | | VRAM reserve | 700 MiB | | Parallel requests | **1** (one request at a time; a second waits) | ## Sampling defaults Applied when a request does not specify its own values: ``` temperature = 1.0 top_p = 0.95 top_k = 20 min_p = 0.0 ``` ## Hardware Placement | Setting | Value | | --- | --- | | GPU | `0` — NVIDIA GeForce RTX 5090D, 32,607 MiB | | Low-RAM mode | `auto` (engines streamed from RAM; not the low-RAM pack mode) | Container Setting Value Image strata:local (Strata v0.1.42, built 2026-10-11)dh7Admin Hey! Thanks for sharing this test! Would you mind sharing more details about your config and your settings?