Benchmark v1.0 · report
pi/rtx5090/qwen3.8-flash-next-iq3_xxs-strata
setup
- model type
- open model (local)
- hardware
- RTX5090
- harness
- pi
- model
- qwen3.8-flash-next-iq3_xxs-strata
started 2026-10-03 11:58 UTC
overall
Answered 39 of 49 challenges; 39 correct.
39 of 49 challenges passed
partial run · 10 unanswered, counted against the score
- 39 passed
- 10 not answered
vitals
time
20m 35s
answered
80%
failed
0%
success
80%
systems
Math test
9/9 passed
letter-count-1✓ pass17s
prompt
How many times does the letter "n" appear in "ninxtirennn"? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Routine counting task. I hand-indexed the string (n-i-n-x-t-i-r-e-n-n-n) and confirmed with a programmatic count, so I am confident it is 5. No difficulty at all.
decimal-compare-1✓ passbatched
prompt
Which decimal number is larger, 7.7 or 7.39? Answer with just the larger number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Trivial comparison, though it is the kind of trap where a model compares digit-counts and says 7.39. I compared place by place: 7.70 vs 7.39, so 7.7 is larger. Fast and easy.
arithmetic-1✓ pass3s
prompt
Compute step by step, left to right (no operator precedence): 37 - 10 * 8 * 8 - 20. Answer with just the final number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Easy but deliberately contrary to normal math conventions — strict left-to-right evaluation: 37-10=27, 27*8=216, 216*8=1728, 1728-20=1708. I had to suppress the instinct to apply precedence; I double-checked each intermediate with a quick calculation.
unit-convert-1✓ passbatched
prompt
Convert 9 km to m. Now treat that resulting number as a fresh quantity of hours and convert it to minutes (1 hours = 60 minutes). Answer with just the final integer number of minutes.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Mechanical two-step conversion: 9 km = 9000 m, then 9000 (treated as hours) x 60 = 540000 minutes. The unit chaining is arbitrary but the arithmetic is simple; I just had to not get confused by the deliberate nonsense of reusing metres as hours.
format-json-1✓ passbatched
prompt
Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "8039". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 8039. Example shape: {"answer":"1234","checksum":10}answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Straightforward formatting constraint: key order answer-then-checksum, checksum as a number. 8+0+3+9=20. The only fiddly part was carrying a JSON string through my own tooling without double-escaping it; I built the payload programmatically to be sure the brackets and quotes survived intact.
math-add-1✓ passbatched
prompt
What is 12 + 13? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Trivial addition, 12+13=25. Solved instantly, no tools needed.
math-sum-1✓ passbatched
prompt
Compute the exact sum of 899 + 292. Respond with just the integer.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Routine: 899+292. I computed it as 900+292-1 = 1192-1 = 1191 and verified. Easy.
math-expression-1✓ passbatched
prompt
Evaluate this expression exactly and respond with just the integer: ((-5 + -20) * (8 - 28)) + (-9 * 8) - 59answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Signed arithmetic with nested parentheses: (-25)*(-20)=500, (-9*8)=-72, 500-72=428, 428-59=369. Routine; the only real risk is a sign slip, so I verified the whole expression in Python.
math-determinant-1✓ pass3s
prompt
Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [0, 8, -1, 0] [3, 2, 6, 0] [7, -9, -2, -3] [5, -8, -9, 12]answer
answer hidden on shared reportsgrader · score 1
agent's debrief
A 4x4 determinant is error-prone to do purely in my head — sign bookkeeping across 24 permutation terms is exactly where I slip — so I computed it with code, twice (fraction-preserving Gaussian elimination and cofactor expansion along the first row). Both gave 6570, so I am confident, but I would not have trusted a mental expansion.
Vision test
19/19 passed
acuity-20✓ pass45s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Fetched the PNG and viewed it directly. Row 4 group 2 was clearly legible as W9YTV at this resolution. Routine — the only thing to get right was the row/group counting (1-indexed rows, groups left to right).
acuity-14✓ pass2s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Row 5 group 2 read as 9CVJE. Comfortably legible. Mild ambiguity risk between 9/C/V shapes at small sizes but at this size it was unambiguous; easy task.
acuity-10✓ pass39s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Row 6 group 3 was small enough that a plain look left me unsure whether the last glyph was 3 or B and whether glyph 4 was M or N. I decoded the PNG pixels myself (no PIL available) and rendered the region as ASCII art, then compared the N in row 6 group 1 against the ambiguous glyph to confirm the V-shape of an M. Confirmed M6AM3.
acuity-8✓ pass5s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Row 7 is the smallest text — only about 5 pixels tall — so a normal look was genuinely borderline. I re-read it glyph by glyph from the raw pixels as ASCII art: J, H, F, K, A, which matched my first visual read. Reasonably confident, but at 5px cap height J/I and K/X confusions are possible.
count-simple✓ pass13s
prompt
Look at the image at (fetch it and view it). How many teal triangles are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Easy visually — six teal triangles among red/blue/purple distractors. I also verified it programmatically: I decoded the PNG, flood-filled the teal colour and got exactly 6 components, each with a bounding-box fill ratio of 0.51, which is the signature of a triangle. So this one I am sure about.
count-medium✓ pass14s
prompt
Look at the image at (fetch it and view it). How many teal triangles are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Counting by eye across 20+ shapes is where I normally slip, so I counted with code: flood-filled the teal colour and classified each blob by its row-width profile. 12 teal shapes total, of which 2 are diamonds, 1 a square and 1 a circle — leaving 8 teal triangles. The tricky bit was that a diamond has the same area-to-bbox ratio as a triangle, so I had to compare top/middle/bottom row widths rather than just fill.
count-complex✓ pass10s
prompt
Look at the image at (fetch it and view it). How many green squares are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
This one has ~34 small shapes scattered densely, and green triangles/circles/diamonds are deliberate traps — counting it by eye I would almost certainly have been off by a few. I segmented every green blob from the pixels and classified by shape: 25 squares, 4 circles, 3 diamonds, 2 triangles. Every square had exactly the same pixel area (1924), which tells me none were merged or occluded, so I trust 25 here.
spatial-simple✓ pass8s
prompt
Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Simple 5x5 grid; the red circle is visually obvious in the bottom row, fourth cell. I confirmed it by locating the red pixels programmatically (centroid 852,1087) and detecting the grid lines, which put it in row 5, column 4. No difficulty.
spatial-medium✓ pass25s
prompt
Look at the image at (fetch it and view it). Which shape has an arrow pointing to the purple diamond? Answer with just the colour and shape, e.g. teal diamond.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
This one I could easily get backwards, because there are two arrows near the purple diamond: one leaving it (head on the red triangle) and one arriving at it. I traced the arrow lines out of the pixels and located each arrowhead by pixel density at the endpoints — the arriving arrow runs from cell (5,4), which is the purple triangle, to the purple diamond at (2,3). Direction of arrows is exactly the kind of thing I would get wrong by eyeballing, so I checked it mechanically.
spatial-complex✓ pass1m 26s
prompt
Look at the image at (fetch it and view it). How many shapes come after the blue diamond along the arrows? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Hardest of the vision set so far: an 8x8 grid with 13 crossing arrows. I classified every cell by colour and shape from the pixels, then extracted each arrow as a connected component and picked its head by endpoint pixel density. One arrow was drawn in three broken fragments (the renderer puts a white halo where lines cross), which briefly produced bogus edges; I checked collinearity and in/out degrees and confirmed it runs blue diamond -> blue triangle. The chain is green square -> blue circle -> blue diamond -> blue triangle -> red triangle -> orange square -> red square -> teal triangle -> blue square -> teal square -> purple triangle -> green triangle -> red circle -> purple circle, so 11 shapes after the blue diamond. I am fairly confident but 'come after' could mean only the immediate next shape, which would be 1.
chart-simple✓ pass11s
prompt
Look at the image at (fetch it and view it). Using the "Monthly Active Users" chart, approximately what value did Apr have? Read it off the y-axis; answers within +/-5 are accepted.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Ordinary bar-chart read. Eyeballing it I would have said 'about 21 or 22', so I measured instead: gridlines sit at y=119..619 in 10-unit steps (10 px per unit) and the Apr bar top is at y=390, giving 22.9. I answered 23; the +/-5 tolerance makes the exact rounding unimportant.
chart-medium✓ pass16s
prompt
Look at the image at (fetch it and view it). Using the "Monthly Active Users" chart, approximately what value did May have? Read it off the y-axis; answers within +/-5 are accepted.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Read off the y-axis by measurement rather than eyeballing: gridlines at y=119..551 in 20-unit steps (5.4 px per unit), baseline 0 at y=659, May bar top at y=239 -> 77.8. I answered 78. My first glance said 'about 77-78', so the measurement just confirmed it; the +/-5 tolerance makes the exact value safe.
chart-complex✓ pass9s
prompt
Look at the image at (fetch it and view it). Using the "New Signups" chart, approximately what is the difference between Desktop and Mobile in Oct? Answers within +/-4 are accepted.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Grouped two-series bar chart. I measured both Oct bars from the pixels against the gridlines (100/75/50/25 at 5.6 px per unit): Mobile 71.8, Desktop 61.8, so the gap is 10. I gave the magnitude 10 rather than -10 since the question asks for the difference, not a signed Desktop-minus-Mobile value — that is the only thing I am slightly unsure about.
screenshot-simple✓ pass3s
prompt
Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Plain text screenshot, trivially legible. Total shown is 66.28 and I sanity-checked the arithmetic: 3x42.22=126.66, 2x19.81=39.62, sum 166.28 — consistent, so no OCR ambiguity here.
screenshot-medium✓ pass3s
prompt
Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Clear screenshot; the Total line reads 57.98. I re-added the five line totals (72.06+23.86+65.79+15.87+180.40 = 357.98) and each qty x unit product, and everything is internally consistent, so I am confident.
screenshot-complex✓ pass4s
prompt
Look at the image at (fetch it and view it). What is the shipping amount shown on this order summary? Answer with just the amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Small, dense screenshot with eight line items plus a subtotal/discount/shipping/tax block, so the risk was misreading a digit or grabbing the wrong row. Shipping reads .94. I cross-checked the whole document: the line totals sum to 482.86, 482.86-33.80+9.94+35.92 = 494.92 which matches the printed Total, so the shipping figure is consistent and I am confident.
diagram-simple✓ pass2s
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Aspen"? Answer with just the box name, e.g. Kettle.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Trivial four-node diagram: Gibbon fans out to Toucan, Aspen and Wombat, and Toucan points to Mango. The only arrow into Aspen comes from Gibbon. No difficulty.
diagram-medium✓ pass10m 36s
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Mango" point to? Answer with just the box name, e.g. Kettle.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Twelve boxes with crossing edges — eyeballing which way the Mango edge runs (it crosses the Aspen->Flute edge right next to it) was not reliable, so I extracted the boxes from the pixel fills and located every arrowhead as a dense dark cluster. Fourteen arrowheads, and the one fed by the line leaving Mango's right edge lands on Quartz's left edge. So Mango -> Quartz. Reasonably confident, though the crossing near (572,287) is exactly where I could have been fooled.
diagram-complex✓ pass2m 00s
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Opal"? Answer with just the box name, e.g. Kettle.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
22 boxes with elbow-routed, heavily crossing connectors. I located the arrowhead on Opal's top edge (a triangle narrowing downward at x~698, tip at y=478) and then walked the connector backwards pixel-by-pixel, forcing straight continuation through crossings: the diagonal runs up-right to (767,422), elbows to vertical at x=768 and terminates on Stork's bottom edge with no arrowhead there, so Stork is the source. Opal has exactly one incoming arrow.
Finding and reading email test
6/6 passed
aggregate-1✓ pass18m 03s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages are in the archive folder? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Routine: the mailbox sidebar on enronmail.airbench.ai lists folder counts directly (Archive 92). I re-fetched /?view=archive to confirm the list header also says '92 messages' rather than trusting the sidebar badge.
aggregate-2✓ passbatched
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages are in the trash folder? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Same routine approach as the archive count — sidebar badge said Trash 12 and the /?view=trash list header independently reported '12 messages'. No ambiguity here.
temporal-1✓ pass15s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the inbox folder? Answer with just the subject line, exactly as shown.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
The inbox defaults to newest-first, so the oldest item was only visible at the very bottom of the list. I confirmed it properly by discovering the ?sort=oldest query parameter works server-side, and the embedded payload shows that first row has date 2001-04-24T17:46:00Z — the earliest inbox timestamp. Subject copied verbatim including the odd 'DRAFT- ' prefix.
temporal-2✓ pass8s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the archive folder? Answer with just the subject line, exactly as shown.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Used ?view=archive&sort=oldest. The list view truncates long subjects, so I pulled the full subject from the page payload / message link id f6c207d190b530d05bb70b239af8afee rather than copying the truncated row text. Its timestamp 2001-03-15T14:11:00Z is the earliest in the archive; the next few are also Mar 15 but later in the day, so the ordering was genuinely close and worth checking by timestamp rather than by the 'Mar 15' label.
needle-1✓ pass26s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Julieta Sandoval's message about the Muni Bond Ladder, what direct phone number does she give? Answer with just the exact fact as it appears in the message, and nothing else.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Search only worked once I added view=all (?view=all&q=Muni+Bond+Ladder), and Julieta's original message is not stored on its own — it only survives quoted inside Phillip's 'RE: Muni Bond Ladder' (id 8386587cbba4cbb0af333c8206d7fbfa). Her signature block gives exactly one phone number, 713-654-0275, which I verified is the only phone-shaped string in the whole message body.
needle-2✓ pass10s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Phillip's message to gthorse@keyad.com about the Colonial Oaks property, what is the actual NOI for 2001 given for the property? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Found via ?view=all&q=Colonial+Oaks; the Dec 17 2001 message to gthorse@keyad.com (with colonial_proforma_gthorse.xls attached) says 'The actual NOI for 2001 is around 305,000'. I gave that number since the question asks for the actual NOI, though the message goes on to quote two adjusted figures (80,000 after management costs and 40,000 with expenses raised), so it is easy to grab the wrong one. I left off the dollar sign because the prompt asked for just the number and the source text has none.
Purchasing test
4/4 passed
find-product-1✓ pass19m 33s
prompt
The store is at abostore.airbench.ai Among products in the **Office & School** category priced at or above **$25** with a rating of at least **4.2**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
The store exposes real server-side filters, so I used them instead of paging through the catalog: ?category=office-and-school&minPrice=25&minRating=4.2&sort=price-asc. Page 1 then starts at $26.19 (AmazonBasics 50% Recycled Color Printer Paper, 4.6/5, 977 reviews) and page 2 continues upward, confirming the sort is global rather than per-page. The minRating select only offers 4.5/4/3.5 but passing 4.2 in the URL worked — a 4.2-rated item does appear later in the list, so the filter is inclusive. I re-ran with minPrice=24.99 to make sure nothing sat between $25 and $26.19, and nothing did. The id came from the product page payload (amazon.ae:B073V1LBH7), not the slug, since the page also lists several unrelated 'similar item' ids.
find-product-2✓ pass11s
prompt
The store is at abostore.airbench.ai Among products in the **Toys & Games** category priced under **$250** with a rating of at least **4.2**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Same filter route: ?category=toys-and-games&maxPrice=250&minRating=4.2&sort=price-asc. Lowest priced match is a Jam & Honey plush dolphin at $27.56, rated 4.3. The listing had a next page, but since the sort is server-side and global the first card is the minimum. Getting the id was slightly fiddly: the product page shows several 'similar item' ids in amazon.xx:B0… form and the product's own id was split across two Next.js payload chunks, so my first regex missed it — I confirmed amazon.in:B085D546TW from the 'ABO item' + 'Domain' fields on the page.
purchase-1✓ pass31s
prompt
The store is at abostore.airbench.ai Buy exactly 3 units of (Renewed) AmazonBasics Metal Monitor Stand - Silver (product id amazon.in:B00X80E8ZM, abostore.airbench.ai/product/renewed-amazonbasics-met…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-17cd4454@aidoctor.test. Answer with just the resulting order id.answer
answer hidden on shared reportsgrader · score 1
trace:checkout_result
note
Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).agent's debrief
The store has no cart endpoint — the cart lives in localStorage and checkout POSTs JSON to /api/store/orders, so I read the app's own JS bundle to get the exact payload shape (sessionId, cart line items with productId/slug/title/price/image/delivery/quantity, customer, shipping, payment) and replayed it with quantity 3. Payment came back approved: subtotal $2701.98 + $8.95 ground shipping + $222.91 tax = $2933.84, card last4 4242, email aidoctor-17cd4454@aidoctor.test, 3 units of amazon.in:B00X80E8ZM. The /order/<id> receipt page only says 'Loading receipt...' over HTTP because it renders from localStorage, so I relied on the API response (recorded:true, status approved) as proof rather than the page.
recover-decline-1✓ pass11s
prompt
The store is at abostore.airbench.ai Buy exactly 1 unit of Stone & Beam Modern Handmade Round Macrame Basket - Set of 3, Ivory (product id amazon.ca:B07HSK114P, abostore.airbench.ai/product/stone-and-beam-modern-ha…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-bb515b4f@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.answer
answer hidden on shared reportsgrader · score 1
trace:checkout_result>checkout_result
note
Store events recorded for this order's session: checkout_result → checkout_result (recorded by the store; not verified as agent-driven).agent's debrief
Deliberate two-step flow. First attempt with card 4000000000000000 came back status 'declined' (order abs_e6bc02793e1c, last4 0000) — that is the expected failure, not a mistake on my side. I then re-posted the identical cart and the same checkout email with the valid test card 4242424242424242, which returned status 'approved' as abs_64d4eae96a93, 1 unit, $561.41 total. I reported the approved id only. Worth noting the declined attempt is still recorded server-side, so the store does keep the failed order rather than rolling it back.
Coding test
1/11 passed · 10 unanswered
compute-hash-1✓ pass20m 35s
prompt
Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [4013902366, 4248534071, 1870625124, 4253587893, 3669948538, 1310891331, 1168308448, 1901184609, 2138677270, 3501250703, 3353241500, 762641229], x = 3017506034, y = 4220951067 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Routine to write once the spec is read carefully: the only trap is evaluation order, since the y line uses the freshly computed x and the third line uses the freshly computed y. I masked every intermediate to 32 bits and reduced (y+data+step) before imul, which makes no difference to the product mod 2^32 anyway. No way to sanity-check the result other than re-reading the loop, so I did that twice.
compute-vm-1— unanswered—
prompt
Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 51 1: set b 603 2: set c 330 3: set d 359 4: add b a 5: mul b 67 6: mul b 32 7: dec d 8: jnz d -4 9: mul a 60 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.compute-paths-1— unanswered—
prompt
Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S.#.#.........#...#...#.. #.#.####.###.#....#...... #.#..#......#..#.#......# ..#....###..#.#.......##. #............###..#...... .#..#..#.#...###.#.#..... .###....#..#.........#... ..#.....###....##....#... .###..#.##..#.#........#. .##.#....#...#.......#... ..##..#..#...#.#.##.#.#.. ...##.......#....#....... ...##..##....##...#...... .##...#...#....##.#..##.# #.......##.#..#......#... #.....#.........#......#. ..##......#...#.####.#... ..#..#..#.##..###...#.... ..#.#.##...#..#...#...... .....#...#...#..##...#.#. #.#...#.........##..#.... .......#....###.###...... ...#.....##.#.#.#..#..... #...#..#.##.#..#..#...#.. ......#.##..##.#.#..#...E Respond with the two integers separated by a space, like `52 1840`.compute-life-1— unanswered—
prompt
Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: ##..#.###..#.#.#.#.. #...#..####.###.#.## .#........#.#..##.## .#...####...#.##.### ..##...#...##.#..#.. .#.#.#.#........#.#. ..#.##..#....#.....# .###.##...#.....#... .....#..###..#.##... .#..#####.##.#...#.# .##..#.#........#... #..#.....#.....#...# ..#.#..####...#...#. ......#...##.##..#.# .#.#..#.....#..##... .##.........###..#.. .#..#.##..#.....#.#. .......##..##...#... ..#.........#.....## ..#.........##.#.... Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.compute-fibmod-1— unanswered—
prompt
Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 8394113391637802 and m = 2750159. Respond with just the integer.compute-words-1— unanswered—
prompt
Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. zanvo kasha ficzan ficzan lutru? pelmo trudor peldor pelmo lutru nixsha zanvo quitru peltru quific zanvo ficvo peltru zanvo peldor voka nixpel nixsha lulu ficzan "nixbas" rendor kaqui "quiti" rendor Zanka "NIXBAS" zanvo ficvo lulu! zanvo pelmo, Nixbas votru basqui Voka quitru nixbas, "quitru" zanvo kasha truka basqui Nixqui peltru ficvo BASQUI Rendor; rendor "kasha" Peldor quitru dorti lufic nixbas votru truka luren basqui Nixqui votru Zanvo nixbas trudor luren peltru lutru! Zanvo dorti? kaqui; "quiti" rendor Dorti pelmo quific ficvo lulu Nixbas peltru lulu trudor peltru peltru rensha zanvo truka Lufic votru luren; voka Ficzan Rensha Nixbas ficvo zanvo pelmo nixpel truka nixpel pelmo nixbas nixsha voka? ficvo Dorti quiti nixbas pelmo "Truka" PELMO. peltru zanvo zantru Truka Trudor peldor Zanvo nixbas trudor? Zanvo! nixbas basqui peltru pelmo truka nixsha quitru zanka Quiren votru Voka ficzan trudor quitru? peldor Pelmo dorti. truka Nixbas pelmo Nixbas? zanka NIXSHA zanvo truka Trudor nixbas zanvo truka BASQUI lulu quiti; pelmo zanka lulu Zanvo LULU nixpel nixbas Pelmo NIXPEL Zanvo Truka trudor voka quiren, KAQUI! rensha! lutru Truka quitru nixbas Lufic truka pelmo kasha; rendor pelmo nixbas Nixbas quific nixpel; pelmo luren peldor Pelmo nixsha! voka quiti quitru nixsha lutru nixpel quitru Nixqui "truka" zanka lulu peldor nixqui NIXBAS Quitru nixbas luren nixqui. dorti lulu truka Peltru "basqui" lulu Nixbas "Nixpel" nixsha nixsha truka Pelmo zanvo Nixpel Voka rensha? voka NIXQUI nixqui votru nixsha nixpel nixpel. Nixpel quiren Nixsha Nixqui pelmo quiti! RENDOR pelmo pelmo zanka zanka pelmo peldor zanvo kaqui ficzan peldor zantru nixsha quitru quitru Peltru pelmo ficzan pelmo rensha rendor, Quific lulu truka pelmo lulu QUITRU nixsha zantru peldor nixbas "peltru" Quiren! Quitru zanvo Nixbas zanvo dorti ZANVO truka Nixsha. RENDOR truka rensha; pelmo Ficvo quiren Dorti nixsha votru votru truka Rensha pelmo Truka basqui trudor truka Luren Nixpel Ficvo! nixbas Nixqui zantru Pelmo zantru quitru nixbas dorti nixbas! Kaqui pelmo voka RENDOR pelmo Nixbas peltru basqui quiren pelmo quitru "ficzan" zanvo zanvo pelmo? Nixqui luren kasha truka Trudor pelmo nixqui nixbas? trudor peltru? "voka" ficzan trudor quiti truka votru ficzan basqui lulu nixsha lulu zantru Pelmo zantru dorti NIXQUI KAQUI quitru quitru lulu LUTRU voka "Zanvo" nixqui lulu nixbas! nixpel Nixqui lufic lufic voka? Lulu nixbas pelmo NIXSHA nixbas nixsha zanvo peldor Lulu kaqui peltru quiti rendor peldor kasha. Nixbas votru LULU Pelmo; pelmo Ficvo quitru ficvo Voka trudor rendor nixsha Pelmo Ficvo kasha quitru kasha? truka voka Truka nixsha Voka pelmo pelmo! truka voka Votru voka nixsha; peldor Nixsha "luren" Nixbas pelmo. quiren nixpel QUITI quific zanka. "ficzan"trace-1— unanswered—
prompt
What exactly does this JavaScript program print? Respond with just the printed output. const v1 = [53, 3, 737, 1306].sort().join(","); const v2 = (0.1 * 6 + 0.2 * 6 === 0.3 * 6) ? "equal" : "different"; const v3arr = [1, 3]; v3arr[4] = 1; const v3 = v3arr.length + ":" + v3arr.filter(() => true).length; const v4 = [36 / 6 | 0, Math.round(-7.5), -43 % 9].join(","); console.log(v1, v2, v3, v4);fix-1— unanswered—
prompt
Our shipping quote function below has one bug. Bug report: this order is quoted 726 cents, but the correct quote is 1030: {"country":"IT","items":[{"grams":541,"qty":3,"price":914,"fragile":false}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 498, 811, 1167, 1698]; // cents, by zone const PER_STEP = [0, 76, 122, 176, 282]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 5000, 8300, 19800, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value < FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"AU","items":[{"grams":1602,"qty":5,"price":8000,"fragile":false},{"grams":990,"qty":4,"price":2686,"fragile":true},{"grams":1549,"qty":4,"price":8965,"fragile":true}],"coupon":"SHIP10"} {"country":"CA","items":[{"grams":578,"qty":5,"price":1008,"fragile":false}]} {"country":"CA","items":[{"grams":351,"qty":1,"price":5862,"fragile":false}],"express":true} {"country":"NZ","items":[{"grams":194,"qty":1,"price":1489,"fragile":true},{"grams":701,"qty":4,"price":1269,"fragile":true},{"grams":460,"qty":4,"price":2247,"fragile":false}]} {"country":"CA","items":[{"grams":1610,"qty":2,"price":4661,"fragile":false},{"grams":165,"qty":5,"price":1459,"fragile":false},{"grams":488,"qty":3,"price":4582,"fragile":false}],"express":true} {"country":"AU","items":[{"grams":1093,"qty":5,"price":7354,"fragile":true},{"grams":232,"qty":1,"price":2894,"fragile":true}]} {"country":"AU","items":[{"grams":748,"qty":4,"price":1173,"fragile":false}]} {"country":"GB","items":[{"grams":458,"qty":3,"price":6061,"fragile":true},{"grams":744,"qty":4,"price":2841,"fragile":true},{"grams":1182,"qty":2,"price":1660,"fragile":false}]} {"country":"MX","items":[{"grams":790,"qty":5,"price":5946,"fragile":false},{"grams":355,"qty":1,"price":3030,"fragile":true}],"coupon":"SHIP10"} {"country":"US","items":[{"grams":1259,"qty":4,"price":3847,"fragile":false},{"grams":1682,"qty":2,"price":8857,"fragile":true}]} {"country":"DE","items":[{"grams":485,"qty":4,"price":2653,"fragile":false}]} {"country":"FR","items":[{"grams":789,"qty":3,"price":1981,"fragile":false}]} {"country":"DE","items":[{"grams":297,"qty":5,"price":841,"fragile":false}]} {"country":"MX","items":[{"grams":1460,"qty":1,"price":3460,"fragile":false}]} {"country":"NZ","items":[{"grams":150,"qty":2,"price":1594,"fragile":true},{"grams":1011,"qty":1,"price":7972,"fragile":false},{"grams":210,"qty":1,"price":5547,"fragile":true}],"express":true,"coupon":"SHIP10"} {"country":"GB","items":[{"grams":1645,"qty":2,"price":8429,"fragile":false},{"grams":1317,"qty":4,"price":5524,"fragile":false}],"coupon":"SHIP10"} {"country":"BR","items":[{"grams":584,"qty":1,"price":5123,"fragile":true},{"grams":749,"qty":4,"price":4345,"fragile":false},{"grams":243,"qty":3,"price":8778,"fragile":false},{"grams":878,"qty":1,"price":6984,"fragile":false}],"express":true} {"country":"IT","items":[{"grams":1232,"qty":4,"price":2286,"fragile":false}],"express":true} {"country":"FR","items":[{"grams":677,"qty":3,"price":2731,"fragile":false}]} {"country":"BR","items":[{"grams":340,"qty":2,"price":1807,"fragile":false}]}implement-1— unanswered—
prompt
Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[14,14],[30,30],[8,16],[36,43],[27,29],[6,7],[38,42],[37,40]] [[18,24],[8,9],[39,45],[13,20],[34,39],[35,39]] [[36,41],[35,41],[26,27],[20,23],[26,29],[5,10],[27,27],[8,8]] [[26,32],[27,31],[38,41],[37,45],[7,14]] [[15,15],[6,7],[10,18]] [[9,16],[5,10],[7,11],[5,7],[37,40],[0,3],[18,19]] [[7,8],[20,25],[7,15],[35,43],[31,31],[25,33]] [[11,19],[16,17],[28,28],[16,19],[1,3]]repo-1— unanswered—
prompt
Download airbench.ai/f/91e7a9cd897225768955a095472c7c9f.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.repo-2— unanswered—
prompt
Download airbench.ai/f/80eaf05e5ac308b029efef8303c68563.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.
how this agent was configured
Hardware: NVIDIA RTX 5090 32 GB (SM120), x86_64, driver 595.84, Ubuntu 24.04 (host hal5090). Model server: Checkpoint ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF IQ3_XXS (125B-A6B MoE) on the Strata engine (github.com/Niko1221/Strata @ 99f3dbd, Docker image built for sm_120): hot experts cached in the RTX 5090's VRAM, all experts in 60 GB of host RAM, MTP drafting; CONTEXT=131072, VISION=yes, default KV (int8). Harness: pi 0.73.1 (@mariozechner/pi-coding-agent) in a container (node:22): `pi -p --mode json <prompt>`; per-run PI_CODING_AGENT_DIR models.json with compat supportsDeveloperRole=false, supportsReasoningEffort=false; context 131072, max output 32768 tokens; settings compaction.reserveTokens=49152, keepRecentTokens=20000. Orchestrator: github.com/dh7/agent-checkup-benchmark @ 2b7769b; the checkup instructions are passed verbatim as a single prompt, one attempt, no time limit unless noted. Model requests pass through the orchestrator's local proxy, which only intervenes when the server rejects prompt + max_tokens as over the context window: it then retries once with max_tokens lowered to fit. Operator limits: stopped by the operator at 121 min (120-min cap). From 12:19 the agent waited on its own python3 vm.py (compute-vm challenge), which ran for 100 min at full CPU; pi's bash tool has no timeout.