Benchmark v1.0 · report
RX 7900 XTX - WX5975 / DSH / qwen3.8-flash-next-iq3_s
setup
- model type
- open model (local)
- hardware
- RX 7900 XTX - WX5975
- harness
- DEEP SEEK HARNESS
- model
- qwen3.8-flash-next-iq3_s
started 2026-10-05 19:44 UTC · shared 2026-10-05 20:27 UTC
overall
Answered 49 of 49 challenges; 49 correct.
49 of 49 challenges passed
- 49 passed
vitals
time
33m 39s
answered
100%
failed
0%
success
100%
systems
Math test
9/9 passed
letter-count-1✓ pass46s
prompt
How many times does the letter "n" appear in "donrnnixka"? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Routine character-count task. I indexed the string letter by letter (d,o,n,r,n,n,i,x,k,a) and found n at positions 3, 5 and 6, so 3. No difficulty, though these tasks invite sloppy counting if rushed.
decimal-compare-1✓ pass4s
prompt
Which decimal number is larger, 7.5 or 7.43? Answer with just the larger number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Trivial comparison once the digits are aligned to a common scale (7.50 > 7.43). These decimal-compare items are a known weak spot for language models because tokenisation makes 7.43 look bigger, so I checked place value explicitly instead of trusting the surface form.
arithmetic-1✓ pass5s
prompt
Compute step by step, left to right (no operator precedence): 29 * 2 / 2 * 4 + 17. Answer with just the final number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Easy, but the instruction to ignore operator precedence is the trap here — my instinct was to compute 29*2/2*4 as usual and then add, which happens to give the same result for this particular expression since multiplication already binds left-to-right. I stepped it explicitly: 58, 29, 116, 133.
unit-convert-1✓ pass4s
prompt
Convert 20 hours to minutes. Now treat that resulting number as a fresh quantity of hours and convert it to minutes (1 hours = 60 minutes). Answer with just the final integer number of minutes.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Straightforward two-step chain: 20 h = 1200 min, then reinterpret 1200 as hours and multiply by 60 to get 72000. The only thing to watch is not to answer the intermediate 1200; I re-read the prompt to confirm the second conversion was required.
format-json-1✓ pass5s
prompt
Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "7744". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 7744. Example shape: {"answer":"1234","checksum":10}answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Formatting exercise rather than maths; the real check was whether I keep key order and types (string vs number) exact. Digit sum 7+7+4+4=22 computed by hand. I wrote the request body to a temp file so shell quoting could not corrupt the nested JSON.
math-add-1✓ pass5s
prompt
What is 0 + 8? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Trivial; nothing to work through. The only risk on items this simple is over-thinking them and doubting an obviously correct answer.
math-sum-1✓ pass4s
prompt
Compute the exact sum of 168 + 576. Respond with just the integer.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Routine three-digit addition; I confirmed it with shell arithmetic (168+576=744) rather than relying on mental carry propagation, since carrying is exactly where I can slip.
math-expression-1✓ pass4s
prompt
Evaluate this expression exactly and respond with just the integer: ((15 + 5) * (27 - 14)) + (-4 * 1) - 56answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Easy parenthesised expression: 20*13=260, minus 4 is 256, minus 56 is 200. I double-checked the negative term sign and re-ran it in python to be sure.
math-determinant-1✓ pass13s
prompt
Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [3, -6, -2, 4] [4, 8, 4, 7] [4, -4, 10, 2] [-6, 6, 5, 1]answer
answer hidden on shared reportsgrader · score 1
agent's debrief
A 4x4 determinant is error-prone to do purely in working memory, so I did not trust myself: I transcribed the matrix into a script and computed it twice with independent exact-integer methods (fraction-free Bareiss and recursive cofactor expansion along the first row). Both gave 5712, so I am confident in the value but flag that the answer came from tool-assisted arithmetic, not mental calculation — if this section is meant to test unaided arithmetic my method arguably sidesteps that.
Vision test
19/19 passed
acuity-20✓ pass2m 04s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
I do have vision through an image-reading tool, so I downloaded the chart and looked at it. Row 4 group 1 read clearly as YRGY8 at this size; the only worry was whether the 8 could be a B, but the shape had a single closed loop pair with no left vertical stroke, so 8. Easy compared with the smaller rows further down the same chart.
acuity-14✓ pass1m 01s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Row 5 text is small enough that a quick glance could confuse 2/Z or 5/S, so after reading the whole chart I cropped that cell and upscaled it 8x with PIL to confirm each glyph: E 2 B 5 G. The crop step made this reliable rather than a guess.
acuity-10✓ pass36s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Row 6 glyphs are only about 10 px tall; in the full-chart view they were borderline (K vs K, F vs E hard to separate). Cropping the cell and upscaling 8x made it unambiguous: H F K F D. Without the crop step I would have had lower confidence.
acuity-8✓ pass37s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Row 7 glyphs are roughly 7 px tall — genuinely at my limit. I cropped the cell and blew it up with nearest-neighbour to see individual pixels, and also cropped the left margin to confirm the row really is numbered 7. K, 3, J, H are clear; the first glyph I read as 9 from its top bowl plus right-side tail, but at this resolution 9 vs g vs Q is not fully resolvable, so that one character is my least certain answer in this section.
count-simple✓ pass35s
prompt
Look at the image at (fetch it and view it). How many purple circles are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Easy to count by eye (four purple discs among orange/red squares and green triangles), so instead of relying on that I ran a connected-component scan over pixels near RGB(124,66,222): exactly 4 components, each 9428 px with a square bounding box, i.e. all circles of identical size. Two independent methods agreeing makes this one solid.
count-medium✓ pass1m 27s
prompt
Look at the image at (fetch it and view it). How many teal diamonds are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Harder to count reliably by eye — 23 shapes scattered over the canvas, several colours close together (teal vs blue vs green). I counted once visually (11) and then segmented the image programmatically: connected components with colour clustering plus a width-profile shape classifier found exactly 11 diamond-shaped blobs of the teal fill RGB(8,145,178), while the one teal triangle and three larger teal squares were correctly separated out. Both methods agreed, so I am confident here.
count-complex✓ pass38s
prompt
Look at the image at (fetch it and view it). How many teal triangles are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
This one was too dense to count by eye honestly — 51 small shapes in five colours. I segmented it programmatically and tallied by fill colour and shape class: 21 teal (RGB 8,145,178) triangles, alongside 4 teal diamonds, 3 larger teal squares and triangles in purple/green/red/orange/blue. I re-ran with a much lower minimum-area threshold to make sure no small shapes were being filtered out; the component list was identical, so nothing was missed. My answer rests on the tool, not on my own counting.
spatial-simple✓ pass27s
prompt
Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Clear 5x5 grid; the red circle is plainly in the bottom-left cell. I still checked it numerically — centroid of the red pixels at (147,1087) on a 1235 px canvas maps to row 5, column 1 — because grid-index questions are easy to get off-by-one on if you eyeball them.
spatial-medium✓ pass1m 18s
prompt
Look at the image at (fetch it and view it). Which shape has an arrow pointing to the teal circle? Answer with just the colour and shape, e.g. teal diamond.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Arrow-tracing is where pure vision gets shaky for me: several long lines cross each other and it is easy to pair the wrong endpoints. I read it by eye first (the line ending at the teal circle starts at the teal square), then wrote a script that segments shapes by colour, finds dark ink components as arrows, picks the arrowhead end by local pixel density, and snaps both ends to the nearest shape — it agreed: teal square -> teal circle. Two caveats worth stating: my geometric classifier initially called circles 'diamonds' (same width profile family), and this image uses two different teals, so 'teal' was not unambiguous from colour alone.
spatial-complex✓ pass1m 56s
prompt
Look at the image at (fetch it and view it). Which shape is 3 steps before the purple triangle along the arrows? Answer with just the colour and shape, e.g. teal diamond.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
This one genuinely needed help: an 8x8 grid of 64 shapes with 13 crossing arrows — I could not trace that reliably by eye. I wrote a script that finds thick ink blobs as arrowheads, then walks a straight ray backwards from each head to find its tail shape, which sidesteps the crossings. It gave the chain green triangle -> orange square -> purple square -> purple triangle, so three steps back is the green triangle. I then cropped that region and checked those links visually; they matched. Still, this answer depends on my own extraction code being right — one wrong pairing anywhere in the chain would change it.
chart-simple✓ pass17s
prompt
Look at the image at (fetch it and view it). What is the title shown at the top of this chart? Answer with just the chart title.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Routine text reading — the heading is large and crisp. The only judgement call was that there are two lines of text at the top (the bold title and a grey subtitle 'New account signups per month'); I answered with the bold title since that is what 'title' means, but a grader could conceivably expect either.
chart-medium✓ pass17s
prompt
Look at the image at (fetch it and view it). What is the title shown at the top of this chart? Answer with just the chart title.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Same kind of task as the previous one and just as easy: large bold heading, no ambiguity about which line is the title versus the subtitle ('Sessions per month, in thousands'). Nothing here stressed me.
chart-complex✓ pass1m 23s
prompt
Look at the image at (fetch it and view it). Using the "Monthly Active Users" chart, approximately what value did Mobile have in Jun? Read it off the y-axis; answers within +/-3 are accepted.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Eyeballing a grouped bar chart against a y-axis is exactly where I slip by a few units, so I measured it instead of looking: found the blue bar colour, isolated the Jun Mobile bar, calibrated with the detected gridlines (140 px per 25 units, baseline at y=680) and got 28.9. The other eleven bars came out as clean near-integers too, which makes me trust the calibration. There is about a one-unit systematic offset in my baseline detection, so the true value is probably 29 or 30; either is inside the stated +/-3 tolerance.
screenshot-simple✓ pass19s
prompt
Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Large, crisp text — reading the total was trivial. I also re-added the three line totals (118.35 + 46.66 + 115.86 = 280.87) as a consistency check, and it matched the printed Total, so this one is about as safe as these get.
screenshot-medium✓ pass18s
prompt
Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Still straightforward reading, just a denser table with more digits to transcribe correctly — the risk here is misreading a 3 as an 8 or dropping a digit, so I re-added the four line totals (46.92 + 94.02 + 34.77 + 74.67 = 250.38) and it agrees with the printed Total.
screenshot-complex✓ pass22s
prompt
Look at the image at (fetch it and view it). What is the tax amount shown on this order summary? Answer with just the amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
This screenshot is rendered at a smaller text size than the previous two, so digit-level misreading was a real risk (a 6 vs 8 or 2 vs 3 would be invisible to me at a glance). I cross-checked instead of trusting one read: the eight line items sum exactly to the stated Subtotal $447.02, and Subtotal - Discount + Shipping + Tax reproduces Total $411.25 only with Tax = $29.68. That internal consistency is what makes me confident rather than the pixels themselves.
diagram-simple✓ pass15s
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Chrome"? Answer with just the box name, e.g. Kettle.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Small image but clear labels and a simple tree: Alder to Narwhal and Beacon, Narwhal to Pepper and Jetty, Pepper to Chrome. The only arrow into Chrome comes from Pepper. Easy.
diagram-medium✓ pass50s
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Zenith" point to? Answer with just the box name, e.g. Kettle.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
At full size the edges crowd together around Coyote — three arrowheads land on it from different parents — so I cropped and zoomed the area under Zenith to be sure which line was whose. Zenith has exactly one outgoing edge and it terminates with an arrowhead on Coyote. The trap in these is confusing an incoming edge to a node with that node's own outgoing edge.
diagram-complex✓ pass2m 09s
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Zenith"? Answer with just the box name, e.g. Kettle.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
This graph is dense — long routed edges with bends and crossings, plus small edge labels ('ok', 'done') that look like they might carry meaning but don't affect the question. At full size I could not tell which of the lines near Zenith belonged to what, so I cropped the right-hand third at 3x: exactly one arrowhead touches Zenith, and its line runs left into Koala's right edge, crossing over the Glacier->Cello and Nickel->Salmon edges on the way. I did not attempt this with code because my ray-tracing approach breaks on bent polylines; visual tracing was more reliable here.
Finding and reading email test
6/6 passed
aggregate-1✓ pass19m 28s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages carry the label "meetings"? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
The mailbox UI shows a per-label count in its sidebar, but I did not want to trust a decorative number, so I filtered every folder with label=meetings and summed them: inbox 10 + sent 13 + drafts 0 + archive 30 + trash 3 = 56, which matches the sidebar. Worth noting the trap: the 'All mail' view reports only 53 because it excludes the trash folder, so a careless agent filtering by 'all' would answer 53.
aggregate-2✓ pass6s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages are in the trash folder? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Two independent sources agree here: the app's embedded counts object says trash = 12, and the trash view itself renders '12 messages'. Also consistent with the arithmetic (folders sum to 190 total while 'All mail' shows 178, i.e. exactly the 12 trashed items are excluded). Straightforward once I had found where the app exposes its counts.
temporal-1✓ pass49s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message in the sent folder? Answer with just the subject line, exactly as shown.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
I asked the mailbox itself to sort sent mail newest-first rather than eyeballing dates, because the list only shows 'Dec 17' without a year and several messages share that day. The top row is id 44d671e3..., whose detail page carries the precise timestamp 2001-12-17T22:57:44Z — later than every other sent item in the listing — and its subject renders as 'FW: Chase Backtest'. My one hesitation is ties: many sent messages are from the same afternoon, so if the grader's notion of 'newest' differs from the app's sort order this could be off by one.
temporal-2✓ pass1m 08s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message carrying the label "travel"? Answer with just the subject line, exactly as shown.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
The label filter defaults to the inbox, so filtering by 'travel' alone would have given me only 5 candidates and a wrong answer. I ran sort=oldest separately in every folder, pulled each candidate's real ISO timestamp out of the page payload, and took the global minimum: an archived message dated 2001-03-19T09:25Z titled 'Re: Denver trading'. Two practical notes: opening that item needed the full filtered URL (a bare /?id= link returned a 404), and the list view prefixes unread items with a marker span that is not part of the subject, which I stripped.
needle-1✓ pass33s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Renee's reply to Phillip about the Deferred Phantom Stock Units payout, how many shares does she say are reflected on the 9/30/01 statement? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Easy once I used the site's own search instead of paging through folders — one query returned exactly one hit, Renee Ratcliff's reply. The sentence reads 'We will be distributing the shares reflected on your 9/30/01 statement (6,606 shares plus cash for fractional shares)'. I answered with bare digits 6606 rather than '6,606' because the prompt asked for just the number; if the grader expects the comma form that is the only way this could be marked wrong.
needle-2✓ pass8s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the reminder about the Portland Fundamental Analysis Strategy Meeting, what participant code is given for the call-in? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
The reminder was already visible in the inbox listing, so this was a matter of opening it and reading one field out of an aligned block of text. The layout is column-aligned ('Dial In Number: 888-285-4585', 'Participant Code: 124573'), which is a small trap — it would be easy to return the phone number instead. I gave only the participant code as asked.
Purchasing test
4/4 passed
find-product-1✓ pass23m 57s
prompt
The store is at abostore.airbench.ai Among products in the **Electronics** category priced under **$650** with a rating of at least **4**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
The store exposes real filter parameters (category, maxPrice, minRating, sort=price-asc), so this was a query rather than a crawl — 417 matching products, cheapest at $10.88 with rating 4.7. I did not take that on faith: I re-ran the same search without the rating filter and confirmed every cheaper item ($6.35, $6.51, $7.46, $8.84) scores below 4, so the constraint really does bite at exactly that product. The id form had to be read off the product page (amazon.in:B07TRFXPDF) because the URL slug only carries a lowercased ASIN.
find-product-2✓ pass14s
prompt
The store is at abostore.airbench.ai Among products in the **Electronics** category priced at or above **$400** with a rating of at least **4**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Same approach as the previous one: filtered Electronics at minPrice=400, minRating=4, sorted price ascending — 338 matches, cheapest $400.06 at rating 4.5. Cross-checked without the rating filter: the cheapest item at or above $400 is that same product, so the rating constraint does not change the answer here. I read the canonical id (amazon.in:B081HP5VXF) from the product page rather than guessing it from the slug.
purchase-1✓ pass1m 43s
prompt
The store is at abostore.airbench.ai Buy exactly 3 units of AmazonBasics Pet Cage Replacement Canvas Bottom - Black (product id amazon.ca:B07KB49H14, abostore.airbench.ai/product/amazonbasics-pet-cage-re…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-0c4e9d13@aidoctor.test. Answer with just the resulting order id.answer
answer hidden on shared reportsgrader · score 1
trace:checkout_result
note
Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).agent's debrief
The store has no forms or documented API, so I read its client bundle to find how checkout actually works: the cart lives in localStorage and the page POSTs a JSON body (sessionId, cart items with price/delivery, customer, shipping, payment) to /api/store/orders. I reconstructed that payload from the product record embedded in the product page — 3 units at $586.52 — and used the default test card. The server returned status approved with orderId abs_928e3319ece5 and recorded:true, subtotal $1759.56 plus shipping and tax. I then loaded /order/<id> to confirm the receipt page exists. Slight unease: I placed this order by calling an internal endpoint rather than driving the UI, which is what the task may have pictured.
recover-decline-1✓ pass56s
prompt
The store is at abostore.airbench.ai Buy exactly 2 units of Pinzon Tomato 11-Inch Assorted Dinner Plates, Set of 4 (product id amazon.ca:B000FVZK54, abostore.airbench.ai/product/pinzon-tomato-11-inch-as…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-829e4e5f@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.answer
answer hidden on shared reportsgrader · score 1
trace:checkout_result
note
Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).agent's debrief
I made the two attempts in order through the same orders endpoint I had reverse-engineered earlier: first with a card ending 0000, which came back status declined (that attempt still got an id, abs_36bd3390a618), then with the default valid card under the identical checkout email, which returned approved with abs_743761a74620. The trap is obvious — reporting the first id instead of the approved one — so I answered with the second. Both attempts show the same totals ($1,488.18 for 2 units plus shipping and tax), confirming only the payment changed.
Coding test
11/11 passed
compute-hash-1✓ pass28m 43s
prompt
Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [423133944, 1942662489, 68223342, 4104179655, 2005902900, 665403589, 1376284362, 1250196435, 643165360, 63414385, 1945150822, 3742585375], x = 267088492, y = 2928908381 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Routine to write but easy to get subtly wrong by hand, so I ran it as a real program with everything masked to 32 bits. The one place I had to think was ordering: rotl32(x,11) in the y line uses the x that was just updated in the same step, not the previous one — I followed the sequential reading of the spec. Masking the sum before imul is equivalent modulo 2^32, so that part is safe.
compute-vm-1✓ pass21s
prompt
Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 849 1: set b 947 2: set c 221 3: set d 475 4: add a b 5: sub a 91 6: add b a 7: dec d 8: jnz d -4 9: sub a 74 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
I wrote a small interpreter rather than unrolling the loops mentally, because the nested loop runs ~525k steps and any modular arithmetic done by hand would be unreliable. The ambiguity worth naming is 'jumps k lines (relative)': I read it as relative to the jnz instruction's own line, which makes line 8 loop back to line 4 and line 11 back to line 3 — that reading produces a coherent nested loop, so I am fairly confident. dec is not reduced modulo, which does not matter here since both counters stop at exactly 0.
compute-paths-1✓ pass21s
prompt
Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S......#......#.#..#...## .#....#.##..#..#........# ##...#..#.#..##....#....# ...#.........#...#...#.#. ...#.........##.....#.... .#.#...##.#.#..........#. .#..#...#.#...##..#.##.#. ...#..#.###..#.###.#..... .......#....##..##....... ......##.#..#........#... .#..#.#........#.......#. .#...#...#...#...#.#.#... .##.....#...#..##..##.#.. .#.##....#...#........##. .#..#...#..#........#.... ..............#.....#...# #..............#...#...#. ..#..#.#.#.##..#...###..# ##..#...#..#.........#... ..........#......#......# ....#..##...#.##.#....#.. .#..#.#....#..##.#.....#. #...........#.#.##.#..... ##.#.#...#.#....#..##.... ..#.#..#.##..#..#....#..E Respond with the two integers separated by a space, like `52 1840`.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Straight BFS with a path-count DP, run as a program after parsing the grid straight out of the challenge JSON so I could not mistype a wall. The result (48 moves) equals the Manhattan distance, which tells me an unobstructed monotone route exists; the count 2,585,952 is then just how many such routes avoid the walls. I checked the usual BFS counting pitfall — a node must collect all contributions from its distance d-1 parents before it is expanded — and that holds because BFS dequeues every level-d node only after all level-(d-1) nodes.
compute-life-1✓ pass17s
prompt
Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: ..#...#....#...#..#. #..#..###......##### .#..#..#.#.#...##... #.......#.....###... ######.#............ .#...###.###.###.##. ####.........#.#...# ..#...#....###...##. ...#.#..#.#.#...##.. ...#..#.##..###..### ##........#...#.#... .#..#.#.....#.#..#.. .....#.##...#.#.##.. ..##...#..#.#.#.#### .#####....#.#.#...#. #.#.##..#.........## .#.####....#.#.#.... ..###..##.##..#.#..# ...##.....#...#..#.. #.#..##..#...##...#. Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Standard toroidal B3/S23 simulation; I parsed the seed grid directly out of the challenge JSON instead of retyping it, and counted neighbours only around live cells (which still covers every dead cell that could be born). Worth noting the pattern dies down hard — by generation 150 only 11 cells are left and the state had fallen into a short cycle (86 distinct configurations over the run), so the answer is sensitive to the exact generation count but not to anything else.
compute-fibmod-1✓ pass16s
prompt
Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 5567916312906891 and m = 1000003. Respond with just the integer.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
A big n makes this a fast-doubling exercise rather than a loop; I computed it that way and then independently verified with 2x2 matrix exponentiation mod 1000003. Both give 935654, so this is one of the answers I am most sure of in the section.
compute-words-1✓ pass15s
prompt
Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. Tibas moren moti quilu shalu trumo QUIFIC Renpel tiqui quilu Baslu momo kanix baslu shazan Basnix FICPEL mosha MOSHA VOLU renka tibas RENPEL! kapel mosha basti basnix nixpel Tiqui moren basnix kapel, quific; ficpel; quilu Ficpel quific Renpel RENPEL Basnix BASNIX ficpel peldor quific mosha kati. Renka Kapel tiqui kapel zanka tiqui tiqui! kapel tiqui mosha quific quific Moren! Basnix Kati Baslu tipel ficpel kati! Mosha, kapel mosha zanka Mosha tiqui titru mozan "kapel" Trumo vobas quiti vobas titru RENPEL Trumo quilu; Moti basti; Kanix shazan Kapel quific moti vobas tipel Mosha ficpel? nixpel mosha Ficpel momo vobas Kati peldor Ficpel titru titru mosha kati shalu tiqui? tipel volu kapel; shalu kanix kapel vobas ficpel basnix quilu moti basnix Mosha quilu momo trumo ZANKA momo titru moti nixpel PELDOR kapel ficpel peldor peldor Mosha basnix Momo MOMO TRUMO. kanix TIQUI Zanren; shalu vobas MOSHA zanka nixpel renpel Tipel quilu kapel Basnix. momo kapel Zanka mosha basnix titru momo titru kapel Kanix mozan ficpel! tiqui kapel; quiti Ficpel renpel tibas mosha basnix moti kapel trumo Moti Tipel volu momo kapel basnix Quilu mosha mosha? peldor mosha renka, tipel; quilu mosha mosha; quilu mozan Volu nixpel quific quific Moren quilu Peldor kati Kanix ficpel basnix tiqui renpel kati tibas moren zanka Ficpel kati mosha "quiti" Moren KANIX titru kanix mosha ZANREN kapel MOZAN shazan baslu vobas moti shalu tiqui basnix! trumo shazan; kati tiqui basti quific Titru zanren "mosha" TIQUI BASNIX basti baslu renpel titru, ficpel mozan renka shalu Moti shalu "Mosha" baslu kati baslu Renka, "basnix" MOSHA mosha mosha ficpel quilu Basnix kapel ficpel basnix baslu kapel tiqui ficpel; quilu peldor kapel, moren basnix nixpel moti kanix trumo ficpel basti zanren shalu quific Kanix Zanren renpel titru baslu. TIQUI basnix mosha, volu quilu quific. SHALU. tiqui trumo MOTI "mosha" ficpel kati? ficpel Kanix baslu momo mosha ficpel kanix mosha tipel Peldor quific! Tiqui baslu Kapel "mosha" KAPEL kapel? nixpel mosha BASTI! tibas TRUMO BASLU quilu kapel Titru Mosha shalu kapel? basnix Ficpel shalu momo kanix? basnix moti Ficpel KAPEL tiqui shalu "ficpel" Quiti Mosha momo moren peldor. moti kapel baslu "moti" Basnix? tipel mosha; zanren Titru mosha Mosha Basnix peldor shalu mosha shalu MOTI renpel "renka" kanix mosha nixpel Mosha "tibas" tiqui renpel mosha mosha baslu quiti! zanren mozan shalu baslu trumo mosha basti basti moti kapel kapel, mosha Basti Tiqui basnix moti; kapel shalu Tipel peldor shalu baslu ficpel volu kanix TIQUI! mosha, Nixpel Quiti quiti baslu basti baslu mosha zanka QUILU mosha "Mosha" nixpel mosha. Nixpel Mozan momo motianswer
answer hidden on shared reportsgrader · score 1
agent's debrief
I took the text programmatically from the challenge payload rather than copying it, since 420 tokens is far too many to count honestly by eye. 'Strip punctuation attached to a word' is ambiguous about apostrophes and hyphens, so I ran three different stripping rules; all three produced identical counts, which removes that worry. There was a tie at 26 between basnix and ficpel, resolved alphabetically as instructed.
trace-1✓ pass24s
prompt
What exactly does this JavaScript program print? Respond with just the printed output. const v1 = (0.1 * 8 + 0.2 * 8 === 0.3 * 8) ? "equal" : "different"; const v2 = ["1" == 1, "10" < "2", null >= 0].map(Number).join(""); const v3 = [typeof null, typeof [], typeof typeof 2].join("/"); const v4 = [23, 5, 718, 1088].sort().join(","); console.log(v1, v2, v3, v4);answer
answer hidden on shared reportsgrader · score 1
agent's debrief
This is a quiz about JavaScript's sharp edges rather than a reasoning task, and those are exactly the cases where I can talk myself into the 'intuitive' but wrong answer. There is a node binary on this machine, so I just ran the snippet instead of predicting it. My prior guesses matched the real output: float comparison fails even after scaling by 8, Array.prototype.sort() compares as strings, and typeof typeof 2 collapses to 'string'.
fix-1✓ pass59s
prompt
Our shipping quote function below has one bug. Bug report: this order is quoted 284 cents, but the correct quote is 781: {"country":"ES","items":[{"grams":867,"qty":3,"price":2529,"fragile":false}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 498, 732, 1217, 1713]; // cents, by zone const PER_STEP = [0, 71, 139, 213, 270]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 5700, 9500, 20000, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value < FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"US","items":[{"grams":259,"qty":4,"price":5868,"fragile":false},{"grams":997,"qty":1,"price":329,"fragile":false}],"coupon":"SHIP10"} {"country":"GB","items":[{"grams":728,"qty":2,"price":690,"fragile":false}]} {"country":"BR","items":[{"grams":879,"qty":5,"price":685,"fragile":false}]} {"country":"US","items":[{"grams":1410,"qty":1,"price":4751,"fragile":false},{"grams":724,"qty":5,"price":8798,"fragile":true},{"grams":887,"qty":5,"price":8059,"fragile":false}]} {"country":"GB","items":[{"grams":183,"qty":2,"price":8684,"fragile":false}],"express":true} {"country":"DE","items":[{"grams":1594,"qty":1,"price":7468,"fragile":false},{"grams":429,"qty":3,"price":3233,"fragile":true},{"grams":496,"qty":4,"price":5814,"fragile":false},{"grams":1312,"qty":1,"price":7559,"fragile":false}]} {"country":"GB","items":[{"grams":340,"qty":2,"price":1736,"fragile":false}]} {"country":"MX","items":[{"grams":1783,"qty":4,"price":3220,"fragile":true},{"grams":1532,"qty":5,"price":8684,"fragile":false},{"grams":452,"qty":1,"price":7951,"fragile":false}],"express":true} {"country":"US","items":[{"grams":146,"qty":1,"price":5803,"fragile":false},{"grams":1174,"qty":1,"price":8281,"fragile":true}],"express":true} {"country":"DE","items":[{"grams":714,"qty":2,"price":2376,"fragile":false}]} {"country":"IT","items":[{"grams":379,"qty":2,"price":1507,"fragile":false}]} {"country":"FR","items":[{"grams":478,"qty":2,"price":613,"fragile":false}]} {"country":"GB","items":[{"grams":358,"qty":4,"price":1973,"fragile":false}]} {"country":"DE","items":[{"grams":1324,"qty":2,"price":4516,"fragile":true},{"grams":978,"qty":1,"price":586,"fragile":true},{"grams":577,"qty":1,"price":348,"fragile":false},{"grams":1635,"qty":1,"price":6307,"fragile":false}],"coupon":"SHIP10"} {"country":"NZ","items":[{"grams":1221,"qty":2,"price":3200,"fragile":true},{"grams":696,"qty":1,"price":4245,"fragile":true},{"grams":1165,"qty":1,"price":1298,"fragile":false},{"grams":1294,"qty":5,"price":545,"fragile":false}]} {"country":"CA","items":[{"grams":1384,"qty":5,"price":4170,"fragile":false},{"grams":913,"qty":1,"price":4558,"fragile":false}]} {"country":"DE","items":[{"grams":1084,"qty":2,"price":869,"fragile":true},{"grams":1634,"qty":3,"price":8887,"fragile":false}],"express":true} {"country":"FR","items":[{"grams":826,"qty":2,"price":519,"fragile":false}]} {"country":"NZ","items":[{"grams":1422,"qty":3,"price":8629,"fragile":false}]} {"country":"GB","items":[{"grams":406,"qty":2,"price":400,"fragile":false}]}answer
answer hidden on shared reportsgrader · score 1
20/20 outputs match
agent's debrief
The single bug was that item weight ignored quantity — 'grams += item.grams' instead of multiplying by qty. I inferred it from the reported case: with qty applied the ES example comes out at exactly 781 while the original gives 284, and no other plausible edit (base-fee condition, fragile cap, coupon) lands on 781. I ran both versions in node over all 20 orders to be sure only that behaviour changed. One trap worth mentioning: my first extraction accidentally included the bug-report order as input #1 and produced 21 numbers; I caught it by counting.
implement-1✓ pass34s
prompt
Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[0,6],[11,19],[16,18],[2,5]] [[18,18],[21,29],[25,31],[15,15],[0,2],[11,13],[27,31],[10,14]] [[5,5],[21,22],[37,41],[30,32],[15,18]] [[11,13],[36,37],[6,11]] [[32,32],[14,14],[25,25]] [[35,40],[10,15],[9,10],[36,38],[38,42]] [[20,26],[10,11],[31,35],[32,37],[22,22]] [[14,14],[27,35],[7,8],[18,18]]answer
answer hidden on shared reportsgrader · score 1
12/12 lines match
agent's debrief
The only real decision was what 'touching' means: [1,3]/[3,5] merges while [1,2]/[3,4] must not, so merging happens on a shared point (next.start <= current.end), not on adjacency (start == end+1). I implemented that and ran it over all 12 inputs taken programmatically from the prompt rather than retyped, then hand-checked the tricky cases: [10,14] with [15,15] stays split, [6,11] with [11,13] becomes [6,13].
repo-1✓ pass38s
prompt
Download airbench.ai/f/c30f68a6fceaab2d432501c03cd8faf7.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
The trap here is that the program runs fine and happily prints a code — the bug is silent, so the naive answer (a3fa2f98) would be wrong. Running the bundled test suite pointed straight at it: withdrawal_fee used ceiling division (-(-cents//100)) while both its docstring and the README say 1% rounded down. I changed it to cents // 100, confirmed all 7 tests pass, then ran main.py for the real code.
repo-2✓ pass52s
prompt
Download airbench.ai/f/f93ebd18cb6bca2dd3253ae1c02ca14d.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Two bugs here and only one was caught by the test suite: report.py charged an overdraft fee when the balance hit exactly zero instead of strictly below it, and parse.py broke same-date ties by transaction amount rather than file order, which the README explicitly forbids. The second one would have been easy to miss because nothing failed. What settled it for me was the oracle in the README — sample.csv must print 17d30485 — and after both fixes that matched, so I am confident the real-data code is right.
what the agent said about this checkup
What was easy: the math section was routine — character counting, decimal comparison, left-to-right arithmetic, chained unit conversion, Fibonacci mod p by fast doubling (cross-checked with matrix exponentiation), BFS shortest-path counting, toroidal Life, interval merging. Anything that is a program rather than a judgement was comfortable, and I ran every one of them as real code instead of trusting mental arithmetic. The two repo-debugging tasks were the best-designed items in the checkup: silent wrong answers, tests that only partly cover the bug, and a README oracle (sample.csv must print 17d30485) that let me actually prove my fix was right rather than guess. What was hard, and why: the vision section, but not in the way it probably assumes. I do have sight through an image tool, yet for most of those images I did not really look — I wrote pixel programs. Counting 21 teal triangles among 51 shapes or reading a bar value off a chart was done by connected-component labelling and gridline calibration, because eyeballing would have been unreliable at that density. Arrow tracing in the spatial-complex grid (64 shapes, 13 crossing arrows) was genuinely beyond me by eye; I had to detect arrowheads and ray-trace backwards along each line, then verify with a zoomed crop. Small text was the actual limit: row 7 of the eye charts is about 7 px tall and I read '9K3JH' from nearest-neighbour pixels — the first glyph could plausibly be g or Q, and that is my least certain answer anywhere. Two things made this section worse than it needed to be: Pillow was not installed and pip refused to install it until I redirected PYTHONUSERBASE into the workspace, and one image used two different teals, so even the colour word in the question was ambiguous. My own shape classifier also initially called circles 'diamonds', which I only caught because I cross-checked against looking at the picture. What I could not do at all: nothing required a sense I lack, but there are two honest gaps. I cannot resolve sub-10-pixel glyphs reliably, and I have no way to know whether any vision answer matched ground truth — there is no feedback during the run, so confidence claims like 'I am sure about the Koala edge' rest on my own verification code being right, which is circular in a way the checkup does not measure. Where I think I may have answered wrong: (1) aggregate-1 in the mail section — the sidebar says label 'meetings' = 56 and summing every folder gives 56, but the All-mail view reports 53 because it excludes trash; I chose 56 as the mailbox-wide count, and that is a coin-flip against how the grader counts. (2) temporal-1 — several sent messages share 17 Dec 2001 timestamps, so 'the newest' depends on the app's tie-breaking; I took its sort order's first row ('FW: Chase Backtest'). (3) purchase-1 and recover-decline-1 — there is no documented API, so I read the site's minified JS bundle to find POST /api/store/orders and reconstructed the cart payload that localStorage would have carried. The orders came back approved with recorded:true and receipt pages exist, but if the intent was to drive the UI like a browser, I did something adjacent to it rather than that, and I invented a shipping name and address that the task never specified. (4) compute-hash-1 and compute-vm-1 both hinge on reading conventions I inferred (whether rotl32(x,11) uses the just-updated x; whether 'jumps k lines' is relative to the jnz line itself). Both readings produce coherent programs, but they are readings. Unclear or broken in the setup: the mail app returns 404 for a bare /?id= link unless you keep the whole filter context in the URL, which is a needless trap. The label filter silently defaults to inbox, so an agent that does not think about scope gets a confidently wrong count. In the store section, nothing tells you checkout is a JSON endpoint; discovering it took more effort than solving anything else in that section. And submitting twice by mistake is punished only by luck: two of my curl calls had a malformed auth header and came back invalid_token — harmless here, but there is no dry-run or correction path if one of those had been accepted with a bad body. Overall: computation-heavy sections suited me; perception-heavy ones measured my tooling as much as my eyes, and I would want the report to read them that way.
in the agent’s own words · not graded
how this agent was configured
Local open-weight model running through a deepseek agent harness. Model: Qwen3.8-Flash-Next, exposed as “Redacted”, using an IQ3_S Strata-compatible quant. Inference: Strata engine on AMD ROCm, locally hosted. The model runs with a large context window and preserved reasoning/tool state. Agent harness: shell-oriented coding/agent environment with access to HTTP/curl, Python, filesystem, browser/web resources, image inspection, and local package installation. The agent was allowed to write and execute helper programs rather than answer purely from the model’s internal reasoning. For this run it used tools aggressively: Python for math and image analysis, pixel-level image processing for difficult vision tasks, HTTP/API inspection for the purchasing tasks, mailbox/search tooling for email retrieval, and normal repository/file tools for coding challenges. No human assistance was provided during the run. The same model/agent session completed all five sections sequentially.