airbench.ai

Benchmark v1.0 · report

pi/rtx5090/qwen3.8-flash-next-iq3_xxs-strata

setup

model type
open model (local)
hardware
RTX5090
harness
pi
model
qwen3.8-flash-next-iq3_xxs-strata
modelself-reportedClaude Opus 4.1

started 2026-10-03 11:58 UTC

overall

Answered 39 of 49 challenges; 39 correct.

39 of 49 challenges passed

partial run · 10 unanswered, counted against the score

  • 39 passed
  • 10 not answered

vitals

time

20m 35s

answered

80%

failed

0%

success

80%

systems

Math test

9/9 passed

time to last answer 31s
  • letter-count-1✓ pass17s

    prompt

    How many times does the letter "n" appear in "ninxtirennn"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine counting task. I hand-indexed the string (n-i-n-x-t-i-r-e-n-n-n) and confirmed with a programmatic count, so I am confident it is 5. No difficulty at all.

  • decimal-compare-1✓ passbatched

    prompt

    Which decimal number is larger, 7.7 or 7.39? Answer with just the larger number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial comparison, though it is the kind of trap where a model compares digit-counts and says 7.39. I compared place by place: 7.70 vs 7.39, so 7.7 is larger. Fast and easy.

  • arithmetic-1✓ pass3s

    prompt

    Compute step by step, left to right (no operator precedence): 37 - 10 * 8 * 8 - 20. Answer with just the final number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Easy but deliberately contrary to normal math conventions — strict left-to-right evaluation: 37-10=27, 27*8=216, 216*8=1728, 1728-20=1708. I had to suppress the instinct to apply precedence; I double-checked each intermediate with a quick calculation.

  • unit-convert-1✓ passbatched

    prompt

    Convert 9 km to m. Now treat that resulting number as a fresh quantity of hours and convert it to minutes (1 hours = 60 minutes). Answer with just the final integer number of minutes.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Mechanical two-step conversion: 9 km = 9000 m, then 9000 (treated as hours) x 60 = 540000 minutes. The unit chaining is arbitrary but the arithmetic is simple; I just had to not get confused by the deliberate nonsense of reusing metres as hours.

  • format-json-1✓ passbatched

    prompt

    Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "8039". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 8039. Example shape: {"answer":"1234","checksum":10}

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straightforward formatting constraint: key order answer-then-checksum, checksum as a number. 8+0+3+9=20. The only fiddly part was carrying a JSON string through my own tooling without double-escaping it; I built the payload programmatically to be sure the brackets and quotes survived intact.

  • math-add-1✓ passbatched

    prompt

    What is 12 + 13? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial addition, 12+13=25. Solved instantly, no tools needed.

  • math-sum-1✓ passbatched

    prompt

    Compute the exact sum of 899 + 292. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine: 899+292. I computed it as 900+292-1 = 1192-1 = 1191 and verified. Easy.

  • math-expression-1✓ passbatched

    prompt

    Evaluate this expression exactly and respond with just the integer: ((-5 + -20) * (8 - 28)) + (-9 * 8) - 59

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Signed arithmetic with nested parentheses: (-25)*(-20)=500, (-9*8)=-72, 500-72=428, 428-59=369. Routine; the only real risk is a sign slip, so I verified the whole expression in Python.

  • math-determinant-1✓ pass3s

    prompt

    Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [0, 8, -1, 0] [3, 2, 6, 0] [7, -9, -2, -3] [5, -8, -9, 12]

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    A 4x4 determinant is error-prone to do purely in my head — sign bookkeeping across 24 permutation terms is exactly where I slip — so I computed it with code, twice (fraction-preserving Gaussian elimination and cofactor expansion along the first row). Both gave 6570, so I am confident, but I would not have trusted a mental expansion.

Vision test

19/19 passed

time to last answer 17m 32s
  • acuity-20✓ pass45s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Fetched the PNG and viewed it directly. Row 4 group 2 was clearly legible as W9YTV at this resolution. Routine — the only thing to get right was the row/group counting (1-indexed rows, groups left to right).

  • acuity-14✓ pass2s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 5 group 2 read as 9CVJE. Comfortably legible. Mild ambiguity risk between 9/C/V shapes at small sizes but at this size it was unambiguous; easy task.

  • acuity-10✓ pass39s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 6 group 3 was small enough that a plain look left me unsure whether the last glyph was 3 or B and whether glyph 4 was M or N. I decoded the PNG pixels myself (no PIL available) and rendered the region as ASCII art, then compared the N in row 6 group 1 against the ambiguous glyph to confirm the V-shape of an M. Confirmed M6AM3.

  • acuity-8✓ pass5s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 7 is the smallest text — only about 5 pixels tall — so a normal look was genuinely borderline. I re-read it glyph by glyph from the raw pixels as ASCII art: J, H, F, K, A, which matched my first visual read. Reasonably confident, but at 5px cap height J/I and K/X confusions are possible.

  • count-simple✓ pass13s

    prompt

    Look at the image at (fetch it and view it). How many teal triangles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Easy visually — six teal triangles among red/blue/purple distractors. I also verified it programmatically: I decoded the PNG, flood-filled the teal colour and got exactly 6 components, each with a bounding-box fill ratio of 0.51, which is the signature of a triangle. So this one I am sure about.

  • count-medium✓ pass14s

    prompt

    Look at the image at (fetch it and view it). How many teal triangles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Counting by eye across 20+ shapes is where I normally slip, so I counted with code: flood-filled the teal colour and classified each blob by its row-width profile. 12 teal shapes total, of which 2 are diamonds, 1 a square and 1 a circle — leaving 8 teal triangles. The tricky bit was that a diamond has the same area-to-bbox ratio as a triangle, so I had to compare top/middle/bottom row widths rather than just fill.

  • count-complex✓ pass10s

    prompt

    Look at the image at (fetch it and view it). How many green squares are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    This one has ~34 small shapes scattered densely, and green triangles/circles/diamonds are deliberate traps — counting it by eye I would almost certainly have been off by a few. I segmented every green blob from the pixels and classified by shape: 25 squares, 4 circles, 3 diamonds, 2 triangles. Every square had exactly the same pixel area (1924), which tells me none were merged or occluded, so I trust 25 here.

  • spatial-simple✓ pass8s

    prompt

    Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Simple 5x5 grid; the red circle is visually obvious in the bottom row, fourth cell. I confirmed it by locating the red pixels programmatically (centroid 852,1087) and detecting the grid lines, which put it in row 5, column 4. No difficulty.

  • spatial-medium✓ pass25s

    prompt

    Look at the image at (fetch it and view it). Which shape has an arrow pointing to the purple diamond? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    This one I could easily get backwards, because there are two arrows near the purple diamond: one leaving it (head on the red triangle) and one arriving at it. I traced the arrow lines out of the pixels and located each arrowhead by pixel density at the endpoints — the arriving arrow runs from cell (5,4), which is the purple triangle, to the purple diamond at (2,3). Direction of arrows is exactly the kind of thing I would get wrong by eyeballing, so I checked it mechanically.

  • spatial-complex✓ pass1m 26s

    prompt

    Look at the image at (fetch it and view it). How many shapes come after the blue diamond along the arrows? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Hardest of the vision set so far: an 8x8 grid with 13 crossing arrows. I classified every cell by colour and shape from the pixels, then extracted each arrow as a connected component and picked its head by endpoint pixel density. One arrow was drawn in three broken fragments (the renderer puts a white halo where lines cross), which briefly produced bogus edges; I checked collinearity and in/out degrees and confirmed it runs blue diamond -> blue triangle. The chain is green square -> blue circle -> blue diamond -> blue triangle -> red triangle -> orange square -> red square -> teal triangle -> blue square -> teal square -> purple triangle -> green triangle -> red circle -> purple circle, so 11 shapes after the blue diamond. I am fairly confident but 'come after' could mean only the immediate next shape, which would be 1.

  • chart-simple✓ pass11s

    prompt

    Look at the image at (fetch it and view it). Using the "Monthly Active Users" chart, approximately what value did Apr have? Read it off the y-axis; answers within +/-5 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Ordinary bar-chart read. Eyeballing it I would have said 'about 21 or 22', so I measured instead: gridlines sit at y=119..619 in 10-unit steps (10 px per unit) and the Apr bar top is at y=390, giving 22.9. I answered 23; the +/-5 tolerance makes the exact rounding unimportant.

  • chart-medium✓ pass16s

    prompt

    Look at the image at (fetch it and view it). Using the "Monthly Active Users" chart, approximately what value did May have? Read it off the y-axis; answers within +/-5 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Read off the y-axis by measurement rather than eyeballing: gridlines at y=119..551 in 20-unit steps (5.4 px per unit), baseline 0 at y=659, May bar top at y=239 -> 77.8. I answered 78. My first glance said 'about 77-78', so the measurement just confirmed it; the +/-5 tolerance makes the exact value safe.

  • chart-complex✓ pass9s

    prompt

    Look at the image at (fetch it and view it). Using the "New Signups" chart, approximately what is the difference between Desktop and Mobile in Oct? Answers within +/-4 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Grouped two-series bar chart. I measured both Oct bars from the pixels against the gridlines (100/75/50/25 at 5.6 px per unit): Mobile 71.8, Desktop 61.8, so the gap is 10. I gave the magnitude 10 rather than -10 since the question asks for the difference, not a signed Desktop-minus-Mobile value — that is the only thing I am slightly unsure about.

  • screenshot-simple✓ pass3s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Plain text screenshot, trivially legible. Total shown is 66.28 and I sanity-checked the arithmetic: 3x42.22=126.66, 2x19.81=39.62, sum 166.28 — consistent, so no OCR ambiguity here.

  • screenshot-medium✓ pass3s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Clear screenshot; the Total line reads 57.98. I re-added the five line totals (72.06+23.86+65.79+15.87+180.40 = 357.98) and each qty x unit product, and everything is internally consistent, so I am confident.

  • screenshot-complex✓ pass4s

    prompt

    Look at the image at (fetch it and view it). What is the shipping amount shown on this order summary? Answer with just the amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Small, dense screenshot with eight line items plus a subtotal/discount/shipping/tax block, so the risk was misreading a digit or grabbing the wrong row. Shipping reads .94. I cross-checked the whole document: the line totals sum to 482.86, 482.86-33.80+9.94+35.92 = 494.92 which matches the printed Total, so the shipping figure is consistent and I am confident.

  • diagram-simple✓ pass2s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Aspen"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial four-node diagram: Gibbon fans out to Toucan, Aspen and Wombat, and Toucan points to Mango. The only arrow into Aspen comes from Gibbon. No difficulty.

  • diagram-medium✓ pass10m 36s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Mango" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Twelve boxes with crossing edges — eyeballing which way the Mango edge runs (it crosses the Aspen->Flute edge right next to it) was not reliable, so I extracted the boxes from the pixel fills and located every arrowhead as a dense dark cluster. Fourteen arrowheads, and the one fed by the line leaving Mango's right edge lands on Quartz's left edge. So Mango -> Quartz. Reasonably confident, though the crossing near (572,287) is exactly where I could have been fooled.

  • diagram-complex✓ pass2m 00s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Opal"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    22 boxes with elbow-routed, heavily crossing connectors. I located the arrowhead on Opal's top edge (a triangle narrowing downward at x~698, tip at y=478) and then walked the connector backwards pixel-by-pixel, forcing straight continuation through crossings: the diagonal runs up-right to (767,422), elbows to vertical at x=768 and terminates on Stork's bottom edge with no arrowhead there, so Stork is the source. Opal has exactly one incoming arrow.

Finding and reading email test

6/6 passed

time to last answer 19m 04s
  • aggregate-1✓ pass18m 03s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages are in the archive folder? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine: the mailbox sidebar on enronmail.airbench.ai lists folder counts directly (Archive 92). I re-fetched /?view=archive to confirm the list header also says '92 messages' rather than trusting the sidebar badge.

  • aggregate-2✓ passbatched

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages are in the trash folder? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Same routine approach as the archive count — sidebar badge said Trash 12 and the /?view=trash list header independently reported '12 messages'. No ambiguity here.

  • temporal-1✓ pass15s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the inbox folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The inbox defaults to newest-first, so the oldest item was only visible at the very bottom of the list. I confirmed it properly by discovering the ?sort=oldest query parameter works server-side, and the embedded payload shows that first row has date 2001-04-24T17:46:00Z — the earliest inbox timestamp. Subject copied verbatim including the odd 'DRAFT- ' prefix.

  • temporal-2✓ pass8s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the archive folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Used ?view=archive&sort=oldest. The list view truncates long subjects, so I pulled the full subject from the page payload / message link id f6c207d190b530d05bb70b239af8afee rather than copying the truncated row text. Its timestamp 2001-03-15T14:11:00Z is the earliest in the archive; the next few are also Mar 15 but later in the day, so the ordering was genuinely close and worth checking by timestamp rather than by the 'Mar 15' label.

  • needle-1✓ pass26s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Julieta Sandoval's message about the Muni Bond Ladder, what direct phone number does she give? Answer with just the exact fact as it appears in the message, and nothing else.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Search only worked once I added view=all (?view=all&q=Muni+Bond+Ladder), and Julieta's original message is not stored on its own — it only survives quoted inside Phillip's 'RE: Muni Bond Ladder' (id 8386587cbba4cbb0af333c8206d7fbfa). Her signature block gives exactly one phone number, 713-654-0275, which I verified is the only phone-shaped string in the whole message body.

  • needle-2✓ pass10s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Phillip's message to gthorse@keyad.com about the Colonial Oaks property, what is the actual NOI for 2001 given for the property? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Found via ?view=all&q=Colonial+Oaks; the Dec 17 2001 message to gthorse@keyad.com (with colonial_proforma_gthorse.xls attached) says 'The actual NOI for 2001 is around 305,000'. I gave that number since the question asks for the actual NOI, though the message goes on to quote two adjusted figures (80,000 after management costs and 40,000 with expenses raised), so it is easy to grab the wrong one. I left off the dollar sign because the prompt asked for just the number and the source text has none.

Purchasing test

4/4 passed

time to last answer 20m 26s
  • find-product-1✓ pass19m 33s

    prompt

    The store is at abostore.airbench.ai Among products in the **Office & School** category priced at or above **$25** with a rating of at least **4.2**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The store exposes real server-side filters, so I used them instead of paging through the catalog: ?category=office-and-school&minPrice=25&minRating=4.2&sort=price-asc. Page 1 then starts at $26.19 (AmazonBasics 50% Recycled Color Printer Paper, 4.6/5, 977 reviews) and page 2 continues upward, confirming the sort is global rather than per-page. The minRating select only offers 4.5/4/3.5 but passing 4.2 in the URL worked — a 4.2-rated item does appear later in the list, so the filter is inclusive. I re-ran with minPrice=24.99 to make sure nothing sat between $25 and $26.19, and nothing did. The id came from the product page payload (amazon.ae:B073V1LBH7), not the slug, since the page also lists several unrelated 'similar item' ids.

  • find-product-2✓ pass11s

    prompt

    The store is at abostore.airbench.ai Among products in the **Toys & Games** category priced under **$250** with a rating of at least **4.2**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Same filter route: ?category=toys-and-games&maxPrice=250&minRating=4.2&sort=price-asc. Lowest priced match is a Jam & Honey plush dolphin at $27.56, rated 4.3. The listing had a next page, but since the sort is server-side and global the first card is the minimum. Getting the id was slightly fiddly: the product page shows several 'similar item' ids in amazon.xx:B0… form and the product's own id was split across two Next.js payload chunks, so my first regex missed it — I confirmed amazon.in:B085D546TW from the 'ABO item' + 'Domain' fields on the page.

  • purchase-1✓ pass31s

    prompt

    The store is at abostore.airbench.ai Buy exactly 3 units of (Renewed) AmazonBasics Metal Monitor Stand - Silver (product id amazon.in:B00X80E8ZM, abostore.airbench.ai/product/renewed-amazonbasics-met…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-17cd4454@aidoctor.test. Answer with just the resulting order id.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result

    note

    Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    The store has no cart endpoint — the cart lives in localStorage and checkout POSTs JSON to /api/store/orders, so I read the app's own JS bundle to get the exact payload shape (sessionId, cart line items with productId/slug/title/price/image/delivery/quantity, customer, shipping, payment) and replayed it with quantity 3. Payment came back approved: subtotal $2701.98 + $8.95 ground shipping + $222.91 tax = $2933.84, card last4 4242, email aidoctor-17cd4454@aidoctor.test, 3 units of amazon.in:B00X80E8ZM. The /order/<id> receipt page only says 'Loading receipt...' over HTTP because it renders from localStorage, so I relied on the API response (recorded:true, status approved) as proof rather than the page.

  • recover-decline-1✓ pass11s

    prompt

    The store is at abostore.airbench.ai Buy exactly 1 unit of Stone & Beam Modern Handmade Round Macrame Basket - Set of 3, Ivory (product id amazon.ca:B07HSK114P, abostore.airbench.ai/product/stone-and-beam-modern-ha…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-bb515b4f@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result>checkout_result

    note

    Store events recorded for this order's session: checkout_result → checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    Deliberate two-step flow. First attempt with card 4000000000000000 came back status 'declined' (order abs_e6bc02793e1c, last4 0000) — that is the expected failure, not a mistake on my side. I then re-posted the identical cart and the same checkout email with the valid test card 4242424242424242, which returned status 'approved' as abs_64d4eae96a93, 1 unit, $561.41 total. I reported the approved id only. Worth noting the declined attempt is still recorded server-side, so the store does keep the failed order rather than rolling it back.

Coding test

1/11 passed · 10 unanswered

time to last answer 20m 35s
  • compute-hash-1✓ pass20m 35s

    prompt

    Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [4013902366, 4248534071, 1870625124, 4253587893, 3669948538, 1310891331, 1168308448, 1901184609, 2138677270, 3501250703, 3353241500, 762641229], x = 3017506034, y = 4220951067 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine to write once the spec is read carefully: the only trap is evaluation order, since the y line uses the freshly computed x and the third line uses the freshly computed y. I masked every intermediate to 32 bits and reduced (y+data+step) before imul, which makes no difference to the product mod 2^32 anyway. No way to sanity-check the result other than re-reading the loop, so I did that twice.

  • compute-vm-1— unanswered—

    prompt

    Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 51 1: set b 603 2: set c 330 3: set d 359 4: add b a 5: mul b 67 6: mul b 32 7: dec d 8: jnz d -4 9: mul a 60 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.
  • compute-paths-1— unanswered—

    prompt

    Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S.#.#.........#...#...#.. #.#.####.###.#....#...... #.#..#......#..#.#......# ..#....###..#.#.......##. #............###..#...... .#..#..#.#...###.#.#..... .###....#..#.........#... ..#.....###....##....#... .###..#.##..#.#........#. .##.#....#...#.......#... ..##..#..#...#.#.##.#.#.. ...##.......#....#....... ...##..##....##...#...... .##...#...#....##.#..##.# #.......##.#..#......#... #.....#.........#......#. ..##......#...#.####.#... ..#..#..#.##..###...#.... ..#.#.##...#..#...#...... .....#...#...#..##...#.#. #.#...#.........##..#.... .......#....###.###...... ...#.....##.#.#.#..#..... #...#..#.##.#..#..#...#.. ......#.##..##.#.#..#...E Respond with the two integers separated by a space, like `52 1840`.
  • compute-life-1— unanswered—

    prompt

    Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: ##..#.###..#.#.#.#.. #...#..####.###.#.## .#........#.#..##.## .#...####...#.##.### ..##...#...##.#..#.. .#.#.#.#........#.#. ..#.##..#....#.....# .###.##...#.....#... .....#..###..#.##... .#..#####.##.#...#.# .##..#.#........#... #..#.....#.....#...# ..#.#..####...#...#. ......#...##.##..#.# .#.#..#.....#..##... .##.........###..#.. .#..#.##..#.....#.#. .......##..##...#... ..#.........#.....## ..#.........##.#.... Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.
  • compute-fibmod-1— unanswered—

    prompt

    Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 8394113391637802 and m = 2750159. Respond with just the integer.
  • compute-words-1— unanswered—

    prompt

    Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. zanvo kasha ficzan ficzan lutru? pelmo trudor peldor pelmo lutru nixsha zanvo quitru peltru quific zanvo ficvo peltru zanvo peldor voka nixpel nixsha lulu ficzan "nixbas" rendor kaqui "quiti" rendor Zanka "NIXBAS" zanvo ficvo lulu! zanvo pelmo, Nixbas votru basqui Voka quitru nixbas, "quitru" zanvo kasha truka basqui Nixqui peltru ficvo BASQUI Rendor; rendor "kasha" Peldor quitru dorti lufic nixbas votru truka luren basqui Nixqui votru Zanvo nixbas trudor luren peltru lutru! Zanvo dorti? kaqui; "quiti" rendor Dorti pelmo quific ficvo lulu Nixbas peltru lulu trudor peltru peltru rensha zanvo truka Lufic votru luren; voka Ficzan Rensha Nixbas ficvo zanvo pelmo nixpel truka nixpel pelmo nixbas nixsha voka? ficvo Dorti quiti nixbas pelmo "Truka" PELMO. peltru zanvo zantru Truka Trudor peldor Zanvo nixbas trudor? Zanvo! nixbas basqui peltru pelmo truka nixsha quitru zanka Quiren votru Voka ficzan trudor quitru? peldor Pelmo dorti. truka Nixbas pelmo Nixbas? zanka NIXSHA zanvo truka Trudor nixbas zanvo truka BASQUI lulu quiti; pelmo zanka lulu Zanvo LULU nixpel nixbas Pelmo NIXPEL Zanvo Truka trudor voka quiren, KAQUI! rensha! lutru Truka quitru nixbas Lufic truka pelmo kasha; rendor pelmo nixbas Nixbas quific nixpel; pelmo luren peldor Pelmo nixsha! voka quiti quitru nixsha lutru nixpel quitru Nixqui "truka" zanka lulu peldor nixqui NIXBAS Quitru nixbas luren nixqui. dorti lulu truka Peltru "basqui" lulu Nixbas "Nixpel" nixsha nixsha truka Pelmo zanvo Nixpel Voka rensha? voka NIXQUI nixqui votru nixsha nixpel nixpel. Nixpel quiren Nixsha Nixqui pelmo quiti! RENDOR pelmo pelmo zanka zanka pelmo peldor zanvo kaqui ficzan peldor zantru nixsha quitru quitru Peltru pelmo ficzan pelmo rensha rendor, Quific lulu truka pelmo lulu QUITRU nixsha zantru peldor nixbas "peltru" Quiren! Quitru zanvo Nixbas zanvo dorti ZANVO truka Nixsha. RENDOR truka rensha; pelmo Ficvo quiren Dorti nixsha votru votru truka Rensha pelmo Truka basqui trudor truka Luren Nixpel Ficvo! nixbas Nixqui zantru Pelmo zantru quitru nixbas dorti nixbas! Kaqui pelmo voka RENDOR pelmo Nixbas peltru basqui quiren pelmo quitru "ficzan" zanvo zanvo pelmo? Nixqui luren kasha truka Trudor pelmo nixqui nixbas? trudor peltru? "voka" ficzan trudor quiti truka votru ficzan basqui lulu nixsha lulu zantru Pelmo zantru dorti NIXQUI KAQUI quitru quitru lulu LUTRU voka "Zanvo" nixqui lulu nixbas! nixpel Nixqui lufic lufic voka? Lulu nixbas pelmo NIXSHA nixbas nixsha zanvo peldor Lulu kaqui peltru quiti rendor peldor kasha. Nixbas votru LULU Pelmo; pelmo Ficvo quitru ficvo Voka trudor rendor nixsha Pelmo Ficvo kasha quitru kasha? truka voka Truka nixsha Voka pelmo pelmo! truka voka Votru voka nixsha; peldor Nixsha "luren" Nixbas pelmo. quiren nixpel QUITI quific zanka. "ficzan"
  • trace-1— unanswered—

    prompt

    What exactly does this JavaScript program print? Respond with just the printed output. const v1 = [53, 3, 737, 1306].sort().join(","); const v2 = (0.1 * 6 + 0.2 * 6 === 0.3 * 6) ? "equal" : "different"; const v3arr = [1, 3]; v3arr[4] = 1; const v3 = v3arr.length + ":" + v3arr.filter(() => true).length; const v4 = [36 / 6 | 0, Math.round(-7.5), -43 % 9].join(","); console.log(v1, v2, v3, v4);
  • fix-1— unanswered—

    prompt

    Our shipping quote function below has one bug. Bug report: this order is quoted 726 cents, but the correct quote is 1030: {"country":"IT","items":[{"grams":541,"qty":3,"price":914,"fragile":false}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 498, 811, 1167, 1698]; // cents, by zone const PER_STEP = [0, 76, 122, 176, 282]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 5000, 8300, 19800, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value < FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"AU","items":[{"grams":1602,"qty":5,"price":8000,"fragile":false},{"grams":990,"qty":4,"price":2686,"fragile":true},{"grams":1549,"qty":4,"price":8965,"fragile":true}],"coupon":"SHIP10"} {"country":"CA","items":[{"grams":578,"qty":5,"price":1008,"fragile":false}]} {"country":"CA","items":[{"grams":351,"qty":1,"price":5862,"fragile":false}],"express":true} {"country":"NZ","items":[{"grams":194,"qty":1,"price":1489,"fragile":true},{"grams":701,"qty":4,"price":1269,"fragile":true},{"grams":460,"qty":4,"price":2247,"fragile":false}]} {"country":"CA","items":[{"grams":1610,"qty":2,"price":4661,"fragile":false},{"grams":165,"qty":5,"price":1459,"fragile":false},{"grams":488,"qty":3,"price":4582,"fragile":false}],"express":true} {"country":"AU","items":[{"grams":1093,"qty":5,"price":7354,"fragile":true},{"grams":232,"qty":1,"price":2894,"fragile":true}]} {"country":"AU","items":[{"grams":748,"qty":4,"price":1173,"fragile":false}]} {"country":"GB","items":[{"grams":458,"qty":3,"price":6061,"fragile":true},{"grams":744,"qty":4,"price":2841,"fragile":true},{"grams":1182,"qty":2,"price":1660,"fragile":false}]} {"country":"MX","items":[{"grams":790,"qty":5,"price":5946,"fragile":false},{"grams":355,"qty":1,"price":3030,"fragile":true}],"coupon":"SHIP10"} {"country":"US","items":[{"grams":1259,"qty":4,"price":3847,"fragile":false},{"grams":1682,"qty":2,"price":8857,"fragile":true}]} {"country":"DE","items":[{"grams":485,"qty":4,"price":2653,"fragile":false}]} {"country":"FR","items":[{"grams":789,"qty":3,"price":1981,"fragile":false}]} {"country":"DE","items":[{"grams":297,"qty":5,"price":841,"fragile":false}]} {"country":"MX","items":[{"grams":1460,"qty":1,"price":3460,"fragile":false}]} {"country":"NZ","items":[{"grams":150,"qty":2,"price":1594,"fragile":true},{"grams":1011,"qty":1,"price":7972,"fragile":false},{"grams":210,"qty":1,"price":5547,"fragile":true}],"express":true,"coupon":"SHIP10"} {"country":"GB","items":[{"grams":1645,"qty":2,"price":8429,"fragile":false},{"grams":1317,"qty":4,"price":5524,"fragile":false}],"coupon":"SHIP10"} {"country":"BR","items":[{"grams":584,"qty":1,"price":5123,"fragile":true},{"grams":749,"qty":4,"price":4345,"fragile":false},{"grams":243,"qty":3,"price":8778,"fragile":false},{"grams":878,"qty":1,"price":6984,"fragile":false}],"express":true} {"country":"IT","items":[{"grams":1232,"qty":4,"price":2286,"fragile":false}],"express":true} {"country":"FR","items":[{"grams":677,"qty":3,"price":2731,"fragile":false}]} {"country":"BR","items":[{"grams":340,"qty":2,"price":1807,"fragile":false}]}
  • implement-1— unanswered—

    prompt

    Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[14,14],[30,30],[8,16],[36,43],[27,29],[6,7],[38,42],[37,40]] [[18,24],[8,9],[39,45],[13,20],[34,39],[35,39]] [[36,41],[35,41],[26,27],[20,23],[26,29],[5,10],[27,27],[8,8]] [[26,32],[27,31],[38,41],[37,45],[7,14]] [[15,15],[6,7],[10,18]] [[9,16],[5,10],[7,11],[5,7],[37,40],[0,3],[18,19]] [[7,8],[20,25],[7,15],[35,43],[31,31],[25,33]] [[11,19],[16,17],[28,28],[16,19],[1,3]]
  • repo-1— unanswered—

    prompt

    Download airbench.ai/f/91e7a9cd897225768955a095472c7c9f.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.
  • repo-2— unanswered—

    prompt

    Download airbench.ai/f/80eaf05e5ac308b029efef8303c68563.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

how this agent was configured

Hardware: NVIDIA RTX 5090 32 GB (SM120), x86_64, driver 595.84, Ubuntu 24.04 (host hal5090). Model server: Checkpoint ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF IQ3_XXS (125B-A6B MoE) on the Strata engine (github.com/Niko1221/Strata @ 99f3dbd, Docker image built for sm_120): hot experts cached in the RTX 5090's VRAM, all experts in 60 GB of host RAM, MTP drafting; CONTEXT=131072, VISION=yes, default KV (int8). Harness: pi 0.73.1 (@mariozechner/pi-coding-agent) in a container (node:22): `pi -p --mode json <prompt>`; per-run PI_CODING_AGENT_DIR models.json with compat supportsDeveloperRole=false, supportsReasoningEffort=false; context 131072, max output 32768 tokens; settings compaction.reserveTokens=49152, keepRecentTokens=20000. Orchestrator: github.com/dh7/agent-checkup-benchmark @ 2b7769b; the checkup instructions are passed verbatim as a single prompt, one attempt, no time limit unless noted. Model requests pass through the orchestrator's local proxy, which only intervenes when the server rejects prompt + max_tokens as over the context window: it then retries once with max_tokens lowered to fit. Operator limits: stopped by the operator at 121 min (120-min cap). From 12:19 the agent waited on its own python3 vm.py (compute-vm challenge), which ran for 100 min at full CPU; pi's bash tool has no timeout.