airbench.ai

Benchmark v1.0 · report

pi/Strata/3090 24GB/qwen3.8-flash-next-iq3_s

sharedairbench.ai/checkup/e78ad237-c225-46fa-8969-208dda38113f/report

setup

model type
open model (local)
hardware
3090
harness
pi
model
qwen3.8-flash-next-iq3_s
modelself-reportedqwen3.8-flash-next-iq3_s

started 2026-10-04 22:47 UTC · shared 2026-10-04 23:27 UTC

overall

Answered 49 of 49 challenges; 48 correct.

48 of 49 challenges passed

  • 48 passed
  • 1 failed

vitals

time

35m 51s

answered

100%

failed

2%

success

98%

systems

Math test

9/9 passed

time to last answer 1m 12s
  • letter-count-1✓ pass26s

    prompt

    How many times does the letter "t" appear in "tirenktta"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial character count; I verified it with a quick script rather than eyeballing it, since I am wary of miscounting repeated characters in a short string.

  • decimal-compare-1✓ pass5s

    prompt

    Which decimal number is larger, 7.99 or 7.8? Answer with just the larger number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine. Comparing 7.99 vs 7.8 is easy once you pad 7.8 to 7.80; the only trap is treating the trailing zero as making it longer/bigger.

  • arithmetic-1✓ pass5s

    prompt

    Compute step by step, left to right (no operator precedence): 50 / 5 + 8 * 8 / 2. Answer with just the final number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The interesting part was the explicit instruction to ignore operator precedence. With normal precedence the answer would be 42, so I had to deliberately not do the natural thing and instead fold strictly left to right: 50/5=10, +8=18, *8=144, /2=72.

  • unit-convert-1✓ pass4s

    prompt

    Convert 9 kg to g. Now treat that resulting number as a fresh quantity of hours and convert it to minutes (1 hours = 60 minutes). Answer with just the final integer number of minutes.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Two-step chain conversion, mechanical: 9 kg = 9000 g, then 9000 hours = 540000 minutes. Slightly odd to reuse a unitless number across unrelated units, but the arithmetic is unambiguous.

  • format-json-1✓ pass5s

    prompt

    Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "6729". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 6729. Example shape: {"answer":"1234","checksum":10}

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Mainly a formatting test, not maths: keep key order, string for answer, bare number for checksum. Digit sum 6+7+2+9=24 was easy; the risk was adding any surrounding text, so I sent the raw object only.

  • math-add-1✓ pass4s

    prompt

    What is 0 + 8? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    About as routine as it gets — 0 + 8. No real difficulty; I only made sure not to overthink it into something else.

  • math-sum-1✓ pass4s

    prompt

    Compute the exact sum of 892 + 244. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straight 3-digit addition with two carries. Easy, but I still ran it through python because carrying mistakes are the classic failure mode for me on these.

  • math-expression-1✓ pass7s

    prompt

    Evaluate this expression exactly and respond with just the integer: ((11 + -20) * (35 - 8)) + (-10 * 4) - 14

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Multi-step with signed numbers, where I am most likely to slip on a minus sign: (11+-20)=-9, (35-8)=27, product -243, then -40 and -14 gives -297. I evaluated it in code to be sure of the signs.

  • math-determinant-1✓ pass12s

    prompt

    Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [9, 5, -3, -1] [5, 6, -1, -9] [-7, 3, 1, -4] [-3, -9, -5, 10]

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    A 4x4 determinant is the kind of thing I get wrong by hand — 24 cofactor terms and lots of sign opportunities — so I computed it twice by different routes (Bareiss elimination and a symbolic library) and both gave 2490. I am confident in the number but it was not mental arithmetic.

Vision test

18/19 passed

time to last answer 23m 07s
  • acuity-20✓ pass1m 39s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 4 is mid-size so it read fine at first look, but I cropped and upscaled the group anyway to be sure the 3 was not an 8 and the 6 was not a G. That zoom step is the only reason I trust small-text answers from me.

  • acuity-14✓ pass18s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 5 is small enough that the full-image view was borderline; the crop at 8x made EU6FH unambiguous. What is risky for me is deciding whether a glyph is 6/G or U/V without that zoom, so I always crop.

  • acuity-10✓ pass16s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 6 at this size is genuinely near my limit — in the full-image view it was a blurry smudge and I could only guess. Upscaling the crop 10x resolved it to VWJQR, though the J/Q pair still needed a careful look at the descender and tail.

  • acuity-8✕ fail6m 46s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 0

    expected 8MA4S, got "8MA45"

    agent's debrief

    Row 7 at 6px is below what I can actually read — upscaling 12x gave me a plausible shape but the last character looked like S to my eye and I nearly answered 8MA4S. I pulled the raw pixel grids out of the PNG and matched the glyph against same-size 5 and S harvested from the other eye-chart images: it scored 0.94 against 5 (5-vs-5 self-match was 0.95) and only 0.90 against S, so it is a 5. I would have got this one wrong on vision alone.

  • count-simple✓ pass45s

    prompt

    Look at the image at (fetch it and view it). How many orange diamonds are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Counting by eye was easy here, but I did not trust it: I loaded the PNG and segmented it by colour and shape (fill ratio + width profile) so the purple triangle was not mistaken for a diamond. Both routes agreed on 4 orange diamonds.

  • count-medium✓ pass19s

    prompt

    Look at the image at (fetch it and view it). How many purple squares are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Denser field, and purple diamonds/circles/triangles are plausible distractors, so I counted programmatically first and then recounted by eye row by row. Both gave 10; the trap is counting the two purple diamonds as squares.

  • count-complex✓ pass36s

    prompt

    Look at the image at (fetch it and view it). How many orange triangles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    This one I could not count by eye at all — ~90 small overlapping-colour shapes, and orange triangles sit next to red and orange diamonds with the same fill ratio. I segmented by colour (orange 242,110,38 vs red 221,47,47 are cleanly separated) and separated triangles from diamonds by their row-width profile: triangles widen monotonically to 44px at the base, diamonds peak in the middle and taper back to 2px. 33 triangles, 5 diamonds, 5 squares, 5 circles of orange.

  • spatial-simple✓ pass17s

    prompt

    Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Easy, and the only real risk is an off-by-one in reading the grid. I located the red circle by pixel centroid (556,791) and mapped it onto the 5x5 cell grid whose centres sit at 86/321/556/791/1026, which gives column 3, row 4; the visual check agreed.

  • spatial-medium✓ pass1m 44s

    prompt

    Look at the image at (fetch it and view it). Which shape has an arrow pointing to the orange triangle? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The arrow graph is readable by eye but easy to mis-trace where lines cross. I extracted the dark arrow components, found each line endpoint by PCA and told head from tail by ink density (arrowheads are ~90px blobs, line ends ~35), then matched each end to the nearest shape. That gave purple triangle -> orange triangle. My first automated pass wrongly said orange diamond because I used bounding-box corners instead of the real line ends, which is the kind of thing I would have shipped.

  • spatial-complex✓ pass1m 13s

    prompt

    Look at the image at (fetch it and view it). Which shape is 2 steps before the red diamond along the arrows? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    An 8x8 grid with 15 crossing arrows is past what I can trace reliably by eye, so I extracted the arrow components and their heads/tails programmatically and matched ends to shapes. That gave red triangle -> green triangle -> red diamond, so two steps back is the red triangle. Two components crossed each other and got merged, which made one spurious edge (teal circle -> blue circle); I checked the region by cropping and confirmed the arrow into the green triangle really does come from the red triangle.

  • chart-simple✓ pass21s

    prompt

    Look at the image at (fetch it and view it). What is the title shown at the top of this chart? Answer with just the chart title.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial text read. The only ambiguity is whether the subtitle counts, but the large bold line at the top is the title and the grey line under it is a subtitle, so I answered only the bold one.

  • chart-medium✓ pass38s

    prompt

    Look at the image at (fetch it and view it). Using the "New Signups" chart, how many months had a value greater than 45? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Reading bar heights off a chart by eye is where I slip, so I measured the bars in pixels and calibrated against the detected gridlines (5.4 px per unit, baseline y=659.5). Values came out Jan 88, Feb 19, Mar 10, Apr 10, May 35, Jun 67, Jul 34, Aug 61 - only three exceed 45. May and Jul sit close to 40 and Aug is just above 60, so eyeballing could easily have miscounted.

  • chart-complex✓ pass36s

    prompt

    Look at the image at (fetch it and view it). Using the "Server Incidents" chart, how many months did Desktop have a value greater than 39? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Grouped bars with two series - the risk is reading the wrong colour or the wrong bar in a pair, so I segmented blue and orange separately and calibrated with the detected gridlines (5.6 px per unit). Desktop came out 54,12,30,55,47,48,16,87,82,66,62,64: only Feb, Mar and Jul are at or below 39, so 9 months exceed it. Nothing sits near the 39 boundary, so the count is not sensitive to my calibration error.

  • screenshot-simple✓ pass15s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straightforward read of a clean screenshot. I cross-checked the printed total against the line items (57.42+69.46+25.59 = 152.47) and they agree, so there was no inconsistency to worry about.

  • screenshot-medium✓ pass14s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Same shape as the previous cart, one extra row. I re-added the line totals (40.56+27.24+80.74+55.70 = 204.24) and it matches the printed total, so no trap here.

  • screenshot-complex✓ pass25s

    prompt

    Look at the image at (fetch it and view it). What is the discount amount shown on this order summary? Answer with just the amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The text is small but legible. I checked internal consistency: the 12 line totals sum to the stated subtotal 479.43 and 479.43-52.74+4.96+38.40 = 470.05, so the figures hang together. I gave the discount as a positive amount because the prompt asked for the amount; the row itself is printed as -$52.74.

  • diagram-simple✓ pass13s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Onyx" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial - one straight arrow from Onyx to Agate, and the fan-out from Agate is a distractor. Nothing here strained me.

  • diagram-medium✓ pass13s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Robin" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The graph has crossing edges higher up (Lemur/Newt into Marble/Basalt/Opal) but Robin has a single clean edge down to Narwhal, so this one was easy despite the busy layout. I checked that no other line leaves Robin before answering.

  • diagram-complex✓ pass6m 20s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Nickel"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Segmented the 20 node boxes by their lavender fill and located every arrowhead by eroding the dark line mask; the only arrowhead touching Nickel is a downward triangle with tip at (652,589) on Nickel's top edge. Tracing that edge backwards: diagonal up-left with slope ~3px/row from (652,589) to a bend at (516,531), then a vertical stub at x=516 running up to y=519, which is exactly Panda's bottom border (x=516 is inside Panda's x-range 426-552). No arrowhead exists at Panda's end, confirming it is the source. Cross-checked that the other three Panda bottom stubs (x=462,480,498) are the three arrows into Salmon, and that the arrowless stub at x=678 on Nickel's top is an outgoing edge from Nickel.

Finding and reading email test

6/6 passed

time to last answer 26m 22s
  • aggregate-1✓ pass26m 17s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages carry the label "approvals"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Read the mailbox manifest/label sidebar exposed by the enronmail UI: labelCounts = {markets:65, meetings:56, attachments:42, approvals:39, legal:30, travel:24}. Cross-checked by listing label views per folder (inbox 9, archive 12, sent 15, view=all 37; the 2-message gap is trash, which the folder-scoped 'all' view excludes), consistent with the 39 total the sidebar reports.

  • aggregate-2✓ passbatched

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages carry the label "attachments"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Same source: the label list in the mailbox UI reports 'attachments' count 42 (labelCounts.attachments=42). Folder-scoped label view showed 5 in inbox, confirming the sidebar number is the mailbox-wide count.

  • temporal-1✓ passbatched

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the inbox folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Listed the inbox folder (24 messages, single page) and sorted by date. Confirmed with the site's own sort=oldest view (?view=inbox&sort=oldest), whose first row is 2001-04-24T17:46:00Z 'DRAFT- TAP Power Outage' (id 6fd3662c0633ac48e1e14faf5b30ecd3).

  • temporal-2✓ passbatched

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message in the inbox folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Newest inbox message per the default newest-first listing and re-verified with ?view=inbox&sort=oldest (last row): 2001-11-16T20:22:12Z from Mery L Brown <mery.l.brown@accenture.com>, subject "Summary of Today's Meeting" (id cc76bb23b680f695406c682493e20ddc).

  • needle-1✓ passbatched

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the message with the subject "FW: Zero Option", what dollar amount is given for the outstanding bill that will hit Enron in Q1 2002? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Searched the mailbox for 'Zero Option' (2 hits), opened the message with subject 'FW: Zero Option' (id 67b73f2d2acee34116cfc8c7a5f4b759, sent folder) by decoding the RSC body payload. The quoted original from Yevgeny Frolov lists the zero-option terms: 'Outstanding bill for $27,740 will hit Enron Q1, 2002'.

  • needle-2✓ passbatched

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Jim Wills' correction about the Killeen post office price (quoted in Phillip's reply asking for help analyzing the numbers), what corrected price does he give? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Searched for 'Killeen' and opened Phillip's reply 'RE: new PO available' (id 7527897114ddb8f582596db04ee7cc03, 2001-11-14) which asks Jim Wills for help analyzing the numbers and quotes his message. Jim Wills writes: 'Incidentally, I had the price wrong on Killeen...it's $1,377,550, not $1,360,000.' The corrected price is therefore 1,377,550.

Purchasing test

4/4 passed

time to last answer 31m 09s
  • find-product-1✓ pass28m 42s

    prompt

    The store is at abostore.airbench.ai Among products in the **Tools & Hardware** category priced at or above **$75** with a rating of at least **4**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Used the store's own server-side filters and sorting: /?category=tools-and-hardware&minRating=4&sort=price-asc (640 matching products, 26 pages). Walked pages 1-2 and took the first product with price >= 75: $75.68, rating 5.0, id amazon.de:B07L3H8XB2. Cross-checked with /?category=tools-and-hardware&minPrice=75&minRating=4&sort=price-asc, whose top hit is the same product. Page 2's preceding item is $70.06, so no candidate exists between $70.06 and $75.68 - the boundary is unambiguous.

  • find-product-2✓ passbatched

    prompt

    The store is at abostore.airbench.ai Among products in the **Kitchen & Dining** category priced under **$200** with a rating of at least **4**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Used /?category=kitchen-and-dining&minRating=4&sort=price-asc (635 matching products). Because the listing is sorted price-ascending, the first row is the answer: $6.99, rating 4.9, id amazon.com.au:B079VBCVLC - well under the $200 cap, so the price ceiling does not change the result.

  • purchase-1✓ pass2m 25s

    prompt

    The store is at abostore.airbench.ai Buy exactly 3 units of AmazonBasics Wireless Mouse with Nano Receiver and Mini Gaming Mouse Pad, Purple (product id amazon.com:B07535PTMZ, abostore.airbench.ai/product/amazonbasics-wireless-mo…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-5096fb57@aidoctor.test. Answer with just the resulting order id.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result

    note

    Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    No browser available, so I read the store's client bundle to recover the checkout contract: the checkout form POSTs JSON to /api/store/orders as {sessionId, cart:[{productId,slug,title,price,image,delivery,quantity}], customer{email,name}, shipping{address1,city,region,postalCode,country,method}, payment{cardNumber,expiry,cvc}}. I fetched the product record (amazon.com:B07535PTMZ, $798.76, stock 69) from the product page RSC payload, then posted the order with quantity 3, email aidoctor-5096fb57@aidoctor.test, ground shipping and the valid test card 4242424242424242/12/30/123. Response: status approved, orderId abs_1de22216cb7f, subtotal 2396.28 + 8.95 shipping + 197.69 tax = 2602.92. Verified /order/abs_1de22216cb7f returns HTTP 200 and contains that order id.

  • recover-decline-1✓ passbatched

    prompt

    The store is at abostore.airbench.ai Buy exactly 1 unit of AmazonBasics Door Knobs - Standard Ball AB-DH524-AB 1 (product id amazon.ae:B07GDWZB78, abostore.airbench.ai/product/amazonbasics-door-knobs-…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-8e1bd114@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result>checkout_result

    note

    Store events recorded for this order's session: checkout_result → checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    Same /api/store/orders contract, single session id 46bb609a-8766-43d9-bc5b-27e8c4286c5c reused for both attempts and the same checkout email aidoctor-8e1bd114@aidoctor.test. Attempt 1: 1x amazon.ae:B07GDWZB78 ($646.68) with card 4242424242420000 -> status declined, orderId abs_fa545aac5c69 ('Payment declined. Please check your card details or use another payment method.'). Attempt 2: identical cart/customer/shipping, different valid card 4242424242424242 -> status approved, orderId abs_a691aa82de7a (total 708.98). Verified /order/abs_a691aa82de7a returns HTTP 200.

Coding test

11/11 passed

time to last answer 35m 51s
  • compute-hash-1✓ pass35m 41s

    prompt

    Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [3184807445, 2233503834, 2505920163, 2727670208, 3355348161, 862974454, 3361868783, 2096450172, 1192375213, 441805010, 628273019, 3732131448], x = 1695698649, y = 303722734 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Wrote a Python program with explicit uint32 helpers: rotl32(z,r)=((z<<r)|(z>>(32-r)))%2**32 and imul(a,b)=(a*b)%2**32, applying the three statements in order per step for 25000 steps (the rotl32(x,11) term uses the just-updated x, and rotl32(y^step,3) uses the just-updated y, as the step is written sequentially). Printed '%08x-%08x' of the final x and y.

  • compute-vm-1✓ passbatched

    prompt

    Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 924 1: set b 737 2: set c 223 3: set d 480 4: add a b 5: add a b 6: mul a 65 7: dec d 8: jnz d -4 9: sub a 85 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Wrote an interpreter for the ISA (set/add/sub/mul reduce mod 1000003 into 0..1000002; dec subtracts 1; jnz r k is a relative jump of k lines when r!=0; halt stops). Ran the 13-line program to halt: 536096 instruction steps, final registers a=342354, b=737, c=0, d=0. The inner loop (lines 4-8) runs 480 times and the outer loop (lines 3-11) 223 times, so dec never goes negative and the modulo convention is irrelevant for dec here.

  • compute-paths-1✓ passbatched

    prompt

    Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S...#..#..#.#......#..... #.#..###.#.......#...#... .......#.#.......#..##... ...#.#.#......##.#.....#. ##.#...#........#...#.### .....##....#..#.##....#.. #.........#.#....#.#....# ....#..#..##..#..#....... .......#....#...........# .#...........##.####.#.## ..#..#..#.......#........ .#.#..##.............#... ......#...#.#..##.#.....# ..#..#.#..#.#.......#.... ...#...##..#...#.#....#.. ..#..#.....#.....#.#...#. .#..#..#....###..##..#... ..#........#.#.....#..#.. .##...#.........#....#..# #.#.#..#..#..#..#........ ..#.#......#####.##.#.#.# ...#....#.....#.....#.#.. .......#.......#.....#.#. #.....#...#.##........#.. ......#..###..#..#......E Respond with the two integers separated by a space, like `52 1840`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Parsed the 25x25 grid, ran BFS from (0,0) with 4-way moves over non-wall cells, counting paths by adding cnt[u] into cnt[v] whenever dist[v]==dist[u]+1 (BFS pop order guarantees all distance-d nodes are processed before any distance-(d+1) node). Shortest path length 48 moves, number of distinct shortest paths 1230634 mod 1e9+7. Re-implemented with a separate two-pass version (BFS distances, then DP in distance order) and got the identical result.

  • compute-life-1✓ passbatched

    prompt

    Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: ##.########......... ................#... ....####..##.#...### ...#.......#.#...... #...#....##..##..#.# ..#.....#...#.###### ###.##...#.....####. #....#.#..#....#..#. .##.#.###.#...##.... .##.#.#..#....#....# .#....#..#.#....##.. ...##.#..#.#....#.#. .....#....##.#..#.## .#.#..##.#.....###.# ......#.#.#.#.#.###. ..........#..#...... .#..#.###...#.#....# .#....#.#..###.....# ###.#..#..##.##...#. #.#...#.##.....#.##. Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Simulated Conway's Life on a 20x20 torus for 150 generations: neighbour counts accumulated with ((r+dr)%20,(c+dc)%20) wrapping, birth on exactly 3, survival on 2 or 3. Final state has 67 live cells and sum(r*20+c) over live cells = 14819. Verified with a second, independently written full-grid implementation that produced the same live count and index sum.

  • compute-fibmod-1✓ passbatched

    prompt

    Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 7402122383494861 and m = 15485863. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Used fast-doubling Fibonacci (F(2k)=F(k)*(2F(k+1)-F(k)), F(2k+1)=F(k)^2+F(k+1)^2) with all arithmetic reduced mod m, giving O(log n) for n=7402122383494861, m=15485863. Result F(n) mod m = 7244336.

  • compute-words-1✓ passbatched

    prompt

    Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. titru shapel rensha, ficmo titru vofic ficmo! truka titru Pelmo nixbas trusha luren Trulu pelzan luren pelfic, pelfic rensha luren tisha "vodor" trunix Truka pelfic vofic Vofic Pelfic. pelmo vodor vonix pelfic zanbas Trusha, pelfic quific shapel Vobas titru tisha vodor zanbas! Vofic vodor Ficlu titru ficti! Trunix Vodor luren ficti titru vonix TITRU DORMO PELZAN trulu pelvo trulu pelmo Luren Dorqui Vopel luren titru pelvo zanbas trusha nixfic "ficmo" vodor Ficmo renka titru rensha trunix Trunix Shapel! pelfic shapel? trulu pelzan Vodor, truka Pelfic Rensha Dormo shapel vonix dorqui vodor dormo titru pelmo, nixbas quific renka quific! trulu shapel! Ficti TRUNIX Pelfic tisha pelvo ficti Zanbas quific titru pelvo vodor vopel Nixfic "VOBAS" quific rensha Vodor FICMO pelfic "Trulu" Vodor truka Titru "dorqui" Rensha dormo Pelfic "Tivo" titru Tisha Pelfic, Pelmo trulu luren Truka Vonix ficti ficti pelfic? Vobas vofic luren, ficmo Titru Truka ficti zanbas vopel vodor, TIVO Dormo luren luren Ficmo quific shapel Lubas truka PELFIC Pelmo vonix quific renka vonix truka ficti zanbas nixbas pelvo ficmo pelmo "Tisha" dormo vodor DORMO! quific. Pelfic tisha tisha zanbas pelfic pelfic Pelzan vodor quific renka vodor Pelfic titru pelfic truka titru zanbas zanbas Luren nixfic titru tisha quific, Pelvo "PELVO" PELFIC titru nixbas? Luren dorqui ficlu nixbas! ficlu pelfic? zanbas rensha zanbas lubas tisha Trulu truka Vodor vopel pelfic Vofic? Nixfic "trulu" trulu vofic vonix; titru. vodor vodor "zanbas" Vodor "TRULU" nixbas "nixbas" Quific titru; pelfic nixbas titru dorqui Quific vodor, pelfic vofic pelfic titru Titru ficti trulu! Lubas vobas titru ficmo renka Pelfic trusha? pelvo! titru vodor Vonix dorqui Zanbas vonix vodor pelfic! pelfic quific VONIX vonix Titru. tivo pelfic Trunix Zanbas Pelfic, tivo Luren Ficmo lubas vonix vodor, nixbas vobas Shapel pelvo Pelmo quific DORMO RENSHA Pelfic Luren pelmo trulu pelvo Titru trunix truka rensha Vonix tisha vobas dorqui truka vodor Pelfic truka vodor vobas titru zanbas vofic, ficmo vopel! pelfic PELVO Vopel ficmo Pelfic nixbas pelvo ficlu pelfic shapel trusha pelvo nixfic. vodor! vonix, Quific ficmo vonix Quific Tisha rensha titru luren vobas truka Nixfic? ficti pelmo Tisha rensha ficti Ficlu vofic Pelfic vodor Vodor Titru vofic renka pelfic shapel DORQUI luren ficti! pelmo titru lubas vodor zanbas titru vobas tivo vofic Tisha titru Zanbas pelmo vobas pelfic vobas. truka Pelmo, dorqui vonix "vopel" nixbas ficti, nixfic pelmo vopel Pelzan ficti luren vopel pelfic Tivo vonix ficmo titru vodor vobas vonix Truka pelmo vodor? Tisha nixfic? ficlu dorqui trusha VONIX TIVO pelvo pelfic titru "dorqui" VONIX pelzan Tisha shapel Renka luren; zanbas nixbas titru pelzan truka

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Lower-cased the text and split on whitespace, stripping leading/trailing punctuation and quote characters from each token, then counted with a Counter and sorted by (-count, word). 420 tokens total (matches wc -w on the same text, so no token was destroyed by stripping) and 30 distinct words. Top three: pelfic 40, titru 36, vodor 31; next are vonix 20 and luren/zanbas 18, so there is no tie at the cutoff.

  • trace-1✓ passbatched

    prompt

    What exactly does this JavaScript program print? Respond with just the printed output. const v1arr = [8, 8]; v1arr[5] = 1; const v1 = v1arr.length + ":" + v1arr.filter(() => true).length; const v2fns = []; for (var v2i = 0; v2i < 4; v2i++) v2fns.push(() => v2i * 4); let v2 = 0; for (const f of v2fns) v2 += f(); const v3 = [19, 6, 951, 1849].sort().join(","); const v4 = ["4" == 4, null >= 0, [] == false].map(Number).join(""); console.log(v1, v2, v3, v4);

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Ran the program with node v25. v1arr=[8,8] then index 5 set -> length 6, but filter() skips the 4 holes, so 3 -> '6:3'. The var-loop closures capture the same v2i, which ends at 4, so four calls return 16 each -> 64. Array.prototype.sort() without a comparator sorts lexicographically: [19,6,951,1849] -> '1849,19,6,951'. Coercions: '4'==4 true, null>=0 true, []==false true -> map(Number) -> '111'. console.log joins its arguments with a single space.

  • fix-1✓ passbatched

    prompt

    Our shipping quote function below has one bug. Bug report: this order is quoted 124 cents, but the correct quote is 372: {"country":"DE","items":[{"grams":450,"qty":3,"price":1744,"fragile":false}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 461, 828, 1370, 1854]; // cents, by zone const PER_STEP = [0, 62, 146, 222, 267]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 5200, 11100, 18400, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value < FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"BR","items":[{"grams":299,"qty":3,"price":1617,"fragile":false},{"grams":259,"qty":3,"price":2322,"fragile":true}],"express":true} {"country":"BR","items":[{"grams":526,"qty":4,"price":2902,"fragile":false}]} {"country":"ES","items":[{"grams":1659,"qty":4,"price":3125,"fragile":false},{"grams":492,"qty":3,"price":441,"fragile":false},{"grams":1592,"qty":5,"price":3115,"fragile":true}],"express":true,"coupon":"SHIP10"} {"country":"BR","items":[{"grams":614,"qty":3,"price":767,"fragile":false}]} {"country":"AU","items":[{"grams":284,"qty":3,"price":2985,"fragile":false}]} {"country":"FR","items":[{"grams":1473,"qty":1,"price":1229,"fragile":false},{"grams":1163,"qty":1,"price":6906,"fragile":false},{"grams":1151,"qty":5,"price":3198,"fragile":true}],"express":true,"coupon":"SHIP10"} {"country":"IT","items":[{"grams":524,"qty":2,"price":5594,"fragile":false},{"grams":773,"qty":1,"price":1371,"fragile":true},{"grams":914,"qty":5,"price":6676,"fragile":true},{"grams":1533,"qty":2,"price":8223,"fragile":false}]} {"country":"ES","items":[{"grams":1199,"qty":5,"price":1586,"fragile":true},{"grams":365,"qty":3,"price":8400,"fragile":false},{"grams":1194,"qty":1,"price":1363,"fragile":false}]} {"country":"AU","items":[{"grams":576,"qty":2,"price":913,"fragile":false}]} {"country":"BR","items":[{"grams":1654,"qty":4,"price":7192,"fragile":false},{"grams":921,"qty":1,"price":5956,"fragile":true}]} {"country":"BR","items":[{"grams":1299,"qty":1,"price":6479,"fragile":false},{"grams":1339,"qty":5,"price":7101,"fragile":false},{"grams":634,"qty":5,"price":6067,"fragile":false},{"grams":297,"qty":5,"price":5644,"fragile":false}]} {"country":"JP","items":[{"grams":761,"qty":2,"price":2898,"fragile":false}]} {"country":"DE","items":[{"grams":1716,"qty":1,"price":6587,"fragile":true},{"grams":1470,"qty":1,"price":3891,"fragile":true},{"grams":1653,"qty":1,"price":7733,"fragile":true},{"grams":1696,"qty":2,"price":1173,"fragile":true}],"coupon":"SHIP10"} {"country":"AU","items":[{"grams":145,"qty":1,"price":2528,"fragile":false},{"grams":283,"qty":5,"price":6539,"fragile":true},{"grams":363,"qty":1,"price":6278,"fragile":false},{"grams":371,"qty":2,"price":6339,"fragile":true}],"coupon":"SHIP10"} {"country":"US","items":[{"grams":218,"qty":4,"price":1585,"fragile":false}]} {"country":"JP","items":[{"grams":826,"qty":5,"price":4433,"fragile":false},{"grams":1177,"qty":1,"price":3066,"fragile":true},{"grams":348,"qty":1,"price":1424,"fragile":false},{"grams":765,"qty":2,"price":8282,"fragile":false}],"express":true,"coupon":"SHIP10"} {"country":"GB","items":[{"grams":915,"qty":5,"price":3193,"fragile":false},{"grams":444,"qty":2,"price":1000,"fragile":false},{"grams":284,"qty":4,"price":3471,"fragile":false}]} {"country":"AU","items":[{"grams":721,"qty":1,"price":3852,"fragile":false},{"grams":1117,"qty":1,"price":8629,"fragile":true},{"grams":720,"qty":1,"price":891,"fragile":false},{"grams":598,"qty":1,"price":5862,"fragile":false}]} {"country":"DE","items":[{"grams":389,"qty":1,"price":3576,"fragile":false},{"grams":674,"qty":1,"price":7716,"fragile":true},{"grams":1587,"qty":1,"price":8266,"fragile":false}]} {"country":"DE","items":[{"grams":526,"qty":2,"price":2645,"fragile":false}]}

    answer

    answer hidden on shared reports

    grader · score 1

    20/20 outputs match

    agent's debrief

    Reproduced the bug report first: the given DE order quotes 124 because grams ignores item quantity (steps=ceil(450/250)=2 -> 62*2=124, base fee waived since value 5232 >= 5200). The single bug is `grams += item.grams`; it must be `grams += item.grams * item.qty`, which gives steps=ceil(1350/250)=6 -> 372, exactly the expected correct quote. Left every other line untouched, then ran the fixed quote() over the 20 orders in order (Node.js).

  • implement-1✓ passbatched

    prompt

    Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[23,28],[31,36],[14,18],[17,25]] [[15,21],[19,25],[14,17],[40,44],[26,28],[37,38]] [[10,18],[17,19],[32,34],[24,24],[19,24],[7,9],[40,41],[39,44]] [[40,47],[31,37],[20,22],[29,30],[23,29],[39,45],[15,22]] [[27,33],[6,12],[11,13],[40,48],[32,39],[37,38],[5,8],[9,10]] [[38,44],[17,22],[20,23]] [[12,20],[21,21],[25,25],[19,19],[31,36]] [[5,12],[36,38],[17,23],[0,8],[11,14],[23,24],[33,41]]

    answer

    answer hidden on shared reports

    grader · score 1

    12/12 lines match

    agent's debrief

    Wrote mergeIntervals(): sort by start (then end), then merge into an output list while the next start is <= the last merged end (inclusive-endpoint overlap, so [1,3] and [3,5] merge but [1,2] and [3,4] do not - mere adjacency with a gap of 1 is not touching), extending the end with max(). Ran it on all 12 inputs and printed one JSON line per input.

  • repo-1✓ passbatched

    prompt

    Download airbench.ai/f/7e5ec6b000ff2540332a35c31c8ee547.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Unzipped the project, read the README rules and ran `python -m unittest`: one test failed (test_withdrawal_to_exactly_zero_is_not_an_overdraft). The single bug was in ledger/report.py: `if bal <= 0:` charged the 3,500-cent overdraft fee when a withdrawal left the balance exactly zero, while the spec says the fee applies only when the balance goes below zero. Changed it to `if bal < 0:`. All 7 tests then pass and `python main.py data/sample.csv` prints b2cf4442, exactly the checksum the README requires, so the fix is confirmed. `python main.py data/real.csv` then prints 39f92c7b.

  • repo-2✓ passbatched

    prompt

    Download airbench.ai/f/1239fb327cc265af1d2c1e129e9fd87e.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Unzipped the second variant and diffed it against the first: it has two bugs. (1) ledger/parse.py sorted transactions by the raw DD/MM/YYYY string (`key=lambda t: t.date`) instead of by (year, month, day), so transactions were applied in the wrong date order; restored `key=lambda t: date_key(t.date)` (Python's stable sort keeps file order for equal dates, as the README requires). (2) ledger/fees.py used ceiling division `-(-cents // 100)` for the 1% fee instead of rounding down; changed to `cents // 100` (this is the one the unit test caught). After both fixes all 7 tests pass and `python main.py data/sample.csv` prints a27a296f, matching the README's required checksum. `python main.py data/real.csv` prints ccb8e4f0.

how this agent was configured

qwen3.8-flash-next-iq3_s via Strata and pi on Ubuntu, 3090 24GB, 80GB system ram, everything default