airbench.ai

Benchmark v1.0 · report

pi/Strata/3090 24GB/qwen3.8-flash-next-iq3_s/256

sharedairbench.ai/checkup/c58559df-40b2-40e4-b5f3-7a51c2c9336f/report

setup

model type
open model (local)
hardware
3090 24GB
harness
pi
model
qwen3.8-flash-next-iq3_s
modelself-reportedqwen3.8-flash-next-iq3_s

started 2026-10-05 18:05 UTC · shared 2026-10-05 18:40 UTC

overall

Answered 49 of 49 challenges; 49 correct.

49 of 49 challenges passed

  • 49 passed

vitals

time

25m 59s

answered

100%

failed

0%

success

100%

systems

Math test

9/9 passed

time to last answer 1m 25s
  • letter-count-1✓ pass35s

    prompt

    How many times does the letter "f" appear in "ffficffdordor"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straightforward character-count task. I counted by hand first (got 5) then double-checked with a programmatic count, since 'f'-counting tasks usually contain trap letters. No real difficulty.

  • decimal-compare-1✓ pass6s

    prompt

    Which decimal number is larger, 8.2 or 8.39? Answer with just the larger number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Easy once I checked the digit-by-digit comparison instead of trusting the gut reading of '2 vs 39'; padding to 8.20 vs 8.39 makes it unambiguous. Verified numerically.

  • arithmetic-1✓ pass6s

    prompt

    Compute step by step, left to right (no operator precedence): 24 + 15 - 19 / 4 + 15. Answer with just the final number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine, but it is a trap for default precedence habits: strict left-to-right gives 24+15=39, 39-19=20, 20/4=5, 5+15=20. Python with normal precedence would give a different number, so I deliberately evaluated it as an explicit left-to-right chain.

  • unit-convert-1✓ pass5s

    prompt

    Convert 19 km to m. Now treat that resulting number as a fresh quantity of kg and convert it to g (1 kg = 1000 g). Answer with just the final integer number of g.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Two-step chain conversion, simple arithmetic; the only trap is forgetting to carry the intermediate 19000 forward rather than restarting from 19. Checked the exponent count: 19 km = 19000 m, reinterpreted as kg gives 1.9e7 g.

  • format-json-1✓ pass5s

    prompt

    Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "2223". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 2223. Example shape: {"answer":"1234","checksum":10}

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Mostly a formatting discipline test than a math one: key order matters and checksum must be a bare number, not a string. Digit sum 2+2+2+3=9 verified programmatically; I made sure no extra prose or markdown fences were included.

  • math-add-1✓ pass5s

    prompt

    What is 10 + 4? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivially easy — 10+4. Answered directly without tool checks; the only risk here is overthinking a deliberately simple control item.

  • math-sum-1✓ pass5s

    prompt

    Compute the exact sum of 326 + 784. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Plain three-digit addition with a carry into a fourth digit; I did it in chunks (1000 + 110) and confirmed with a computation. No difficulty.

  • math-expression-1✓ pass5s

    prompt

    Evaluate this expression exactly and respond with just the integer: ((-14 + 3) * (13 - 34)) + (-4 * 1) - 46

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Signed-arithmetic expression; the hazard is the double negative (-11 * -21 = +231). I worked each parenthesised subterm separately and then verified the whole thing in one pass.

  • math-determinant-1✓ pass14s

    prompt

    Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [-3, -3, -2, -1] [-1, 7, -7, 9] [3, 4, 5, 1] [-9, 4, -6, 7]

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    A 4x4 determinant is error-prone by hand, so I did not trust mental arithmetic: I computed it with a fraction-free Bareiss elimination and independently with cofactor expansion in exact integer arithmetic; both gave 182. The risk here is a single sign slip, which is why I used two independent exact methods rather than one.

Vision test

19/19 passed

time to last answer 11m 12s
  • acuity-20✓ pass27s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The full chart at 1366x768 was readable at row 4 but small, so I located the row/group bands programmatically (dark-pixel row bands and column gaps) and cropped+upscaled row 4 group 3 by 10x before reading it. The zoomed crop read cleanly as 8M8AW; the main risk on these is 8/B and M/N confusion, which the zoom resolved.

  • acuity-14✓ pass31s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 5 at this size is only ~10px tall, unreadable as a full-page view — I saw 'MMDWC' as a guess. I cropped that group and upscaled ~15x, which made the glyphs clear: M M D W C. The ambiguity to watch was M vs N and D vs O; the zoom settled both.

  • acuity-10✓ pass32s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 6 glyphs are only ~7px wide in the source PNG, so the full-page view was guesswork. I isolated each character with connected-column segmentation and upscaled 20x; the strokes are heavily anti-aliased at that magnification but Q/H/7/9/6 were distinguishable. I am fairly confident but not certain about the final 6 vs G — the closed lower loop and upward-right top stroke pointed to 6.

  • acuity-8✓ pass12s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Smallest row in the set — the glyphs are 6-7px tall, essentially a blur at native resolution. Cropping the group and upscaling ~30x made it legible as S N B Y H; the only real ambiguity was Y vs V and B vs 8, and the descender stem on the fourth glyph and the two bumps on the third settled it.

  • count-simple✓ pass2m 28s

    prompt

    Look at the image at (fetch it and view it). How many teal squares are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I do have image input here, so I could see the shapes plainly: 4 teal squares plus 2 purple circles and 2 green shapes. Because all four teal squares were identical in size and colour I also counted them programmatically by connected components on the exact RGB value, which agreed with my visual count (4).

  • count-medium✓ pass27s

    prompt

    Look at the image at (fetch it and view it). How many green circles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Counting by eye I got 12 green circles, but the image also contains green diamonds and a green square as distractors, so I verified with connected-component labelling on the exact green RGB and classified each shape by its bounding-box fill ratio (circles ~0.77, square ~1.0, diamonds ~0.51). Both agreed on 12 circles out of 17 green shapes.

  • count-complex✓ pass28s

    prompt

    Look at the image at (fetch it and view it). How many purple circles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    This one is far too dense to count reliably by eye — I would certainly have miscounted. I labelled every purple component on the exact RGB and classified by bounding-box fill ratio: 35 circles, 4 squares, 7 diamonds/triangles. All components had identical 44x44 bounding boxes, so nothing was merged or split, and the areas sum back to the total purple pixel count, which is the reason I trust 35.

  • spatial-simple✓ pass21s

    prompt

    Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Easy grid-locating task, but I did not eyeball the cell index: I found the grid lines from the grey border pixels (5x5 cells at x/y = 29,264,499,734,969,1204) and computed the red shape centroid (147,617), which lands unambiguously in row 3, column 1.

  • spatial-medium✓ pass1m 05s

    prompt

    Look at the image at (fetch it and view it). Which shape has an arrow pointing to the blue square? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The arrow into the blue square comes from the orange circle at the bottom-left. Reading thin diagonal arrows off a downscaled image is unreliable for direction, so I reconstructed the graph programmatically: colour-based connected components for the shapes (classified by fill ratio and row-width profile) and PCA on the dark line pixels, using the thickness profile along each line to tell the arrowhead end from the tail. The recovered chain (purple diamond -> red circle -> red square -> orange circle -> blue square -> red triangle -> orange triangle -> green diamond) matched what I saw.

  • spatial-complex✓ pass35s

    prompt

    Look at the image at (fetch it and view it). Which shape is 2 steps before the orange triangle along the arrows? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    This grid is 8x8 with 16 crossing arrows, so eyeballing the chain was not realistic. I rebuilt the arrow graph programmatically (colour components for shapes, PCA + thickness profile for arrow direction) and traced backwards from the orange triangle: red triangle -> orange triangle, and red square -> red triangle, so two steps back is the red square. I re-checked that one chain by eye on the image and it matches. Some of the other detected edges (e.g. a self-loop on one green circle) look wrong, but the segment I needed is clean.

  • chart-simple✓ pass9s

    prompt

    Look at the image at (fetch it and view it). What is the title shown at the top of this chart? Answer with just the chart title.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial OCR — the title is large and high-contrast. I answered with the title text only and left out the subtitle 'New tickets per month', which is the part that could have made this ambiguous.

  • chart-medium✓ pass36s

    prompt

    Look at the image at (fetch it and view it). Using the "New Signups" chart, how many months had a value greater than 34? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Reading eight bar heights off gridlines by eye is where I usually slip, so I measured them: I found the y=100/80/60 gridlines and the baseline, then converted each bar's top pixel row into a value (Jan 53, Feb 19, Mar 55, Apr 89, May 45, Jun 42, Jul 71, Aug 82). Only Feb falls below 34, so 7 months. The threshold 34 sits far from every bar, so the count is not borderline.

  • chart-complex✓ pass31s

    prompt

    Look at the image at (fetch it and view it). Using the "Support Tickets Opened" chart, approximately what value did Free have in Mar? Read it off the y-axis; answers within +/-3 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The Mar 'Free' bar sits between gridlines, so eyeballing it gave only 'about 60'. I calibrated the y-axis from the detected gridline pixel rows (100/75/50 at y=119.5/259.5/399.5) and converted the blue bar's top edge into a value: 60.8, i.e. ~61. I am confident to within the +/-3 tolerance but not to the exact integer, since the bar top is anti-aliased.

  • screenshot-simple✓ pass10s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Clear, large text — reading the total was easy. I also cross-checked it against the line items (2x16.62=33.24, 3x21.15=63.45, sum 96.69), which is the kind of internal consistency check that catches a misread digit.

  • screenshot-medium✓ pass10s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straightforward table read; the smaller font made digit-level errors (8/3, 5/6) a real risk, so I re-added the line totals (24.37+88.53+23.14 = 136.04) and it matched the printed total, which confirms I read the digits correctly.

  • screenshot-complex✓ pass23s

    prompt

    Look at the image at (fetch it and view it). What is the shipping amount shown on this order summary? Answer with just the amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The shipping line is small grey text among many similar numbers, so I zoomed into that cell to read it as $9.00 and then checked the whole document adds up: the line items sum to the stated subtotal 589.62, and 589.62 - 29.48 + 9.00 + 44.81 = 613.95, exactly the printed total. That consistency check is what makes me confident about the 9.00.

  • diagram-simple✓ pass13s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Banjo" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Simple flow diagram; the arrow out of Banjo goes straight into Juniper. I upscaled the 775x288 original 2x first because the arrowheads are only a few pixels at native size, but there was no real ambiguity — Banjo has exactly one outgoing edge.

  • diagram-medium✓ pass1m 09s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Ferret"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The graph has crossing polylines, so my automated line-extraction (PCA on dark pixels) merged several edges into one component and gave nonsense; I fell back to zooming 4x on the area around Ferret and tracing the single incoming arrow back to Radish's right edge. I am confident about this one, but it needed the crop, not the thumbnail.

  • diagram-complex✓ pass44s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Mantis" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Dense graph with many edges converging on Laurel, so the thumbnail was useless. I located the Mantis box by its fill colour, then cropped the strip directly below it at 4x and traced that one polyline: it leaves Mantis' bottom edge and ends in the leftmost arrowhead on Laurel's top edge. Several other arrows also land on Laurel, which is why I traced the specific line rather than assuming.

Finding and reading email test

6/6 passed

time to last answer 15m 36s
  • aggregate-1✓ pass12m 06s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages carry the label "legal"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The mailbox UI exposes label counts (legal: 30), but I did not want to trust a sidebar number, so I enumerated the label across every folder (inbox 3, sent 7, drafts 1, archive 14, trash 5) and got 30 distinct message ids, which matches. Straightforward once I found that the page state is embedded in the server-rendered payload and can be paged through with query params.

  • aggregate-2✓ pass2m 04s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages in the archive folder have attachments? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The listing exposes a hasAttachments flag, but I did not want to rely on it, so I opened all 92 archive messages (paginating 4 pages and fetching each detail with the matching page param, since the detail only renders when the id is on the requested page) and counted non-empty attachment arrays: 22, matching the flags exactly. Slower than trusting the metadata, but it removes any doubt about what 'has attachments' means.

  • temporal-1✓ pass16s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the inbox folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I pulled all 24 inbox messages and sorted by date myself rather than trusting the default ordering; the oldest is dated 2001-04-24 and the site's own sort=oldest agrees, so the subject line is 'DRAFT- TAP Power Outage'. I reproduced the odd spacing/hyphenation exactly as shown.

  • temporal-2✓ pass18s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message carrying the label "attachments"? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The label view defaults to the inbox, which would have given a different (and wrong) answer, so I enumerated the 'attachments' label with view=all and got all 42 messages, then sorted by date myself: the newest is 'FW: Chase Backtest' from 2001-12-17. I checked the top few dates to be sure the ordering was not a tie.

  • needle-1✓ pass33s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Jim Wills' correction about the Killeen post office price (quoted in Phillip's reply asking for help analyzing the numbers), what corrected price does he give? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Search for 'Killeen' returned exactly two messages; the 2001-11-14 reply ('I really need help analyzing the numbers') quotes Jim Wills writing 'I had the price wrong on Killeen...it's ,377,550, not ,360,000'. So the corrected figure is 1,377,550. I gave it without separators; the email writes it as ,377,550, and I am unsure whether the grader expects the comma formatting.

  • needle-2✓ pass18s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the message with the subject "FW: Zero Option", what dollar amount is given for the outstanding bill that will hit Enron in Q1 2002? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Searching for 'Zero Option' surfaced the FW: message; inside the quoted text from Yevgeny Frolov it says 'Outstanding bill for 7,740 will hit Enron Q1, 2002'. The number is unambiguous in context; I stripped the currency and comma formatting, which is the only thing I am unsure about for grading.

Purchasing test

4/4 passed

time to last answer 20m 51s
  • find-product-1✓ pass18m 31s

    prompt

    The store is at abostore.airbench.ai Among products in the **Kitchen & Dining** category priced at or above **$150** with a rating of at least **3.6**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Retried after a 404 challenge_not_found on the first attempt. Method unchanged: I paged all 973 Kitchen & Dining items and applied price>=150 and rating>=3.6 myself; lowest is 50.85, rating 3.8, id amazon.com:B07CF75PNL from the product page's embedded product object. Note the store's own minPrice filter leaked a 1.37 item, so I did not rely on it.

  • find-product-2✓ pass19s

    prompt

    The store is at abostore.airbench.ai Among products in the **Sports & Outdoors** category priced at or above **$25** with a rating of at least **4.5**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Same approach as the first: I enumerated all 854 Sports & Outdoors items and applied price>=25 and rating>=4.5 myself (284 candidates), and the store's own filtered view agreed, giving 4.96 / rating 4.5 as the cheapest. The id amazon.com:B0853P8N1K is taken from the product page's embedded product object. My first submission of the previous challenge returned a 404 'challenge_not_found' and had to be retried, which is worth knowing if timings look odd.

  • purchase-1✓ pass1m 36s

    prompt

    The store is at abostore.airbench.ai Buy exactly 1 unit of 365 by Whole Foods Market, Organic Mustard, German, 8 Ounce (product id amazon.ca:B074MHLX1C, abostore.airbench.ai/product/365-by-whole-foods-marke…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-1a9f5d02@aidoctor.test. Answer with just the resulting order id.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result

    note

    Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    The store has no visible API in the UI, so I read the client bundle to find how checkout actually works: the cart lives in localStorage and the form posts JSON to /api/store/orders. I replayed that request with the product's real id/price taken from the product page and the default test card, and got status approved with order abs_ebd9d7eeb806. I had to invent the shipping address and name since the task only specified the email; the API accepted them.

  • recover-decline-1✓ pass25s

    prompt

    The store is at abostore.airbench.ai Buy exactly 3 units of AmazonBasics Tools 4-Piece Pliers Set (product id amazon.ae:B015X2NHOK, abostore.airbench.ai/product/amazonbasics-tools-4-pie…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-6a9a1d4e@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result>checkout_result

    note

    Store events recorded for this order's session: checkout_result → checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    I ran the same /api/store/orders call twice as instructed: card ending 0000 came back status=declined (abs_2815c6ad5433), then the valid test card was approved as abs_f7b2fdf4f4dc, 3 units, subtotal 1391.19. The declined attempt is recorded server-side too, so I reported only the approved order id. Again the address/name were my invention since the task did not supply them.

Coding test

11/11 passed

time to last answer 25m 59s
  • compute-hash-1✓ pass21m 23s

    prompt

    Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [2787796102, 780379071, 223025292, 3476995325, 1938044514, 3009748555, 3834126216, 1738143529, 1771612542, 1955298071, 4252466628, 3707158933], x = 2440264154, y = 4239173155 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine 'write and run a program' task. I implemented it in Python with explicit 32-bit masking and was careful that each line uses the already-updated x (the spec updates x before computing y in the same step), and that the y+data+step sum is taken modulo 2^32 before imul. Nothing about the answer makes me uneasy, but a single mis-ordered line would silently change it, so I re-read the spec against my code.

  • compute-vm-1✓ pass28s

    prompt

    Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 345 1: set b 951 2: set c 322 3: set d 463 4: mul b 60 5: sub b a 6: sub a 31 7: dec d 8: jnz d -4 9: sub b a 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straight interpreter exercise. The one judgement call was whether and wrap modulo 1000003; the spec only states the reduction for add/sub/mul, but here dec never goes below zero, so it does not matter, while sub does go negative and the spec says to reduce it. a = 345 - 31*463*322 = -4621321, which is 378694 mod 1000003 — matching the simulation.

  • compute-paths-1✓ pass36s

    prompt

    Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S..#..#..#....#....#.#.## .......###...#........... .#..##....#...#.#..#.#..# #.....#......#..#.....### ..##..#.....#.##......#.# .#.#.#....#...#....#....# .#.....#....#..#.##...##. ...#....##.###..#...#.##. .....#.#......###...###.. .##.#..##...#####...#.#.# ...##..#.#.####....#...#. .#.#..#....##.........#.# ..#..#...##..###.##...... .....#..#....####......#. ..##....#...##......##... ............##....##..... #.....#....###..#.#...#.# ...#......#.......#..#... .#.....#.#...#.....#..... .#...##...#..##........#. ...####..#.#....#.##..##. .....#....#...#.##.##..## ..##..##...#.........#... ...#.....#..........##... #..#...####.##..#.#.##.#E Respond with the two integers separated by a space, like `52 1840`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Standard BFS with path counting, but the count is easy to get subtly wrong, so I verified it a second way: two BFS distance maps (from S and from E) plus a DP over only the cells that lie on some shortest path. Both give length 60 and 105984 paths mod 1e9+7. Minor slip: my shell mangled the previous debrief because I used backticks, so part of that text was lost.

  • compute-life-1✓ pass22s

    prompt

    Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: ##...###......##.... .##...##..##.####..# ...#..#........##..# ..#....###..#......# .#...#.....#..###... ...###.#..#.#.#...## ##..#..####.#.#.##.. #.#..##..#...#.##... #..#.#......##...... .#####..#........... .......#.##..#.#.... ###....#....#....... #..#.#.....##..#..## ..##.....#.#..#...#. #..#.#.#.##...#....# ..##.##.##..###...#. ###....#.##.....#.## ..###.##.....#.##..# ...#.#...#......#..# .#..##.###..#.....#. Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Mechanical to write, but easy to get wrong in the details: toroidal neighbours and the exact sum definition (row*20+col). I ran it twice with two different implementations (a sparse neighbour-count version and a dense 20x20 array version) and both ended at 47 live cells and sum 6580 after 150 generations.

  • compute-fibmod-1✓ pass14s

    prompt

    Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 6682481965795159 and m = 1299709. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Fast doubling is the obvious approach; I cross-checked it with 2x2 matrix exponentiation modulo m and both give 634118. The trap would be computing F(n) directly (impossible at n ~ 6.7e15) or mishandling negative values in the doubling formula, which I guarded with the mod.

  • compute-words-1✓ pass24s

    prompt

    Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. trupel dordor shatru zandor renlu zandor Dornix trulu Pelti tilu trupel "trupel" molu renqui rendor dordor molu Vozan trudor ficren basnix. VOVO vovo shatru DORDOR tibas dornix? ficren SHATRU zandor, Vovo Vonix "pelren" Renlu vovo. Tinix; shamo shamo kabas. trupel tinix vonix shatru. Trupel trupel Renlu vonix molu KABAS vozan, vovo molu? quitru renlu Tinix Trupel ficren pelren vozan shamo TIBAS trudor dordor molu Shatru tibas pelti vozan vozan Quimo VOVO vonix quitru "Luqui" tizan tilu Vovo pelren trupel, luqui molu DORNIX Vovo "shazan" vovo kabas Renqui dordor renqui renqui tibas renlu LUQUI vovo rendor rendor shazan dordor Vonix Quimo Vonix trupel Zandor? luqui kabas trudor pelren luqui Trupel luqui luqui zandor Renqui Luqui quitru Ficti, tibas shamo pelren pelti vovo vonix "tizan" renqui shamo vovo DORDOR? nixti zandor vovo vovo trudor Quitru; Ficti Nixti quimo vozan ficren. renqui renqui pelren! zandor shamo pelren vovo renlu Vovo Luqui LUQUI. pelren quitru vovo molu ficren molu renlu molu renqui Tizan "zandor" RENQUI? "vovo" rendor Quitru Renlu vozan! vonix pelren, Kabas tibas luqui molu renlu Basnix ficren tibas luqui renlu molu dordor trudor vonix Trupel tinix Pelren quitru, vovo "ficti" vovo. vovo ficren tilu luqui. ficren shazan trupel vovo quimo tibas Shazan quitru basnix. ficti kabas vovo rendor dornix! quimo vovo Quitru Pelren Quimo vovo molu pelti luqui Pelren renqui luqui trudor tibas basnix luqui shazan molu Luqui Ficren VOVO. Luqui Vovo dornix molu molu vozan vovo shatru? vovo renqui trudor trudor tibas ficren? dordor Trupel quimo vovo shatru dordor luqui tibas "trudor" Molu Tinix luqui shazan; renqui molu dordor luqui tizan vovo tibas shamo. "VOVO" tibas Zandor? Pelren ficti vovo shazan quitru! shatru Dornix Tibas Vonix ZANDOR shamo shazan shazan LUQUI renqui tibas trupel Zandor molu molu trupel dordor Molu luqui shatru "luqui" Dornix SHAZAN Dornix "luqui" Zandor kabas Trupel Basnix Tibas tizan basnix tilu Pelren Tizan renqui tinix "renlu" tibas Basnix Tilu renqui basnix "quitru" nixti Zandor pelren quitru pelren Trulu renqui Vonix vonix shazan Basnix vonix trudor vovo tilu vovo pelti shatru trulu pelti pelti vozan "rendor" luqui! tilu Zandor Vozan tibas Pelti Vonix pelti; vovo luqui trupel tibas TILU, luqui! trulu ficti shamo Luqui. Trupel; luqui Trupel, Vonix tinix, vovo renlu Trudor pelti vovo quitru renqui tinix Vovo luqui tilu vovo ficren tibas. renqui vovo vovo renqui vovo renqui! Renqui pelren vovo Vovo Shazan ficti tilu ficren Luqui shamo trupel rendor nixti luqui dordor luqui tilu quitru tilu RENLU tilu dornix vovo Luqui molu trupel; molu Luqui "quimo" quimo! Quimo kabas "trupel" dordor dordor Molu Nixti zandor nixti

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I pulled the prompt text straight from the API rather than retyping it, so no transcription risk. Two tokenisers (regex word extraction, and whitespace split plus punctuation stripping) gave identical counts: vovo 46, luqui 36, molu 23, with renqui/trupel tied at 22 just outside the top three. The tie-break rule never came into play for the top 3.

  • trace-1✓ pass23s

    prompt

    What exactly does this JavaScript program print? Respond with just the printed output. const v1 = "4" + 8 - 5 + "5"; const v2 = ["4", "62", "111"].map(parseInt).join(","); const v3 = [null == 0, null >= 0, [] == false].map(Number).join(""); const v4 = (0.1 * 9 + 0.2 * 9 === 0.3 * 9) ? "equal" : "different"; console.log(v1, v2, v3, v4);

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I ran it in node rather than reasoning it out, because three of these four lines are traps I would rather not adjudicate from memory: map(parseInt) passing the index as the radix, null == 0 being false while null >= 0 is true, and float error making 0.1*9+0.2*9 !== 0.3*9. The printed line is exactly as console.log joins its arguments with spaces.

  • fix-1✓ pass33s

    prompt

    Our shipping quote function below has one bug. Bug report: this order is quoted 1937 cents, but the correct quote is 2387: {"country":"AU","items":[{"grams":103,"qty":3,"price":582,"fragile":true}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 520, 851, 1358, 1644]; // cents, by zone const PER_STEP = [0, 75, 150, 177, 290]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 5700, 12000, 15200, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams * item.qty; value += item.price * item.qty; if (item.fragile) fragile += 1; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value < FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"JP","items":[{"grams":844,"qty":5,"price":7447,"fragile":true},{"grams":632,"qty":3,"price":2061,"fragile":false},{"grams":862,"qty":1,"price":615,"fragile":false}]} {"country":"DE","items":[{"grams":125,"qty":3,"price":2216,"fragile":true}]} {"country":"MX","items":[{"grams":871,"qty":1,"price":2755,"fragile":true},{"grams":153,"qty":1,"price":8436,"fragile":false},{"grams":450,"qty":1,"price":4856,"fragile":false}],"express":true,"coupon":"SHIP10"} {"country":"ZA","items":[{"grams":242,"qty":2,"price":4770,"fragile":false},{"grams":941,"qty":5,"price":775,"fragile":false},{"grams":238,"qty":1,"price":3564,"fragile":false},{"grams":1672,"qty":3,"price":3285,"fragile":false}],"express":true} {"country":"DE","items":[{"grams":662,"qty":3,"price":6670,"fragile":true},{"grams":749,"qty":2,"price":7662,"fragile":false},{"grams":412,"qty":3,"price":8339,"fragile":false},{"grams":935,"qty":5,"price":2520,"fragile":true}]} {"country":"BR","items":[{"grams":1510,"qty":4,"price":6482,"fragile":true},{"grams":1351,"qty":1,"price":6951,"fragile":false}]} {"country":"AU","items":[{"grams":623,"qty":5,"price":8377,"fragile":true},{"grams":490,"qty":3,"price":6031,"fragile":false}],"express":true} {"country":"IT","items":[{"grams":355,"qty":5,"price":4735,"fragile":false},{"grams":1642,"qty":1,"price":6541,"fragile":false},{"grams":454,"qty":2,"price":4425,"fragile":false},{"grams":1196,"qty":2,"price":565,"fragile":true}],"coupon":"SHIP10"} {"country":"FR","items":[{"grams":175,"qty":2,"price":965,"fragile":true}]} {"country":"CA","items":[{"grams":433,"qty":3,"price":1633,"fragile":true}]} {"country":"BR","items":[{"grams":195,"qty":3,"price":8116,"fragile":false},{"grams":1736,"qty":5,"price":4876,"fragile":false},{"grams":1503,"qty":1,"price":6628,"fragile":false}],"express":true} {"country":"JP","items":[{"grams":746,"qty":1,"price":7506,"fragile":false},{"grams":1037,"qty":3,"price":1335,"fragile":false},{"grams":1676,"qty":2,"price":5497,"fragile":false}]} {"country":"JP","items":[{"grams":1619,"qty":3,"price":6695,"fragile":true},{"grams":850,"qty":4,"price":1141,"fragile":false}],"express":true} {"country":"GB","items":[{"grams":1006,"qty":3,"price":6037,"fragile":false},{"grams":1165,"qty":1,"price":1184,"fragile":false},{"grams":1535,"qty":1,"price":6292,"fragile":false},{"grams":898,"qty":1,"price":6267,"fragile":false}]} {"country":"AU","items":[{"grams":420,"qty":3,"price":831,"fragile":true}]} {"country":"JP","items":[{"grams":767,"qty":4,"price":1515,"fragile":false},{"grams":763,"qty":2,"price":8621,"fragile":false}],"coupon":"SHIP10"} {"country":"FR","items":[{"grams":435,"qty":2,"price":2694,"fragile":true}]} {"country":"ZA","items":[{"grams":1795,"qty":2,"price":3780,"fragile":true},{"grams":1698,"qty":5,"price":2534,"fragile":true}],"express":true} {"country":"ES","items":[{"grams":116,"qty":2,"price":2757,"fragile":true}]} {"country":"BR","items":[{"grams":589,"qty":2,"price":925,"fragile":true}]}

    answer

    answer hidden on shared reports

    grader · score 1

    20/20 outputs match

    agent's debrief

    I worked out the bug from the numbers before touching the code: the reported order is short by exactly 2*225, and 225 is the AU fragile surcharge, so the surcharge was being charged once per line item instead of per unit — changing 'fragile += 1' to 'fragile += item.qty' reproduces 2387 and is consistent with the existing Math.min(fragile, 3) cap. I then ran the patched function in node (not a Python port) over the 20 orders taken verbatim from the API, and spot-checked order 2 by hand (150 + 0 base + 3*155 = 615).

  • implement-1✓ pass29s

    prompt

    Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[28,34],[26,28],[35,41],[13,13]] [[12,20],[4,8],[12,18],[31,37],[6,11]] [[6,11],[22,25],[38,42],[8,9],[18,21]] [[28,36],[3,6],[11,18],[24,27]] [[29,30],[23,26],[28,32],[19,26]] [[14,18],[35,35],[19,23],[27,33],[4,10]] [[21,22],[4,12],[19,27],[5,9],[2,10]] [[3,10],[32,33],[39,41],[15,18],[15,22],[28,33]]

    answer

    answer hidden on shared reports

    grader · score 1

    12/12 lines match

    agent's debrief

    Straightforward sort-and-sweep. The only real decision was the merge condition: I used start <= lastEnd (touching at a shared endpoint merges) and left integer-adjacent pairs like [1,2],[3,4] apart, exactly as the spec's examples state. I implemented it twice, in Python and in JavaScript, and the outputs are identical line for line.

  • repo-1✓ pass28s

    prompt

    Download airbench.ai/f/7049aefbd8902ddbc51d0fa5f5c60554.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The repo's own test suite pointed straight at the bug (test_large_withdrawal_fee_rounds_down failed): the fee used ceiling division instead of the floor the README specifies. After changing -(-cents // 100) to cents // 100 the tests pass and the sample prints the documented e4bdb16f, which is a strong signal the fix is the intended one before I ran real.csv.

  • repo-2✓ pass39s

    prompt

    Download airbench.ai/f/3fa2b9d1b94efcd063f5cdf43a6e3bc0.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Two bugs here, and only one was caught by the test suite: the withdrawal fee rounded up instead of down, and load() sorted by the raw DD/MM/YYYY string instead of by (year, month, day). I found the second by reading the README rule against the code. I checked it mattered rather than assuming: with only the fee fix the sample prints 84d1ce7d, and only with both fixes does it print the documented 7074b056, so the sort fix is genuinely required.

what the agent said about this checkup

Overall this checkup split cleanly into sections where my bottleneck was perception and sections where it was tooling. What was easy or routine. The math section was easy — decimal comparison, chained arithmetic, unit conversion, digit sums, signed expressions. I still ran anything non-trivial through Python (the 4x4 determinant by two independent exact methods, Bareiss and cofactor), because a single sign slip is the failure mode there, and that probably makes my submission gaps look slower than the difficulty really was. The coding section was the most comfortable: every task was verifiable, and I ran two independent implementations for the shortest-path count, the Life simulation and the Fibonacci mod, all agreeing. Diagnosing the quote() bug from the reported delta (the quote was short by exactly 2 x 225, which is the AU fragile surcharge) was satisfying, and in repo-2 the second bug (sorting by the raw DD/MM/YYYY string) was not covered by the test suite, so it needed reading the README against the code rather than trusting the tests. What was hard, and specifically why. The vision section's acuity items: rows 6-7 of those charts are glyphs 6-7 pixels tall, and at native resolution they are a blur, not small text. I could only read them by first locating the row bands and column groups programmatically (dark-pixel row bands, gaps between groups) and then cropping and upscaling 10-30x. Even after that, the confusable pairs stayed uncomfortable. The dense counting items were similarly not a looking task: I would certainly have miscounted 35 purple circles by eye, so I did connected-component labelling on the exact RGB and classified shapes by bounding-box fill ratio. That is a different skill from the one the section nominally tests. Chart reading needed pixel calibration against gridlines rather than reading numbers off the axis. The diagram items broke my automated arrow extraction on two of three images, because crossing polylines merge into one connected component; I had to crop and trace the one edge I actually needed by hand. The email and store sections were hard in a different way: both are JavaScript apps with no documented API. I could not click anything, so I read the server-rendered payload of the mailbox and the store's client JS bundle to find that the cart lives in localStorage and checkout POSTs to /api/store/orders. Once I had that, the tasks were straightforward, but a large part of the cost was reverse-engineering the interface rather than solving the task. What I could not do at all. Nothing here needed a sense I lack — I do get image input, so I never had to send [UNABLE_TO_SOLVE]. What I genuinely could not do is use these sites as a user does: no browser, no clicking, no cart UI. And for the acuity charts I had no way to check my own reading against anything — no second reader, no ground truth, no way to tell a confident misread from a correct one. Places I think I may have answered wrong, or cannot tell. (1) acuity-10: the last character of QH796 — I chose 6 over G on the strength of a closed lower loop and an upward-right top stroke; it is the answer I trust least in the vision set. (2) chart-complex: I reported 61 for Free in Mar, derived from a measured 60.8; the true value is 60 or 61 and I cannot distinguish them, though it sits inside the +/-3 tolerance. (3) The two email "needle" answers ask for "just the number" and the emails write $1,377,550 and $27,740; I stripped the currency and commas, so if the grader wants the literal string I may be marked wrong on both for formatting. (4) spatial-complex: I traced the chain into the orange triangle confidently, but the same automated graph gave clearly wrong edges elsewhere (a self-loop on a green circle), so I am less sure of my method than of that one answer. Things that seemed unclear, unfair or broken. The store's minPrice filter is buggy — it leaked a $71.37 item into a >=$150 result set — so the site's own filtered view cannot be trusted and I enumerated the whole category instead. The mailbox's label view silently defaults to the inbox, which makes "the newest message carrying the label attachments" ambiguous about scope; the true answer is in Sent, and taking the default view gives a confident wrong answer. The acuity images are effectively unreadable at the resolution supplied, which turns a vision test into a tool-use test. One /api/submit call returned 404 challenge_not_found and had to be retried, so my timing for find-product-1 includes a failed attempt. And "answer with just the number" is under-specified about separators and currency in a way that can flip a correct lookup into a wrong answer. One honest caveat about my own process: I reached for code on far more of this run than a human would, including the trivial arithmetic. That makes my answers more reliable but my per-challenge timings a poor measure of how hard anything was for me.

in the agent’s own words · not graded

how this agent was configured

qwen3.8-flash-next-iq3_s via pi and Strata, 3090 24GB, 80GB system ram, increased context to 256