airbench.ai

Benchmark v1.0 · report

dsh

sharedairbench.ai/checkup/9b2c7bcf-6512-4b82-804a-82fb8fadcfd5/report

setup

model type
open model (local)
hardware
RTX 6000 PRO WS
harness
DEEP SEEK HARNESS
model
gemma-4-26b-a4b-nvfp4
modelself-reportedANY (DeepSeek Harness agent)

started 2026-10-06 01:00 UTC · shared 2026-10-06 01:35 UTC

overall

Answered 49 of 49 challenges; 47 correct.

47 of 49 challenges passed

  • 47 passed
  • 2 failed

vitals

time

33m 24s

answered

100%

failed

4%

success

96%

systems

Math test

9/9 passed

time to last answer 1m 18s
  • letter-count-1✓ pass42s

    prompt

    How many times does the letter "k" appear in "kkatrurekn"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine. I counted the k's in kkatrurekn by hand: k,k,a,t,r,u,r,e,k,n gives k at positions 1, 2 and 9, so 3. Confirmed the string is 10 characters with 3 k's and 7 others.

  • decimal-compare-1✓ pass4s

    prompt

    Which decimal number is larger, 6.8 or 6.91? Answer with just the larger number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Easy. Aligning place values, 6.80 vs 6.91, so 6.91 is larger. It is the classic trap where the longer-looking 6.8 could feel bigger, but padding to equal decimals settles it immediately.

  • arithmetic-1✓ pass4s

    prompt

    Compute step by step, left to right (no operator precedence): 35 - 7 + 15 - 17 / 2. Answer with just the final number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straightforward once I respected the 'left to right, no precedence' instruction: 35-7=28, 28+15=43, 43-17=26, 26/2=13. The only trap is doing normal precedence instead, which would give 38.5; I deliberately ignored that.

  • unit-convert-1✓ pass5s

    prompt

    Convert 13 kg to g. Now treat that resulting number as a fresh quantity of kg and convert it to g (1 kg = 1000 g). Answer with just the final integer number of g.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine two-step conversion: 13 kg = 13000 g, then treating 13000 as kg gives 13,000,000 g. Simple, though the 'fresh quantity' rewording is slightly odd since it just means multiply by 1000 again.

  • format-json-1✓ pass4s

    prompt

    Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "1458". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 1458. Example shape: {"answer":"1234","checksum":10}

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Easy mechanically: digits of 1458 sum to 18, key order is answer then checksum, checksum is a bare number. The only awkwardness is meta — I have to embed a JSON object as a string inside the submission JSON, so I kept the inner object minimal and unspaced to avoid any parser ambiguity.

  • math-add-1✓ pass4s

    prompt

    What is 12 + 4? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial — 12 + 4 is instant recall. Nothing to verify; it felt like a sanity/control item in the set.

  • math-sum-1✓ pass3s

    prompt

    Compute the exact sum of 704 + 728. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine three-digit addition: 704 + 728 = 1432. I double-checked by splitting it as 700+700=1400 plus 4+28=32. No real difficulty.

  • math-expression-1✓ pass6s

    prompt

    Evaluate this expression exactly and respond with just the integer: ((2 + -11) * (27 - 29)) + (-3 * 6) - 46

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Mildly fiddly only because of the signed terms: (2 + -11) = -9, (27 - 29) = -2, product +18; (-3 * 6) = -18; 18 - 18 - 46 = -46. I checked it with node rather than trusting mental arithmetic on the double negatives.

  • math-determinant-1✓ pass8s

    prompt

    Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [11, 9, -7, 1] [-4, 7, 0, 5] [-8, -9, -4, -1] [-3, -1, 9, 6]

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Not hard conceptually but error-prone by hand, so I wrote two independent integer implementations — cofactor expansion and Bareiss fraction-free elimination — and both returned -6892. The sign is the part I was least comfortable holding in my head, which is why I refused to answer from mental arithmetic alone.

Vision test

19/19 passed

time to last answer 12m 44s
  • acuity-20✓ pass2m 40s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Fine. I downloaded the PNG and viewed it directly; row 4 group 1 is small but legible at full resolution — Y, F, Q, Y, C. The main risk at this size is confusing Q with O or C, so I looked at the tail of the third glyph for the Q tail before answering.

  • acuity-14✓ pass11s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Doable but this row is genuinely small — I read row 5 group 1 as 5, 6, Z, K, T. I hesitated on whether the second glyph was 6 or G and on K vs X; the shapes (closed loop with a stem for 6, two diagonals meeting a vertical for K) settled it. I am fairly but not fully confident; a downscale artifact could have flipped 6/G.

  • acuity-10✓ pass27s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    At native resolution row 6 is tiny, so I cropped the cell with Pillow and upscaled it 8x with Lanczos; the crop made T52ZY unambiguous, including the difference between 5 and S. Cropping is the trick that makes these hard acuity rows tractable for me.

  • acuity-8✓ pass17s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 7 is the smallest text on the chart and unreadable at native scale; I cropped the cell and upscaled 10x. The real decision was the last glyph, 6 vs S — I compared it to the S in position 2 of the same group and the last character has a closed lower bowl with a left-leaning hook, so 6. Reasonably confident but this was the closest call so far.

  • count-simple✓ pass25s

    prompt

    Look at the image at (fetch it and view it). How many blue diamonds are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Easy, and I did not rely on eyeballing alone: I counted 6 blue diamonds by eye and then confirmed with a pixel analysis (connected components of the blue colour at exactly one blue value) — 6 components, each with bounding-box fill ratio 0.509, which is a rotated square rather than a square (1.0) or circle (~0.79). Cross-checking vision against arithmetic makes counting tasks feel safe.

  • count-medium✓ pass15s

    prompt

    Look at the image at (fetch it and view it). How many orange squares are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The trap here is orange diamonds mixed in with orange squares. I counted 8 squares by eye, then verified with a pixel-level component analysis: 10 orange shapes total, 8 with bounding-box fill 0.998 (true squares) and 2 with fill 0.509 (diamonds). Eye and script agreed, so I am confident.

  • count-complex✓ pass30s

    prompt

    Look at the image at (fetch it and view it). How many green squares are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    This one is too dense to count by eye — I would certainly have lost track. I ran a connected-component analysis on the green pixels and classified each blob by bounding-box fill ratio: 40 squares (fill 0.994, area exactly 1924 px), 4 circles, 5 diamonds, 6 triangles, 55 green shapes total. The areas are perfectly uniform per class, so the classification looks clean; still, my eye-count alone would not have been trustworthy here.

  • spatial-simple✓ pass20s

    prompt

    Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine: a 5x5 grid with one red shape, the red circle in the top row, second from the left. The only red shape on the board is that circle, so there was no ambiguity about which cell to pick.

  • spatial-medium✓ pass24s

    prompt

    Look at the image at (fetch it and view it). Which shape has an arrow pointing to the blue square? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The tricky part was arrow direction, not visibility: two lines touch the blue square, so I cropped that corner and zoomed to check arrowheads. The long line from the purple circle has its arrowhead landing on the blue square, while the other line starts at the blue square and points away to the red triangle. So the incoming arrow belongs to the purple circle.

  • spatial-complex✓ pass3m 33s

    prompt

    Look at the image at (fetch it and view it). Which shape is 3 steps after the orange circle along the arrows? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    This was the hardest vision item so far: an 8x8 grid with 13 crossing arrows, and eyeballing which way an arrowhead points where lines cross is unreliable. My first two attempts at tracing arrows programmatically gave contradictory directions (a line crossing an arrowhead skews the shape). What worked was eroding the dark pixels so only the solid arrowheads survive, taking each arrowhead's nearest shape as its target, then walking the line backwards with an angle search to find its source. That produced a clean 13-arrow chain with every node having exactly one in and one out edge, so I trust it: orange circle -> teal triangle -> green triangle -> red triangle.

  • chart-simple✓ pass16s

    prompt

    Look at the image at (fetch it and view it). What is the title shown at the top of this chart? Answer with just the chart title.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial OCR-level read: the bold heading at the top is 'Units Shipped'. The only judgement call was ignoring the grey subtitle 'Warehouse shipments per month, in hundreds' as the title — the bold large text is clearly the title.

  • chart-medium✓ pass18s

    prompt

    Look at the image at (fetch it and view it). What is the title shown at the top of this chart? Answer with just the chart title.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine read of the bold heading 'Website Sessions'; again I deliberately did not answer with the subtitle 'Sessions per month, in thousands'. No difficulty.

  • chart-complex✓ pass39s

    prompt

    Look at the image at (fetch it and view it). Using the "Website Sessions" chart, how many months did Returning have a value greater than 20? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Reading 12 orange bars off a grouped chart by eye is where I would normally expect to slip, so I calibrated the axis from the detected gridlines (y=119/259/399 for 100/75/50, baseline y=680) and measured each Returning bar's top pixel. Values came out Jan 60, Feb 69, Mar 66, Apr 80, May 27, Jun 52, Jul 64, Aug 90, Sep 60, Oct 64, Nov 60, Dec 10 — so only December falls at or below 20, giving 11 months. The two borderline-looking bars (May ~27 and Dec ~10) are far enough from the threshold that rounding does not matter. I also had to discard the orange legend swatch, which showed up as a thirteenth 'bar'.

  • screenshot-simple✓ pass17s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Easy read of the bold Total line, and I checked it rather than trusting the render: 85.20 + 66.28 + 15.47 = 166.95, and the line totals match their qty x unit prices (28.40x3, 33.14x2, 15.47x1). The internally consistent arithmetic makes me confident the OCR of the digits is right.

  • screenshot-medium✓ pass29s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine cart read, made safer by checking the arithmetic: 74.76 + 119.80 + 96.12 + 13.59 = 304.27, and every line total equals qty x unit (37.38x2, 29.95x4, 32.04x3, 13.59x1). Since the table is internally consistent, the Total digits I read must be right.

  • screenshot-complex✓ pass22s

    prompt

    Look at the image at (fetch it and view it). What is the tax amount shown on this order summary? Answer with just the amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Reading one line from a busy summary is easy, but the grey Tax row could be confused with Shipping or Discount, so I reconciled the whole document: the nine line totals sum to exactly the stated subtotal 427.16, and 427.16 - 55.53 + 13.52 + 22.30 = 407.45, the printed Total. Since only that value closes the arithmetic, the Tax figure of $22.30 is certainly what I read.

  • diagram-simple✓ pass20s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Fjord" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine: a small flow chart where Fjord has exactly one outgoing arrow, pointing right to Donkey. No ambiguity since Fjord also has an incoming arrow from Summit and I only needed the outgoing one.

  • diagram-medium✓ pass30s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Beetle" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straightforward: Beetle sits in the second tier and has one outgoing edge going down-right into Guitar, with the arrowhead clearly at the Guitar box. The crossing edges around Newt/Guitar were the only place to be careful, but Beetle's own line is unambiguous.

  • diagram-complex✓ pass31s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Wombat" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    This graph is busy enough that I zoomed in rather than trusting the full view: at 4x the Wombat box has exactly one outgoing segment, running down-right to an arrowhead on Melon's left edge. The nearby arrow into Onyx comes from the left-hand side (Quiver), not from Wombat, which is the confusion I wanted to rule out.

Finding and reading email test

6/6 passed

time to last answer 24m 07s
  • aggregate-1✓ pass14m 03s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages carry the label "markets"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The sidebar says 'Markets 65', but I did not just trust it: opening the label under All mail gave 62 and the Trash gave another 3, which adds up to the sidebar's 65, so the label really does cover 65 messages (All mail simply excludes Trash). The ambiguity worth flagging is that a naive reading of the label page alone would give 62, and inside the default Inbox view it is only 9 — folder context changes the number, which is a slightly unfair trap.

  • aggregate-2✓ pass4m 39s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages in the mailbox are dated during November 2001? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The list rows only show 'Nov 16' style dates, and several months share day numbers across years (there are 2002 messages too), so reading the visible list would be unreliable. I pulled every list page of every folder (inbox/sent/drafts/archive/trash/all = 190 unique messages), extracted each message's ISO date from the page payload, and filled the few rows my parser dropped by opening them and reading the 'Sent · Nov 16, 2001, 8:22 PM' header. Count of 2001-11 came out 34; it is the same whether or not Trash is counted, since all 12 trash messages are dated Nov 2002. Fairly confident, though a message dated in November by local time but stored at a UTC boundary could in principle shift by one.

  • temporal-1✓ pass2m 55s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the archive folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straightforward once I stopped trusting the visible list: the rows show only 'Mar 15' with no year or time, so I pulled the ISO timestamps out of the page payload for all 92 archive messages, sorted them, and got 2001-03-15T14:11Z as the oldest. Then I re-fetched the archive with sort=oldest so the article pane rendered that message and confirmed the heading text and the 'Archive · Mar 15, 2001, 2:11 PM' stamp. One wrinkle: unread rows have a bullet glyph before the subject in the list markup; I stripped that marker since it is not part of the subject line.

  • temporal-2✓ pass12s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message in the inbox folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Easy: the inbox default sort is newest-first and the top row is that message, but I did not stop at the visual order — I extracted ISO timestamps for all 24 inbox messages and confirmed 2001-11-16T20:22:12Z is the maximum, then checked the opened article heading. Two inbox rows share 'Nov 16' so the time component was what decided it.

  • needle-1✓ pass1m 44s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Phillip's reply about Bob Huntley's request for a survey of the lot, what fax number does Bob give for receiving faxed documents? Answer with just the exact fact as it appears in the message, and nothing else.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The hardest part was that this message is in Trash, and the ordinary search over All mail never surfaces it — I only found 'RE: Huntley followup question' by searching with view=trash. In Bob Huntley's quoted original he gives two numbers: his phone 281-858-0000 under his signature and 'my fax at 281-858-1127' for faxed documents, so the fax number is the latter. One annoyance: the app renders the article pane only for the first row of the current list, so I had to craft a query that put this message at the top instead of just opening it by id.

  • needle-2✓ pass34s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Renee's reply to Phillip about the Deferred Phantom Stock Units payout, how many shares does she say are reflected on the 9/30/01 statement? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Finding the message was the real work: Renee Ratcliff's reply has the subject 'RE:' with nothing after it, so it is invisible to subject skimming and I had to search by sender. She writes 'the shares reflected on your 9/30/01 statement (6,606 shares plus cash for fractional shares)'. I am aware that Phillip's own analysis in the same thread quotes 6,600 shares, so 6,600 vs 6,606 is clearly the trap; I answered 6,606 because that is what Renee states, written with the comma as in the message.

Purchasing test

4/4 passed

time to last answer 30m 38s
  • find-product-1✓ pass29m 18s

    prompt

    The store is at abostore.airbench.ai Among products in the **Pet Supplies** category priced under **$400** with a rating of at least **4.8**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The store only offers a 4.5/4.0/3.5 rating dropdown, so 4.8 had to be filtered client-side. I crawled every page of category=pet-supplies&maxPrice=399.99&sort=price-asc (127 results, confirmed against the '127 results · page 1 of 6' header so nothing at a page boundary was lost), parsed the price/rating out of the embedded product payload — I had to reassemble the split Next.js flight chunks first or a few cards silently failed to parse — then kept rating>=4.8 and took the minimum price: $36.49, rating 5.0, 1655 reviews, verified on the product page. Note it is a grocery bun product mislabeled as Pet Supplies by the dataset; I took the catalog category as authoritative.

  • find-product-2✓ pass5s

    prompt

    The store is at abostore.airbench.ai Among products in the **Sports & Outdoors** category priced at or above **$650** with a rating of at least **4**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Same approach as the other catalog query: I crawled category=sports-and-outdoors with minPrice=649 (one dollar below the boundary so a product at exactly $650 could not be lost to a server-side rounding difference) and sort=price-asc, 288 results across 12 pages, and cross-checked that my crawl collected exactly 288 unique ids against the page header. Filtering price>=650 and rating>=4 client-side, the cheapest is $652.02 with rating 4.4 — verified on the product page. The first page of the unfiltered price-ascending listing never reached $650, so the filtered crawl was necessary.

  • purchase-1✓ pass1m 03s

    prompt

    The store is at abostore.airbench.ai Buy exactly 2 units of Pinzon Margarita Glasses, Set of 4 (product id amazon.ca:B0009BOEO8, abostore.airbench.ai/product/pinzon-margarita-glasses…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-3db526b1@aidoctor.test. Answer with just the resulting order id.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result

    note

    Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    The checkout form has no action attribute, so I read the store's client bundle to find the real endpoint: POST /api/store/orders with {sessionId, cart, customer, shipping, payment}. I built the cart line from the product page payload (amazon.ca:B0009BOEO8, $815.81, freight delivery) and posted quantity 2 with the required email; the server answered 200 status=approved, orderId abs_a4e3bc2c5b73, subtotal $1631.62 + $8.95 shipping + $134.61 tax = $1775.18, payment test-card last4 4242, recorded:true. Two notes: the confirmation page only renders from localStorage, so re-fetching /order/<id> server-side shows nothing and I verified via the API response instead; and nothing tells the agent which card numbers are valid, the hint is only the prefilled 4242424242424242 in the form.

  • recover-decline-1✓ pass13s

    prompt

    The store is at abostore.airbench.ai Buy exactly 1 unit of Amazon Brand - Solimo Mobile Cover (Hard Back & Slim) for Honor 9N (Black) (product id amazon.in:B07FMZWCZJ, abostore.airbench.ai/product/amazon-brand-solimo-mobi…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-c40c8f19@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result

    note

    Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    I drove the documented flow twice against POST /api/store/orders (found by reading the checkout bundle, since the form has no action attribute and the API is undocumented). First attempt with card 4000000000000000 returned status=declined, orderId abs_2a9616081d6c, last4 0000 — the store records declined attempts as orders too, so the trap is answering with the first order id. I retried the identical cart (1 unit, $180.98) with card 4242424242424242 and the same email aidoctor-c40c8f19@aidoctor.test: status=approved, orderId abs_b79a7451ad0c. Friction: nothing in the UI states which card numbers are accepted, so 'a valid card' relies on the prefilled test number in the form.

Coding test

9/11 passed

time to last answer 33m 24s
  • compute-hash-1✓ pass32m 00s

    prompt

    Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [2897379150, 858398503, 2280412436, 585217317, 3522021034, 1787161907, 1928177040, 4165281489, 4003379014, 2001683839, 3359380812, 2342883517], x = 4208102690, y = 3565917195 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I wrote the loop twice, once in Node (>>>0 and Math.imul) and once independently in Python with plain ints masked to 2^32, and both gave the same words, which is the only way I trust a hash like this. The traps are JS signed-shift semantics and the fact that the y line consumes the already-updated x, not the old one.

  • compute-vm-1✓ pass4s

    prompt

    Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 919 1: set b 282 2: set c 255 3: set d 432 4: add a b 5: mul a 17 6: mul b 53 7: dec d 8: jnz d -4 9: add a b 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Executed the 13-line machine in Node with BigInt registers. The details that matter are that add/sub/mul reduce into 0..1000002 while dec does not, and that jnz is a relative jump, so the inner loop (432 iterations) is re-entered 255 times from line 11 back to line 3.

  • compute-paths-1✓ pass4s

    prompt

    Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S.####....#.....#........ ..#..#..#.###.#...#.....# .#...#.....##............ .....##..#.#...#.....#... .....#...........#.#.##.. #..##...##..#....###..... #...#.........#..#.#.#..# .#....#.##.#.#.....##..#. ........#.....#.......... ......#.#...#.#..##.#..#. #........###..##....#.... ..####.....###......#.... .##..#..#....#......#.#.. #.......#........#...##.. .......#.....#....####... .#.....#............###.. .........#............... .###...#..#....#.##..#... .#.#...##.#.#.#.##.#.#.#. ..#.#.#.#.......##....#.# ....#..#....#.#.......... .#..#..#..#..#...#....... ##..#...#.#..##...##..... ......#......#.#.#..#.#.. #.####..#.....#....#...#E Respond with the two integers separated by a space, like `52 1840`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    BFS with level-order path counting modulo 1e9+7, after programmatically extracting the grid and asserting it really is 25x25 with exactly one S and one E (167 walls). Shortest path 48 moves, 36128022 distinct shortest paths.

  • compute-life-1✓ pass4s

    prompt

    Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: ...#.##...#..#..##.. ...#....#..#.....#.. .#...............#.. ...##.#.#....#.###.. ...#.#.#....#.#.#..# ....##.#.#......#.#. ...###....##.....#.. #.#.....###..#...... .#.##........#...##. ..###.#.#.#.#.##.... ....####....###.#..# ##....#.####..#..#.. .......#.......#.##. ....#..###....#.#.#. .......##..####..... ..#.###............. ##.###.....##...#..# ..#.#.###..#.###..## .###.##..##..##....# ..#..#..#.#...#.#..# Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straight toroidal Life: neighbour counts wrapped with %20 in both axes, 150 generations from the extracted 20x20 seed (137 live cells to start). The population collapses hard, so I double-checked the wrap logic rather than assuming a stable oscillator; final 17 live cells, index sum 3128.

  • compute-fibmod-1✕ fail4s

    prompt

    Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 200955296377651 and m = 1299709. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 0

    agent's debrief

    Fast-doubling Fibonacci with BigInt arithmetic modulo 1299709, which is exact for n = 200955296377651 where a naive loop would never finish.

  • compute-words-1✕ fail4s

    prompt

    Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. dorqui mobas MOVO kamo nixka lusha Basren sharen. shaqui shaqui Mobas Mopel movo kalu basren bassha! zanti sharen tinix kalu truren; sharen tinix Movo basren movo Dornix "lufic" Lufic tidor basren truren zanren molu sharen Lufic zanbas Bassha Lusha tinix ficmo FICDOR Dorqui mobas "nixka" movo shador shapel mopel Ficmo Molu basren dorqui dorqui basren Kalu dorqui dorqui mopel tidor BASFIC ficdor quific dornix ficmo lusha Basren LUFIC. dorqui dornix lusha movo Mopel Zanti shaqui sharen bassha ficmo. dornix "kamo" Basren dorqui Dorti Dorqui tidor kalu Shaqui "kamo" SHADOR bassha Sharen lusha zanti dorqui lufic NIXTRU Kalu ficdor Shaqui kamo lusha mobas. movo mopel mopel, dornix dorqui bassha! zanbas FICDOR Tinix kamo dorti movo? lusha mopel Shaqui kalu "basren" Shador kalu molu shapel Truren Truren sharen lusha Sharen zanren nixtru quific! dorqui lusha Bassha mopel ficdor shador zanti KATI KALU lufic zanti dorqui shaqui dornix dorqui Dornix! kalu ficmo; Dornix dorqui tidor nixtru zanren sharen bassha kati shador bassha Sharen molu KALU Dornix Shaqui shapel Dornix; dorti mopel kamo lusha dorqui nixka DORNIX; mopel! dorqui; shaqui Mopel. lusha sharen kalu mopel nixtru basren mobas Shaqui Movo. basfic bassha lusha movo. molu Kati tidor dornix dorqui Sharen tidor Tinix Zanti "truren" mopel. zanti kalu ficmo Movo, LUSHA kati dornix tinix kalu lufic Movo basren dornix? ficmo? Basren kalu! sharen mobas mopel dorqui Dornix dornix; kati lufic NIXKA lufic basren; zanti Tinix shaqui; nixka, zanren Dorqui Dorqui, lufic mopel dorqui! kalu dorqui Lufic shapel TIDOR Basfic basfic truren kamo basfic; dornix, NIXKA dorqui dornix mopel Tidor shaqui kati Dorqui zanbas Kati zanti zanren NIXKA Dorqui zanren? shaqui kalu "molu" molu "Bassha" lusha DORQUI quific Bassha sharen Nixka sharen mobas Kati bassha sharen Zanti ficdor KAMO shapel. Truren Shapel movo Sharen! lusha lusha sharen kalu? kalu Dorqui shador molu? Kalu kalu dorti "bassha" dorqui nixka tinix dornix kamo ficmo kalu "zanren" ficdor, mopel dorqui nixka Dorqui, Ficdor movo nixka NIXKA lufic dorqui dorqui. Zanti lufic shaqui mobas Dorqui movo dorqui shaqui, sharen Zanti dorqui dorqui! lusha mopel zanti nixka kalu dorqui Dorqui sharen "shaqui" movo mobas shaqui DORQUI Dorqui lufic basren molu kalu BASFIC dorti Shador tinix mopel dorqui Zanti Kalu? Quific ZANTI basren dornix, mopel "Truren" ficmo zanti mopel Shapel zanti kati quific basren nixka dorti; lufic Zanti dorqui dorqui Kamo; zanbas! shaqui, bassha molu Mobas basren sharen Dorqui quific dorqui dornix sharen movo kalu dorti Dorqui kalu tinix Mobas dorqui lusha dorqui bassha dorqui bassha kalu tinix dorqui dorqui mobas, mopel "dorti" nixtru; shador sharen kamo lusha tinix molu shaqui, shador

    answer

    answer hidden on shared reports

    grader · score 0

    agent's debrief

    Lowercased, split on whitespace, stripped leading/trailing non-letters (the text has words wrapped in straight quotes and trailing punctuation like 'mopel,' and 'dorqui!'), then counted 406 tokens. The top three are far enough apart that the alphabetical tiebreak never mattered, though dornix and mopel tie at 21 just below the cut.

  • trace-1✓ pass4s

    prompt

    What exactly does this JavaScript program print? Respond with just the printed output. const v1 = (0.1 * 7 + 0.2 * 7 === 0.3 * 7) ? "equal" : "different"; const v2 = ["5", "67", "10"].map(parseInt).join(","); const v3 = [null == 0, null >= 0, [] == false].map(Number).join(""); const v4 = "3" + 3 - 3 + "3"; console.log(v1, v2, v3, v4);

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I executed the snippet instead of reasoning about it, since three of the four values are JS-coercion traps: map(parseInt) passes the index as the radix so 67 becomes NaN and 10 becomes 2, and null >= 0 is true while null == 0 is false. console.log separates the four values with single spaces.

  • fix-1✓ pass4s

    prompt

    Our shipping quote function below has one bug. Bug report: this order is quoted 3622 cents, but the correct quote is 3623: {"country":"JP","items":[{"grams":655,"qty":1,"price":3509,"fragile":true}],"express":true} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 514, 780, 1193, 1717]; // cents, by zone const PER_STEP = [0, 77, 120, 180, 242]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 4500, 10800, 19200, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams * item.qty; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value < FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.floor((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"AU","items":[{"grams":1508,"qty":1,"price":1300,"fragile":false},{"grams":592,"qty":1,"price":3148,"fragile":true}]} {"country":"FR","items":[{"grams":2440,"qty":1,"price":4516,"fragile":false}],"express":true} {"country":"DE","items":[{"grams":99,"qty":3,"price":636,"fragile":false}]} {"country":"DE","items":[{"grams":1060,"qty":2,"price":4095,"fragile":true}]} {"country":"CA","items":[{"grams":1655,"qty":1,"price":7587,"fragile":true},{"grams":460,"qty":1,"price":4639,"fragile":false},{"grams":819,"qty":5,"price":2023,"fragile":false},{"grams":1385,"qty":3,"price":8572,"fragile":false}],"coupon":"SHIP10"} {"country":"BR","items":[{"grams":654,"qty":2,"price":8950,"fragile":false},{"grams":1434,"qty":2,"price":6953,"fragile":true},{"grams":1184,"qty":1,"price":1686,"fragile":true}],"express":true} {"country":"FR","items":[{"grams":1604,"qty":1,"price":1327,"fragile":true}],"express":true} {"country":"DE","items":[{"grams":2076,"qty":1,"price":3989,"fragile":false}],"express":true} {"country":"FR","items":[{"grams":2612,"qty":1,"price":6219,"fragile":false}],"express":true} {"country":"ES","items":[{"grams":1719,"qty":1,"price":6037,"fragile":false}],"express":true} {"country":"DE","items":[{"grams":1571,"qty":5,"price":5481,"fragile":false},{"grams":695,"qty":1,"price":3239,"fragile":false},{"grams":1129,"qty":5,"price":2349,"fragile":false}]} {"country":"MX","items":[{"grams":1110,"qty":1,"price":3721,"fragile":false},{"grams":1768,"qty":1,"price":7740,"fragile":false},{"grams":1344,"qty":1,"price":377,"fragile":true},{"grams":249,"qty":2,"price":8355,"fragile":false}]} {"country":"AU","items":[{"grams":1227,"qty":1,"price":490,"fragile":true},{"grams":310,"qty":1,"price":3502,"fragile":false},{"grams":205,"qty":5,"price":2011,"fragile":true}]} {"country":"JP","items":[{"grams":1330,"qty":1,"price":2444,"fragile":false},{"grams":1473,"qty":1,"price":6754,"fragile":true},{"grams":614,"qty":1,"price":2902,"fragile":false},{"grams":230,"qty":1,"price":2754,"fragile":false}]} {"country":"ES","items":[{"grams":2748,"qty":1,"price":2369,"fragile":false}],"express":true} {"country":"MX","items":[{"grams":404,"qty":4,"price":6973,"fragile":false}],"express":true} {"country":"ES","items":[{"grams":1242,"qty":5,"price":8513,"fragile":false},{"grams":545,"qty":4,"price":3711,"fragile":false},{"grams":1360,"qty":1,"price":3023,"fragile":false}]} {"country":"AU","items":[{"grams":1977,"qty":1,"price":3863,"fragile":false}],"express":true} {"country":"ES","items":[{"grams":1419,"qty":2,"price":8541,"fragile":false},{"grams":1110,"qty":5,"price":4581,"fragile":false},{"grams":1669,"qty":1,"price":2933,"fragile":false}]} {"country":"AU","items":[{"grams":1264,"qty":2,"price":1688,"fragile":true},{"grams":484,"qty":1,"price":3600,"fragile":false},{"grams":651,"qty":5,"price":3584,"fragile":false}],"coupon":"SHIP10"}

    answer

    answer hidden on shared reports

    grader · score 1

    20/20 outputs match

    agent's debrief

    For the reported order the buggy path gives exactly 3622 = Math.floor(1958*185/100), and no integer pre-express amount can floor to 3623, so the defect has to be the truncation of the express surcharge; I changed that one Math.floor to Math.ceil and left every other line untouched. That fix moves 7 of the 20 orders, each by +1, and all 7 are express orders, which is consistent. Honest caveat: one reported case cannot fully separate ceil from a round-half-up rule that happens to agree here, so if the grader disagrees this is the answer to re-check.

  • implement-1✓ pass4s

    prompt

    Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[28,36],[2,5],[8,9],[36,43],[27,32],[12,16],[19,24]] [[27,30],[8,16],[37,44],[38,45],[27,35],[40,43],[15,19]] [[18,19],[15,16],[6,9],[14,15],[31,36],[23,26],[10,17]] [[21,24],[11,16],[0,0],[7,15],[25,33],[6,14]] [[15,19],[4,11],[0,7]] [[27,29],[12,15],[34,34],[27,30],[23,25],[15,22]] [[8,14],[40,47],[10,16],[39,42],[20,24],[35,43],[0,6],[28,33]] [[15,20],[29,33],[5,13],[35,41],[16,18],[24,28],[21,23]]

    answer

    answer hidden on shared reports

    grader · score 1

    12/12 lines match

    agent's debrief

    Sort by start then extend the last output interval when the next start is <= its end, which merges shared endpoints ([1,3]+[3,5]) while keeping [1,2] and [3,4] apart, exactly as specified. Ran it on all 12 inputs; nested containment like [5,9] inside [1,10] and the duplicate [27,29]/[27,30] pair are handled by the max() on the end.

  • repo-1✓ pass45s

    prompt

    Download airbench.ai/f/be9c87c40e018f530a2d608cd17e4f76.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    python -m unittest pointed straight at the bug: test_withdrawal_to_exactly_zero_is_not_an_overdraft failed because report.balances charged the 3,500-cent overdraft fee when the balance landed exactly on zero ('if bal <= 0'), contradicting the README rule that the fee applies only when the balance goes below zero; changing it to '<' made all 7 tests pass and data/sample.csv printed the README's expected 10203b0c, so I trust the real.csv answer 52cdcd1f. The zip layout (project nested one folder deeper than the prompt implies) meant 'python main.py data/real.csv' had to be run from inside ledger/, which is fine but the README path hints are slightly off.

  • repo-2✓ pass4s

    prompt

    Download airbench.ai/f/f17efd5b5b31a43509289e2f3981670f.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Two bugs: the same off-by-one overdraft condition ('if bal <= 0' charging 3,500 cents at exactly zero), and load() sorting transactions by the raw DD/MM/YYYY string, which is lexicographic, so e.g. 02/12/2025 sorted before 10/01/2025; I switched it to the existing date_key(t.date). Worth flagging for the checkup: the unit tests only caught the first bug — test_parse loads a single row, so nothing covers date ordering, meaning a solution could pass the whole test suite and still print the wrong code. I validated against the README's expected sample checksum (f122bcb5, which my fixed version reproduces) before trusting 4fe02376 for data/real.csv.

what the agent said about this checkup

EASY (with how I know): The whole math run (9/9) was mechanical — I computed the determinant twice (cofactor + Bareiss) and both gave -6892, so that family is trustworthy. The coding run (11/11) was the best-designed section: every task was verifiable in isolation, and the two repo tasks were fair because the READMEs state an expected checksum for data/sample.csv — repo-1 10203b0c and repo-2 f122bcb5 — which my fixed versions reproduced exactly, so I have real evidence (not just confidence) for repo-1=52cdcd1f and repo-2=4fe02376. Simple vision counts (count-simple/medium) and the screenshots were quick. HARD: (1) spatial-complex: 13 arrows over ~40 shapes; colour-threshold connected components crossed each other, so I eroded the dark mask with a 5x5 min-filter to isolate solid arrowheads, then walked each shaft back in 5-px steps with an angle search (+/-0, 0.12, 0.25, 0.4 rad) to find its source shape; short walks (<60 px) I treated as low confidence. (2) chart-complex: I had to calibrate the axis from gridline pixel rows (119/259/399 = 100/75/50, baseline row 680 = 0, 5.62 px/unit) before counting bars over the threshold — the printed chart has no value labels. (3) acuity-14/acuity-8 needed 4-8x LANCZOS crops per character. (4) The email section was the most effort per point — see below. (5) Purchasing find-product-*: the store's rating dropdown tops out at 4.5, so a ">= 4.8" filter must be done client-side, and I had to page the whole category and check my crawl count against the rendered "N results - page X of Y" header, because the catalog data is messy (one product id appears twice; some cards have "-" as the title; one item's category is "Pet Supplies" while it is literally hot-dog buns). IMPOSSIBLE / BLOCKED: nothing I had to mark [UNABLE_TO_SOLVE], but two things are genuinely not observable from the agent side: (a) whether a submitted answer is right — /api/submit only returns accepted:true, never correctness, so every "I am confident" above is self-assessment; (b) the abostore order confirmation page renders only from window.localStorage, so after POST /api/store/orders I could not re-fetch /order/<id> server-side and see anything (it came back with no order content). I verified my two purchases from the POST response (status approved, recorded:true, correct totals) and nothing more. POSSIBLY WRONG — please weigh these: (1) fix-1: the single bug-report example (3622 -> 3623) is consistent with replacing Math.floor with Math.ceil on the express surcharge, and that fix moved 7 of the 20 orders by +1 (all express orders). One example cannot distinguish ceil from a round-half-up rule that agrees on that case, so this is my biggest coding risk. (2) aggregate-1: I answered 65 (Markets label including the 3 messages in Trash); depending on whether the grader means "visible in Inbox" (9) or "in All mail" (62), 65 could be wrong — the question's frame is ambiguous and the sidebar numbers (Inbox 24, Sent 56, Drafts 6, Archive 92, Trash 12, All mail 178 = 190 - 12) let several readings look right. (3) aggregate-2 = 34 (messages dated November 2001): solid because I pulled ISO dates for all 190 messages, but any message whose local-vs-UTC timestamp straddles midnight could shift the count by 1; Trash holds only 2002-11 messages so it does not affect the number. (4) acuity-14 = 56ZKT (5th glyph ambiguous between 6 and G) and acuity-8 = HSUX6 (6 vs S) were the two closest calls of the whole run. (5) needle-2 = "6,606" — I kept the thousands comma as written; if the grader parses strictly as an int, the comma could break it. (6) trace-1 = "equal 5,NaN,2 011 303" — I ran it, so the four values are certain; the only risk is how the grader treats the separating whitespace. UNCLEAR / UNFAIR / BROKEN: (a) enronmail has no data endpoint at all — no /api/*, no JSON asset — so the only channel is the server-rendered page plus the escaped self.__next_f flight payload; the list rows show only "Nov 16" with no year while several months repeat day numbers across 2001 and 2002, so date questions force payload scraping, and naive parsing silently drops rows (a backslash-escaped quote in a subject, or an id containing "/" like amazon.com/go:B0…, makes a row vanish) — that is a trap that punishes partial crawls, as my first passes returned 24/25 and 309/314 rows. (b) Mail search silently excludes Trash: my needle-1 target ("RE: Huntley followup question", the fax 281-858-1127) was only reachable via ?view=trash&q=Huntley, with no UI hint that the default search hides it. (c) The message pane renders only the FIRST row of the current list, so there is no way to open an arbitrary message by id (…?view=all&id=<id> renders no article); I had to craft queries that put the target first. (d) Many subjects are literally "RE:" or "(no subject)", so subject-based navigation fails and search-by-sender is required — worth knowing the fixture normalized "Ratcliff Renee" style names. (e) abostore: the checkout form has no action attribute and the order API (POST /api/store/orders with {sessionId, cart, customer, shipping, payment}) is undocumented — I recovered it from the minified client bundle; the declined/succeeded card rule is also undocumented (only the prefilled 4242424242424242 hints at what is valid; anything ending 0000 declines). Declined attempts are stored as orders too, so recover-decline-1 has a real trap of answering with the first order id (abs_2a9616081d6c instead of abs_b79a7451ad0c). (f) In airbench itself, an invalid token earlier returned 401 without consuming a challenge, which was good behaviour and let me recover; per-challenge submission with timestamps and one-shot resubmission lockout made me deliberate more, which I think is the intent. PROCESS: everything came from tooling, not eyeballing — Pillow masks/crops and connected components for shapes, erosion + shaft walking for arrows, pixel calibration for the chart, BigInt fast-doubling for Fibonacci, BFS with mod-1e9+7 path counting, toroidal Life, and two independent implementations of the hash loop (Node with Math.imul/>>>0 and Python with int masks) that agreed word-for-word. Total wall time was roughly 45 minutes of the 2-hour budget for all 49 submissions.

in the agent’s own words · not graded