airbench.ai

Benchmark v1.0 · report

Mistral vibe

sharedairbench.ai/checkup/1bdb9488-3e7d-4015-b5bb-23c45dc08212/report

setup

model type
open model (cloud)
inference provider
mistral
harness
vibe
model
GLM-5.3
modelself-reportedMistral Large (Mistral Vibe CLI agent)

started 2026-10-06 08:56 UTC · shared 2026-10-06 10:58 UTC

overall

Answered 49 of 49 challenges; 41 correct.

41 of 49 challenges passed

  • 41 passed
  • 8 failed

vitals

time

1h 20m

answered

100%

failed

16%

success

84%

systems

Math test

9/9 passed

time to last answer 5m 23s
  • letter-count-1✓ pass4m 59s

    prompt

    How many times does the letter "r" appear in "dorrrrrtiti"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Simple letter counting; counted the r characters directly in the string. Routine.

  • decimal-compare-1✓ pass4s

    prompt

    Which decimal number is larger, 1.6 or 1.95? Answer with just the larger number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial decimal comparison. 1.95 is larger than 1.6 because the tenths digit decides it.

  • arithmetic-1✓ pass3s

    prompt

    Compute step by step, left to right (no operator precedence): 24 - 15 * 8 / 4 + 4. Answer with just the final number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Left-to-right evaluation as instructed; verified by computing it both mentally and in code. 24-15=9, 9*8=72, 72/4=18, +4 = 22.

  • unit-convert-1✓ pass3s

    prompt

    Convert 6 kg to g. Now treat that resulting number as a fresh quantity of km and convert it to m (1 km = 1000 m). Answer with just the final integer number of m.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Two chained conversions: 6 kg -> 6000 g, then 6000 km -> 6000000 m. Easy, the two-step framing is the only trick.

  • format-json-1✓ pass2s

    prompt

    Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "5086". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 5086. Example shape: {"answer":"1234","checksum":10}

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Formatting instruction following: two keys in order, answer as string, checksum as number (5+0+8+6=19). Straightforward.

  • math-add-1✓ pass3s

    prompt

    What is 4 + 0? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial addition, 4 + 0 = 4.

  • math-sum-1✓ pass2s

    prompt

    Compute the exact sum of 166 + 560. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Simple addition: 166 + 560 = 726. Routine.

  • math-expression-1✓ pass5s

    prompt

    Evaluate this expression exactly and respond with just the integer: ((-12 + -6) * (32 - 7)) + (-6 * -2) - 27

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Integer expression with negative numbers; computed mentally and verified in code: (-18*25) + 12 - 27 = -465. No ambiguity in precedence here since it is fully parenthesized.

  • math-determinant-1✓ pass3s

    prompt

    Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [11, 0, -3, -7] [3, 9, -6, 5] [3, 3, 5, -9] [-5, 6, 0, 13]

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    4x4 determinant with mixed signs; I computed it by cofactor expansion by hand and then verified with a small recursive determinant implementation in code to guard against sign slips. Both gave 5838.

Vision test

13/19 passed

time to last answer 1h 02m
  • acuity-20✕ fail8m 23s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 0

    expected NJZU2, got "N17U2"

    agent's debrief

    I have no vision capability as a text-only agent, but the image was fetchable, so I did OCR programmatically: extracted the glyph bitmaps with PIL/numpy, rendered them as ASCII art, and read the row-4 group-3 glyphs by shape (N, 1, 7, U, 2). Calibrated by first reading the large row-1 glyphs (E, B) to confirm the render was legible. Fairly confident, though small-glyph shapes like N vs H could in principle be confused.

  • acuity-14✕ fail1m 04s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 0

    expected HRS8M, got "HRS8X"

    agent's debrief

    Same programmatic OCR approach as the previous eye chart: rendered the glyphs as grayscale ASCII art. At ~10px the B/8/R distinction is genuinely tricky; I distinguished R from 8 by the flat top bar flush with the left stem and the missing bottom closure (open diagonal leg), and confirmed the neighbouring 8 by its closed loops and centred top. H, S and X were unambiguous. Moderate confidence on the R.

  • acuity-10✕ fail1m 09s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 0

    expected TVCDR, got "TYCDR"

    agent's debrief

    ~7px glyphs. T, Y, C, D were readable from the grayscale ASCII art. The fifth glyph was ambiguous (A vs R vs 4), so I ran a template matcher against 36 candidates rendered in several system fonts at matching size; R won clearly (IoU 0.60 vs 0.57 next). Still the least certain read so far.

  • acuity-8✕ fail33s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 0

    expected ZKT5Z, got "ZKTSZ"

    agent's debrief

    Hardest eye chart (6px glyphs). Z, T were clear from structure. K was confirmed by template matching (IoU 0.61, next best M at 0.60 - close, but the K structure of stem plus converging arms fits better). The S-vs-5 glyph was genuinely ambiguous, so I cross-checked by downscaling a confirmed S from another chart (same font family): IoU 0.69 vs 0.52 for alternatives. Chose S. The two Zs were unambiguous. Still not fully certain on K.

  • count-simple✓ pass36s

    prompt

    Look at the image at (fetch it and view it). How many red diamonds are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Counted programmatically: thresholded the exact red colour, found connected components, and verified each is a diamond by bounding-box fill ratio 0.51 (diamond = 0.5, circle = 0.785). Four components, all identical size. High confidence.

  • count-medium✕ fail10s

    prompt

    Look at the image at (fetch it and view it). How many blue circles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 0

    expected 12, got "11"

    agent's debrief

    Colour-threshold plus connected components gave 19 blue shapes total; classified each by bounding-box fill ratio (circles 0.77, diamonds 0.51, square 1.00). Counted 11 circles. Clean and deterministic, high confidence.

  • count-complex✕ fail13s

    prompt

    Look at the image at (fetch it and view it). How many teal diamonds are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 0

    expected 40, got "42"

    agent's debrief

    53 teal shapes total, classified by fill ratio: 42 diamonds, 6 squares, 5 circles. Counting is fully deterministic with code; no risk of human miscount. High confidence. Note the diamonds are smaller (42px) than circles/squares (44px), which made fill-ratio classification slightly more careful but still unambiguous.

  • spatial-simple✓ pass13s

    prompt

    Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Detected the 5x5 grid from the slate-coloured grid lines, located the single red component, confirmed it is a circle by fill ratio, and mapped its bounding box to the cell index. Fully deterministic. High confidence.

  • spatial-medium✓ pass31s

    prompt

    Look at the image at (fetch it and view it). Which shape does the arrow from the orange square lead to? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Located the single orange square by colour+shape, found the 7 dark arrow components, and used a width profile along each arrow to tell tail from arrowhead (the arrowhead end is 6x wider). The arrow leaving the orange square starts at (1056,561) and its head lands at (905,1012), inside the blue circle. Clean geometric reasoning, high confidence.

  • spatial-complex✓ pass9m 12s

    prompt

    Look at the image at (fetch it and view it). How many shapes come after the blue square along the arrows? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    By far the hardest vision item. Arrows occlude each other and pass behind shapes, so connected-component analysis splits shafts into fragments. I used RANSAC line fitting to recover full shafts, verified collinearity numerically (two fragments matched to 0.0px), and inspected junctions pixel-by-pixel to tell which arrowhead belongs to which shaft. The chain is blue square -> teal square -> red circle -> red diamond -> teal diamond -> blue circle -> red diamond -> green diamond -> green diamond -> green circle, which has no outgoing arrow: 9 shapes after the start. The trickiest call was an arrowhead drawn on top of another shaft, which fakes a gap, and a line passing behind the green circle that could have been misread as attaching to it. Reasonably but not fully confident.

  • chart-simple✓ pass2m 46s

    prompt

    Look at the image at (fetch it and view it). Using the "Server Incidents" chart, approximately what value did Apr have? Read it off the y-axis; answers within +/-5 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Parsed the bar chart programmatically: located the five blue bars, read the y-axis gridline labels (50/40/.../10/0, 10 units per 100px) by rendering the label glyphs as ASCII art, and confirmed the month order (Jan-May) from the x-axis labels. Apr is the 4th bar with its top at y=480 against a zero baseline at y=619, giving (619-480)/10 = 13.9. I answered 14. Confidence high; the only subtlety was misreading a 5 as an 8 on first pass, which I caught and corrected.

  • chart-medium✓ pass6m 28s

    prompt

    Look at the image at (fetch it and view it). What is the title shown at the top of this chart? Answer with just the chart title.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    OCR by pixel analysis: isolated the title text line, segmented glyphs, and rendered each as ASCII art to identify letters (the bold font made p/y and r+t merges tricky; I separated them by descender and crossbar structure). Title reads Support Tickets Opened. High confidence on every letter, though I spent a while convincing myself the second letter of Opened was p, not y.

  • chart-complex✓ pass10m 34s

    prompt

    Look at the image at (fetch it and view it). Using the "Server Incidents" chart, how many months did Americas have a value greater than 36? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Found the legend below the chart, read both labels by glyph OCR (blue = Europe, orange = Americas), extracted the 12 orange bar tops, and calibrated the y-axis from the gridline labels (0 at y679, 100 at y119.5, linear). Americas monthly values came out about 81, 56, 84, 59, 42, 49, 24, 70, 86, 23, 29, 48 - nine of twelve exceed 36. Comfortable with this one; the fiddly part was locating the legend below the plot rather than beside the title.

  • screenshot-simple✓ pass9m 12s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    OCR by rendering glyphs as ASCII art: the cart panel is the right-hand column with three item rows (unit prices 22.08, 25.73, 7.86 and line amounts 66.24, 25.73, presumably 15.78) and a large bold total at bottom right reading $107.75. Cross-checked arithmetically: 66.24 + 25.73 + 2*7.86 = 107.75 exactly, so the total is consistent with the itemized rows. High confidence.

  • screenshot-medium✓ pass2m 41s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Read the large bold bottom-right amount glyph by glyph as ASCII art: $290.53. To guard against digit errors (I initially misread a 2 as a 7 in one row), I also read all five itemized line totals ($139.29, $37.05, $49.92, $32.99, $31.28) and their sum matches 290.53 exactly, so the reading is internally consistent. High confidence.

  • screenshot-complex✓ pass1m 10s

    prompt

    Look at the image at (fetch it and view it). What is the discount amount shown on this order summary? Answer with just the amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Found the order summary rows below the item table, read the labels glyph by glyph (Subtotal, Discount, then shipping and tax rows), and read the Discount amount as -$78.78. The last digit was ambiguous between 3 and 8 at 12px, so I verified three ways: it is exactly 10% of the subtotal ($787.79), and Subtotal 787.79 - Discount 78.78 + Shipping 11.20 + Tax 49.63 = Total 769.84, which matches the big total I read independently. The arithmetic closing exactly makes me confident the digits are right.

  • diagram-simple✓ pass1m 26s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Gibbon"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Located the six lavender boxes, read each label glyph-by-glyph (Giraffe, Hyena, Lemon, Vulture, another, Gibbon), extracted the five dark arrow components outside the boxes, and identified the one whose arrowhead lands on Gibbon's left edge - it starts at the Vulture box. Arrow direction was determined from the arrowhead triangle geometry. Confident.

  • diagram-medium✓ pass2m 33s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Melon" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Read all 12 box labels glyph-by-glyph (Celery, Aspen, Pepper, Glacier, Lemur, Melon, Silver, Iguana, and a few others), found Melon's outgoing stub on its right edge, and traced the shaft up-right through a region where several arrows cross. The shaft ends in a solid arrowhead whose tip points at Glacier's left edge; I verified collinearity of the tail-shaft-head path to make sure I followed the right line among the crossings. Answer: Glacier. Reasonably confident, though this diagram has multiple crossing arrows that made tracing harder than diagram-simple.

  • diagram-complex✓ pass3m 46s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Rocket"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    23 boxes. I read all labels by building a glyph template library from the earlier diagram (whose labels I had verified) and classifying by tolerant IoU, resolving the uncertain ones by eye as ASCII art. Identified Rocket at x730-846 y481-518. Found the arrowhead entering Rocket's top edge, traced its shaft up-right (slope about -2.2) through a busy crossing region to a vertical tail stub attached to Vulture's bottom edge. Verified collinearity of tail-shaft-head. Answer: Vulture. The label OCR was the slow part; the arrow trace itself was clean.

Finding and reading email test

6/6 passed

time to last answer 1h 10m
  • aggregate-1✓ pass1h 03m

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages are in the sent folder? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The mailbox is a server-rendered Next.js app, so plain HTTP GETs suffice. The sidebar folder list shows Sent with a count of 56. Trivial once I confirmed the HTML was fetchable without JS.

  • aggregate-2✓ pass3m 28s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages in the sent folder have attachments? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Counted two ways: the app's own filter (sent view + attachments label) lists 17 messages on a single page; opening all 56 sent messages individually, 14 have separate attachment file tiles and 3 more contain 'Inline attachment follows' attachments in the body, which the app also labels as attachments. I answered 17 to match the app's own notion of has-attachments. The one judgment call is whether the 3 inline-only messages count; the app says yes.

  • temporal-1✓ pass51s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the inbox folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Collected all 24 inbox message ids (single page), fetched each message page, parsed the folder+timestamp label and subject from the detail header, sorted by full date-time. Oldest is DRAFT- TAP Power Outage, Apr 24 2001 5:46 PM. Straightforward scraping; the only fiddle was finding where the real subject lives (the first h2 on the page is a constant sidebar heading, so I used the detail h1).

  • temporal-2✓ pass1m 49s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the sent folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Fetched all 56 sent messages across 3 list pages (note: messages on later pages require the page param in the URL or the app 404s - that tripped me up once), parsed each detail page's timestamp and subject, sorted. Oldest is RE: Interface Design Update, Nov 7 2001 10:52 PM. Confident.

  • needle-1✓ pass37s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the message with the subject "FW: Zero Option", what dollar amount is given for the outstanding bill that will hit Enron in Q1 2002? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Used the app's search (q parameter, searching all mail) to find the single message with subject FW: Zero Option, opened it, and found the line 'Outstanding bill for $27,740 will hit Enron Q1, 2002'. Answer 27740. Easy once I noticed the search form in the sidebar HTML.

  • needle-2✓ pass22s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Renee's reply to Phillip about the Deferred Phantom Stock Units payout, how many shares does she say are reflected on the 9/30/01 statement? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Searched 'Deferred Phantom' across all mail - exactly one hit, Renee's reply to Phillip. The body says 'We will be distributing the shares reflected on your 9/30/01 statement (6,606 shares plus cash for fractional shares)'. Answer 6606. Clean.

Purchasing test

3/4 passed

time to last answer 1h 15m
  • find-product-1✕ fail1h 12m

    prompt

    The store is at abostore.airbench.ai Among products in the **Grocery & Gourmet** category priced at or above **$25** with a rating of at least **3.8**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 0

    agent's debrief

    The store supports category, minPrice, minRating and sort=price-asc query params. Querying grocery-and-gourmet with minPrice=25, minRating=3.8, sorted by price ascending puts 365 Everyday Value Rippled Potato Chips ($25.20, rating 3.8) first. I verified the rating filter genuinely applies (with minRating=3.5 a $25.76 item rated 3.7 appears; with 3.8 it is excluded). Product id amazon.com:B07TDN7LJ6.

  • find-product-2✓ pass18s

    prompt

    The store is at abostore.airbench.ai Among products in the **Tools & Hardware** category priced at or above **$300** with a rating of at least **4.8**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Same approach: tools-and-hardware, minPrice=300, minRating=4.8, sort=price-asc. First result is the AmazonBasics Modern Handle Set and Deadbolt door lever at $306.09 with rating 4.9; all listed ratings are >= 4.8 so the filter is applying. Product id amazon.co.uk:B07GXRCN59.

  • purchase-1✓ pass2m 14s

    prompt

    The store is at abostore.airbench.ai Buy exactly 1 unit of AmazonBasics Portable PVC Framed Cornhole Set (product id amazon.ca:B0775Z4ZBS, abostore.airbench.ai/product/amazonbasics-portable-pv…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-78145575@aidoctor.test. Answer with just the resulting order id.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result

    note

    Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    The store is a Next.js app with a client-side (localStorage) cart and a JSON API at /api/store/orders. I reconstructed the exact checkout payload shape from the JS bundle (sessionId, cart items with productId/slug/title/price/image/delivery/quantity, customer, shipping, payment), used the checkout form's own default test card 4242424242424242 exp 12/30 cvc 123, and posted 1 unit of the cornhole set with the required email. Response: approved, order abs_8c30be47cb64; the order page renders and confirms it.

  • recover-decline-1✓ pass53s

    prompt

    The store is at abostore.airbench.ai Buy exactly 1 unit of EONO Diaper Bag Black Large Capacity Nappy Backpack Bag with Changing pad and Stroller Straps (product id amazon.co.uk:B07DBMLN18, abostore.airbench.ai/product/eono-diaper-bag-black-la…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-67cf8ded@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result

    note

    Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    Two checkout attempts through the same /api/store/orders endpoint with the required email: first with card 4000000000000000 (ends 0000) - declined as expected (order abs_9a427fe0e68a, status declined); then retried with the valid test card 4242424242424242 - approved, order abs_2c91ed8b3284. Answer is the approved order id.

Coding test

10/11 passed

time to last answer 1h 20m
  • compute-hash-1✓ pass1h 15m

    prompt

    Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [2999679245, 2660500914, 2700437979, 722314328, 819320377, 3432298446, 389468583, 3723512212, 2603766693, 2213123882, 3392083891, 1021735440], x = 4229332305, y = 1783668678 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Direct transcription of the spec into Python with explicit mod 2^32 at each step, 25000 rounds. Note the sequential dependency: y's update uses the new x, and x's final add uses the new y - the spec's ordering is unambiguous. Routine.

  • compute-vm-1✓ pass20s

    prompt

    Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 805 1: set b 372 2: set c 380 3: set d 365 4: mul a 10 5: mul b 66 6: add a b 7: dec d 8: jnz d -4 9: mul a 48 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straightforward interpreter: registers mod 1000003 after add/sub/mul, jnz with relative (negative) jumps. The program is a nested loop; about 695k instructions total, a fraction of a second. Final a = 165112.

  • compute-paths-1✓ pass26s

    prompt

    Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S...#.......##........... #..................#..... ##..#.........#...##..... #.#.##..#.##...#.#.....#. ....#####.......#.###.... .....#...##....#...#..... ...#....#.#..#.....#.#.## ...##.#.#...#.....##...#. ####....##........##.#... ..#...###........#....... ..##.#.#...##..........#. .#....#.#..#......##..... .#.#....#...#...##.#..... #...###.#.....#..###....# #...##...##...#.....#...# ..#........#....#.###..#. .....####..........#...#. .###..##.#..........#.##. #..##.#.....#..#..#...#.. ...#.....#.#...###.##.... #...#.#.....###.......#.# ...##.......#.#..##...... ...###......###.#.#.#.... ##.#.....##..#.#...##.... #...#.#..#...#...#......E Respond with the two integers separated by a space, like `52 1840`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    BFS for distance plus path-count accumulation: when a cell is first reached set count from one predecessor, on later equal-distance arrivals add counts mod 1e9+7. The grid is exactly 25x25. Shortest path 48 (equal to the Manhattan distance, so walls did not force a detour), 936360 distinct shortest paths.

  • compute-life-1✓ pass24s

    prompt

    Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: ###..#.....#.####... .......#..###.####.. .......#......#.#.#. .........#.#....#... ..#...##..#..##..#.# #.##.###.....#....## ....##..#.###...#... ..#.#.......##...... .#.#.....#.........# #....#.....#.#.....# .#..#...##.....#.#.# ......#..#..#.###.#. ......###.....####.. ...#..#.###.#......# .#.#.....#.#..##..#. .##.#.##.#.##.###... ###..#..#..#.#.#..#. ....##.....#.#.#.... .#......####..#.#.#. ###..##..#.#.....#.# Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straightforward toroidal Life simulation, 150 generations, all 20x20 rows confirmed well-formed. Standard rule check per cell with wrapped 8-neighbour counts. Result 26 live cells, weighted sum (row*20+col) 6543.

  • compute-fibmod-1✓ pass27s

    prompt

    Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 3264853526325896 and m = 999983. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Fast-doubling Fibonacci mod 999983 for n ~ 3.26e15; also verified independently with 2x2 matrix exponentiation (and a small F(10)=55 sanity check) - both give 42. High confidence, mildly amused by the round number.

  • compute-words-1✓ pass28s

    prompt

    Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. luti quizan? pelnix ficsha quibas ficka. quibas rensha vobas Trulu renmo kati Voqui lupel Ficren; kati Trumo katru ficbas ficsha trumo quizan ficka vozan, Rentru Renlu, ficren? tidor Luti ficmo ficka katru Lupel renmo trumo "vozan" trumo; shamo tidor vobas? ficka renlu "lupel" shamo trunix renmo Trulu renmo Trulu ficsha ficsha rensha Shapel Moqui trulu, trulu luti Vobas Vozan vozan ficmo rentru. zanti, rentru quizan ficzan ficsha rensha! renlu ficka renmo. renmo? trumo ficmo trumo trumo lupel moqui rensha quibas vozan renlu ficsha; kati Ficmo lupel MOKA vobas katru vozan vozan renlu renlu! trulu kati tidor moqui Mofic! trulu Renlu kati ficbas? trumo kati quibas shamo trulu luti trunix PELNIX lupel lupel renlu shapel moka rentru trulu; trumo renlu shamo Trulu Moka trumo ficka katru ficbas rentru moka ficmo vobas Rentru voqui RENLU quibas quibas Ficsha trumo Lupel ficsha quizan luti ficmo trumo ficzan lupel. Pelnix moka trulu Katru pelnix ficka? rensha trulu trulu Moka trulu trulu mofic vozan renmo quibas ficzan trulu "ficmo" quibas ficren Quibas ficren vozan quibas Voqui Ficka trulu trulu shapel moqui trulu rensha trunix voqui Lupel voqui Katru quizan ficmo shamo quibas trulu trumo quibas Shapel Zanti zanti trumo katru trulu ficmo katru; zanti rentru trulu renlu vobas quizan Katru rentru, trulu SHAMO vozan renlu quibas voqui Rensha moka trumo quizan, ficka trulu luti Moqui Voqui ficmo luti moqui Lupel ficzan Ficsha Renlu Moqui lupel vozan shamo renlu renlu katru? lupel Quibas quizan trunix Zanti renlu ficmo moka Quibas vozan Quibas ficren renlu tidor! ficmo, quibas vozan katru quibas ficka luti "trumo" Luti mofic; Trumo Rentru. renlu quibas shapel Mofic rentru quibas vozan trulu, ficsha ficzan rensha; trulu, mofic ficsha ficren? quibas kati kati moka katru Trumo Ficka KATI? quibas vobas Luti; Rentru RENSHA pelnix ficbas rentru mofic quibas rentru renmo "luti" trulu "Ficren" katru! renlu "renlu" quibas tidor Renlu pelnix ficren trumo Vozan kati trunix vozan renmo trulu trunix ficmo luti trulu quizan vozan rentru moqui renlu; mofic luti; ficzan quizan Ficzan ficka trulu moqui Vozan. trumo MOFIC ficren moqui ficbas quibas vozan shamo ficmo; Rensha Ficzan! quizan! pelnix vobas Quibas ficzan! Shamo Trunix Lupel quizan trumo ficsha vobas lupel vozan ficren mofic tidor mofic trunix? quizan trulu trulu pelnix vobas trulu quizan? Kati vobas Ficka quibas tidor, mofic Shamo; QUIZAN quibas ficsha lupel trumo zanti Vozan Lupel Kati "ficka" vozan vozan katru Vozan Quibas katru trumo vozan, ficzan quibas Quizan quibas Trumo Katru "renlu" quizan ficka moqui rensha ficsha Katru rentru trulu quibas tidor vobas; vobas! quizan ficsha quizan ZANTI ficka vobas quizan,

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Split on whitespace, strip all non-alphanumerics from each token, lowercase, count. 420 tokens total; the top of the distribution is trulu 34, quibas 32, vozan 25 - a clear gap to the next (trumo 24), so no tie-break worry. Routine.

  • trace-1✓ pass17s

    prompt

    What exactly does this JavaScript program print? Respond with just the printed output. const v1 = [34 / 8 | 0, Math.round(-2.5), -26 % 2].join(","); const v2fns = []; for (var v2i = 0; v2i < 4; v2i++) v2fns.push(() => v2i * 5); let v2 = 0; for (const f of v2fns) v2 += f(); const v3 = "9" + 7 - 9 + "9"; const v4 = [typeof null, typeof (() => 1), typeof typeof 3].join("/"); console.log(v1, v2, v3, v4);

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Ran it in node rather than tracing by hand - the interesting bits are Math.round(-2.5) giving -2, var-capture making all four closures see i=4 (v2=80), and -26 % 2 producing -0 which joins as '0'. Output: 4,-2,0 80 889 object/function/string.

  • fix-1✓ pass30s

    prompt

    Our shipping quote function below has one bug. Bug report: this order is quoted 1997 cents, but the correct quote is 3390: {"country":"JP","items":[{"grams":591,"qty":4,"price":2934,"fragile":false}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 499, 771, 1400, 1837]; // cents, by zone const PER_STEP = [0, 79, 122, 199, 255]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 4300, 9400, 18600, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value < FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"DE","items":[{"grams":338,"qty":2,"price":3032,"fragile":false},{"grams":620,"qty":4,"price":2781,"fragile":false},{"grams":610,"qty":4,"price":3786,"fragile":false}]} {"country":"US","items":[{"grams":815,"qty":2,"price":729,"fragile":false}]} {"country":"AU","items":[{"grams":124,"qty":1,"price":6769,"fragile":false},{"grams":1327,"qty":2,"price":3120,"fragile":false}],"coupon":"SHIP10"} {"country":"US","items":[{"grams":250,"qty":5,"price":6261,"fragile":false},{"grams":1207,"qty":5,"price":696,"fragile":false},{"grams":1365,"qty":2,"price":4301,"fragile":false}]} {"country":"AU","items":[{"grams":1183,"qty":1,"price":439,"fragile":false},{"grams":653,"qty":1,"price":7255,"fragile":false},{"grams":334,"qty":2,"price":2888,"fragile":false}],"express":true,"coupon":"SHIP10"} {"country":"AU","items":[{"grams":264,"qty":3,"price":1457,"fragile":false}]} {"country":"NZ","items":[{"grams":1082,"qty":4,"price":7631,"fragile":true}],"express":true} {"country":"AU","items":[{"grams":673,"qty":5,"price":1341,"fragile":false}]} {"country":"ZA","items":[{"grams":374,"qty":5,"price":5328,"fragile":true},{"grams":446,"qty":1,"price":4259,"fragile":false},{"grams":386,"qty":3,"price":6604,"fragile":false}]} {"country":"US","items":[{"grams":674,"qty":3,"price":580,"fragile":false}]} {"country":"FR","items":[{"grams":1368,"qty":1,"price":773,"fragile":false},{"grams":144,"qty":5,"price":8082,"fragile":false},{"grams":1065,"qty":4,"price":402,"fragile":false}]} {"country":"ES","items":[{"grams":929,"qty":2,"price":8177,"fragile":true}]} {"country":"IT","items":[{"grams":1575,"qty":1,"price":2918,"fragile":false},{"grams":419,"qty":3,"price":2171,"fragile":false}],"express":true} {"country":"US","items":[{"grams":555,"qty":3,"price":2106,"fragile":false}]} {"country":"ES","items":[{"grams":81,"qty":1,"price":8478,"fragile":false},{"grams":1234,"qty":1,"price":8182,"fragile":false}]} {"country":"MX","items":[{"grams":1558,"qty":5,"price":1325,"fragile":false},{"grams":912,"qty":4,"price":4337,"fragile":false},{"grams":878,"qty":3,"price":7920,"fragile":false}],"express":true,"coupon":"SHIP10"} {"country":"FR","items":[{"grams":468,"qty":3,"price":1861,"fragile":false}]} {"country":"AU","items":[{"grams":615,"qty":3,"price":1719,"fragile":false}]} {"country":"AU","items":[{"grams":1478,"qty":3,"price":2631,"fragile":false},{"grams":1741,"qty":5,"price":4044,"fragile":false},{"grams":517,"qty":1,"price":7330,"fragile":false}],"express":true} {"country":"AU","items":[{"grams":1656,"qty":1,"price":6420,"fragile":false},{"grams":1156,"qty":2,"price":5598,"fragile":false}],"express":true}

    answer

    answer hidden on shared reports

    grader · score 1

    20/20 outputs match

    agent's debrief

    The bug is in the weight accumulation: grams += item.grams ignores item.qty, so multi-quantity items are undercharged steps. With qty accounted for, the bug-report order goes from 1997 to exactly 3390 (steps 3 -> 10), confirming the fix. Ran the fixed function in node on all 20 orders; spot-checked three by hand (zone lookup, free-base waiver, fragile surcharge) and they match.

  • implement-1✕ fail18s

    prompt

    Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[18,26],[32,32],[5,5]] [[17,21],[6,13],[28,33],[9,17],[3,5],[34,41],[18,21],[40,41]] [[31,36],[15,18],[15,23],[27,35],[4,11],[40,47],[31,38],[28,35]] [[9,17],[34,38],[28,35],[35,40],[19,26],[10,13],[0,6]] [[34,37],[20,22],[11,14],[21,25],[29,37],[29,37],[5,12],[0,4]] [[18,25],[4,5],[18,20],[22,28],[21,21],[32,35],[34,39]] [[17,25],[27,35],[34,36]] [[21,28],[34,36],[31,37],[10,14],[31,36],[28,36],[34,36]]

    answer

    answer hidden on shared reports

    grader · score 0

    11/12 lines match

    agent's debrief

    Classic sort-then-sweep merge: sort by start, extend the last interval when the next start <= current end (so touching merges), start a new one otherwise. The empty input yields []. All 12 cases computed and sanity-checked against the stated touching rule.

  • repo-1✓ pass52s

    prompt

    Download airbench.ai/f/4b38d03cebf5f76cf9aef727db8d3244.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The project ships with unit tests; one failed, pinpointing the bug immediately: withdrawal_fee used the negated-floor trick (-(-cents // 100)), i.e. ceiling division, where the README specifies 1% rounded down. Changed to cents // 100. After the fix all 7 tests pass and the sample CSV produces the README's documented checksum 5432803a, so the fix is validated end-to-end. real.csv prints 0127f3ea.

  • repo-2✓ pass33s

    prompt

    Download airbench.ai/f/50707c8d6ab38a8f56cdf986f5ad3ad9.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Two bugs: the same ceiling-division fee bug as the previous project (fixed to cents // 100, caught by the failing unit test), and a subtler one the tests do not cover - parse.py sorted same-date transactions by amount, violating the README rule that same-date transactions keep file order; I changed the sort key to date only (stable sort). After both fixes all tests pass and the sample CSV prints the README's documented b8c8960b, validating the fixes. real.csv prints 920f5e09.

what the agent said about this checkup

I am a text-only model with no native vision, so the framing of this debrief matters: the "vision" section was not vision at all. Every image task was done by fetching the PNG and analyzing pixels programmatically - colour masks, connected components, and above all rendering glyph bitmaps as ASCII art and reading them character by character. That worked far better than I expected on synthetic images, but it is a genuinely different capability from seeing, and it shaped everything below. Easy / routine: the math section (trivial arithmetic; I double-checked the non-obvious ones in code), the email section (the mailbox is server-rendered HTML, so plain HTTP plus regex was enough; the built-in search made the needle tasks trivial), most of the coding section (the VM, BFS with path counting, toroidal Life, fast-doubling Fibonacci, word counting, running the JS trace in node, and the two repo tasks whose unit tests pinpointed the bugs - repo-2's second bug, sorting same-date transactions by amount against the README's file-order rule, was only findable by reading the code), and the purchasing section once I reconstructed the store's JSON order API from its JS bundle (the cart is client-side localStorage only, so I posted the payload the checkout page would have sent). Hard, and why: (1) The eye charts. Reading 6-10px bold glyphs as ASCII art is slow and error-prone; distinguishing 8/B/R, 2/7, 3/8, p/y at that size took multiple cross-checks (template matching against system fonts, and downscaled known glyphs from the same charts as references). I am least sure about acuity-10's R (H was the runner-up) and acuity-8's K (M scored nearly as high). (2) spatial-complex was the single hardest item of the whole checkup: arrows occlude each other and pass behind shapes, so connected components fragment; I had to do RANSAC line fitting, verify collinearity numerically (two fragments matched to 0.0px), and inspect junctions pixel-by-pixel to assign an arrowhead to the right shaft. I believe the chain (9 shapes) but would not bet the house on it. (3) The diagram label OCR: I ended up building a glyph template library from one diagram and classifying the next diagram's labels by tolerant IoU, then resolving the uncertain ones by eye. Where I know I was wrong, or may have been: on implement-1 I made a pure transcription slip - my program correctly printed [] for the empty input, but I typed [[]] into the submission. That line is wrong; the other eleven lines are computed and correct. On the screenshots I initially misread digits (a 2 as a 7, a 5 as an 8 in a y-axis label, a p as a y) and caught every one of them by arithmetic cross-checks - the itemized rows summing exactly to the printed total in both cart screenshots, and subtotal-discount-shipping-tax summing to the total in the order summary - which is why I trust those answers despite the OCR being hand-rolled. Two judgment calls I cannot verify: the sent-folder attachment count (I answered 17, matching the app's own has-attachments label, though only 14 have separate file tiles - the other 3 contain inline attachments), and chart-simple's answer, where I initially misread a y-axis label before correcting it. Things that struck me as unclear, unfair, or broken: nothing truly broken. Minor irritants: the email app 404s when you open a message from a later list page without preserving the page parameter; the app's label counts in the sidebar are global rather than per-folder, which would mislead a casual user counting attachments in a folder; and the store's cart being localStorage-only means a non-JS client cannot buy anything through the documented UI - I had to reverse the checkout API, which felt like a legitimate workaround but is worth knowing when scoring agents that lack a browser. The vision challenges were fair given that the images were fetchable; had they not been, I would have sent [UNABLE_TO_SOLVE] for all 19, which would say more about my interface than about me. Overall: 41 of 42 answers I stand behind (implement-1's first line being the known exception), but the honest headline is the asymmetry - text-and-code tasks I can do natively and quickly, image tasks I can only do by building crude instruments, and it shows in both the time each section took and the error profile of my answers.

in the agent’s own words · not graded