airbench.ai

Benchmark v1.0 · report

opencode/openrouter/mimo-v2.6-flash

setup

model type
open model (cloud)
inference provider
openrouter
harness
opencode
model
mimo-v2.6-flash
modelself-reportedmimo-v2.6-flash

started 2026-10-07 07:14 UTC

overall

Answered 49 of 49 challenges; 49 correct.

49 of 49 challenges passed

  • 49 passed

vitals

time

58m 24s

answered

100%

failed

0%

success

100%

systems

Math test

9/9 passed

time to last answer 1m 59s
  • letter-count-1✓ pass35s

    prompt

    How many times does the letter "a" appear in "dorbasaaren"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial letter count; I verified it with a quick script rather than by eye, which felt routine.

  • decimal-compare-1✓ pass4s

    prompt

    Which decimal number is larger, 7.53 or 7.2? Answer with just the larger number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straightforward decimal comparison, 7.53 vs 7.2 — no ambiguity, answered immediately.

  • arithmetic-1✓ pass5s

    prompt

    Compute step by step, left to right (no operator precedence): 49 / 7 - 12 - 6 + 17. Answer with just the final number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Left-to-right evaluation was clearly flagged as no-precedence, so I just chained the ops; easy but easy to get wrong if you apply normal precedence.

  • unit-convert-1✓ pass15s

    prompt

    Convert 3 GB to MB. Now treat that resulting number as a fresh quantity of hours and convert it to minutes (1 hours = 60 minutes). Answer with just the final integer number of minutes.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    GB to MB is ambiguous (1000 vs 1024); I took the decimal convention, 3000 MB -> 180000 minutes, but I am not fully sure that matches the intended answer.

  • format-json-1✓ pass10s

    prompt

    Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "5365". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 5365. Example shape: {"answer":"1234","checksum":10}

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Pure format drill: keys in order, number type not string, digit sum 5+3+6+5=19. Routine once I read the shape example carefully.

  • math-add-1✓ pass7s

    prompt

    What is 5 + 19? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial addition, answered instantly.

  • math-sum-1✓ pass15s

    prompt

    Compute the exact sum of 666 + 462. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine addition; I double-checked the carry mentally. Easy.

  • math-expression-1✓ pass10s

    prompt

    Evaluate this expression exactly and respond with just the integer: ((3 + 11) * (11 - 18)) + (-6 * 10) - 20

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Nested arithmetic with negatives; straightforward but I took a moment to trace each bracket. -178.

  • math-determinant-1✓ pass17s

    prompt

    Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [-1, -9, 7, 0] [-8, -4, 9, 4] [6, -8, 6, 4] [-2, -5, 4, -4]

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    4x4 determinant is not something to do by hand reliably, so I computed it two independent ways (cofactor expansion and Bareiss) and both gave 1140.

Vision test

19/19 passed

time to last answer 35m 07s
  • acuity-20✓ pass2m 37s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Read the chart directly, then cropped and magnified row 4 group 2 to confirm; the zoom made it unambiguous.

  • acuity-14✓ pass23s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Same routine as before: cropped row 5 group 2 and magnified it; 3A95F was clear at 6x.

  • acuity-10✓ pass44s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 6 is small but still legible; a 8x crop of group 3 confirmed GAF85.

  • acuity-8✓ pass1m 07s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Smallest row on the chart; needed a 16x zoom to separate the letters, but 3FPPB was clear once magnified.

  • count-simple✓ pass1m 28s

    prompt

    Look at the image at (fetch it and view it). How many red diamonds are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I counted six; I cross-checked by isolating pure-red components programmatically, which also returned 6 (an orange triangle had initially slipped into a looser red filter).

  • count-medium✓ pass42s

    prompt

    Look at the image at (fetch it and view it). How many orange triangles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Counting by eye was error-prone here because orange diamonds and circles share the palette; I segmented orange blobs and separated triangles from diamonds by their bottom-row width, getting 13.

  • count-complex✓ pass1m 19s

    prompt

    Look at the image at (fetch it and view it). How many orange triangles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Too dense to eyeball reliably, so I segmented orange blobs and classified each by shape signature (bottom-row width); 30 triangles plus 4 diamonds, 3 circles and 1 square, which I then spot-checked against the picture.

  • spatial-simple✓ pass45s

    prompt

    Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Easy: the only red object in the 5x5 grid is the circle in the bottom row, third column.

  • spatial-medium✓ pass2m 14s

    prompt

    Look at the image at (fetch it and view it). Which shape does the arrow from the blue triangle lead to? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Several arrows cross this grid, so I traced dark line components programmatically and confirmed by zooming: the line leaving the blue triangle ends in an arrowhead pointing at the green circle.

  • spatial-complex✓ pass12m 58s

    prompt

    Look at the image at (fetch it and view it). Which shape is 3 steps after the orange diamond along the arrows? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    This needed real tracing: I segmented the dark arrow strokes, fit each line, and located arrowheads pixel-by-pixel. Chain: orange diamond -> red diamond -> purple triangle -> orange triangle.

  • chart-simple✓ pass21s

    prompt

    Look at the image at (fetch it and view it). What is the title shown at the top of this chart? Answer with just the chart title.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine: the bold title at top-left reads Website Sessions (the subtitle is separate).

  • chart-medium✓ pass31s

    prompt

    Look at the image at (fetch it and view it). Using the "Units Shipped" chart, how many months had a value greater than 30? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Read all eight bars against the gridlines; only July (~18) is under 30, so seven months qualify.

  • chart-complex✓ pass1m 10s

    prompt

    Look at the image at (fetch it and view it). Using the "Units Shipped" chart, approximately what value did Mobile have in Apr? Read it off the y-axis; answers within +/-3 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I measured the bar pixel top against the gridlines rather than eyeballing it: Apr Mobile comes out at ~36.9, so I answered 37.

  • screenshot-simple✓ pass1m 03s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straight read-off of the bold Total; the three line totals also sum to 111.99, so I was confident.

  • screenshot-medium✓ pass33s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Same layout as the simple one; five line items summing exactly to the printed total, so this was routine.

  • screenshot-complex✓ pass50s

    prompt

    Look at the image at (fetch it and view it). What is the tax amount shown on this order summary? Answer with just the amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Denser summary with subtotal/discount/shipping/tax rows; Tax reads 6.30 and the arithmetic to the total checks out.

  • diagram-simple✓ pass28s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Walrus" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Simple tree diagram; the single arrow leaving Walrus goes straight down to Moose.

  • diagram-medium✓ pass36s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Trout"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Only one arrow arrives at Trout, coming from Fjord; the other Fjord edges go to Celery and Marble.

  • diagram-complex✓ pass5m 19s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Prism" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The diagram is dense, so I walked the stroke pixel-by-pixel from Prism's right edge: east, up, east along the top, then down-right past Piano to an arrowhead on Vortex's left border.

Finding and reading email test

6/6 passed

time to last answer 42m 09s
  • aggregate-1✓ pass41m 11s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages include gthorse@keyad.com in the To field? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I scraped every message's To header across all folders: 12 include gthorse@keyad.com (3 of them in Trash). I counted the whole mailbox rather than just the All mail view, since another answer in this section lives in Trash.

  • aggregate-2✓ pass11s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages are in the sent folder? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straightforward: the Sent folder header and my own scrape of its pages both give 56 messages.

  • temporal-1✓ pass5s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the sent folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Sorted Sent oldest-first and cross-checked by parsing every message's full timestamp; the earliest is Nov 7, 2001 with that subject.

  • temporal-2✓ pass7s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the archive folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Archive sorted oldest-first puts Mar 15, 2001 first; I confirmed with the per-message timestamps as well.

  • needle-1✓ pass11s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Phillip's reply about Bob Huntley's request for a survey of the lot, what fax number does Bob give for receiving faxed documents? Answer with just the exact fact as it appears in the message, and nothing else.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Found Phillip's reply ('I do have a survey of the sight at home') and read Bob's quoted original underneath; his fax for documents is 281-858-1127 (his direct line 281-858-0000 sits just below, which was a near-miss trap).

  • needle-2✓ pass24s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Renee's reply to Phillip about the Deferred Phantom Stock Units payout, how many shares does she say are reflected on the 9/30/01 statement? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Renee Ratcliff's reply states the shares reflected on the 9/30/01 statement are 6,606 (plus cash for fractions); I answered just the number.

Purchasing test

4/4 passed

time to last answer 48m 45s
  • find-product-1✓ pass44m 34s

    prompt

    The store is at abostore.airbench.ai Among products in the **Fashion** category priced under **$50** with a rating of at least **4.2**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Filtered Fashion to <=0 and rating >=4.2 (22 hits), sorted price ascending, and took the first: The Drop Preston Belt Bag at .87, 4.8 stars. ID pulled from the product page.

  • find-product-2✓ pass17s

    prompt

    The store is at abostore.airbench.ai Among products in the **Automotive** category priced under **$300** with a rating of at least **4.2**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Same approach for Automotive under 00 with rating >=4.2: only 3 hits, cheapest is the AmazonBasics car vacuum at 5.46.

  • purchase-1✓ pass2m 59s

    prompt

    The store is at abostore.airbench.ai Buy exactly 2 units of AmazonBasics Compostable Food Trays, 5 Compartment, 500-Count (product id amazon.ca:B075HQHQPP, abostore.airbench.ai/product/amazonbasics-compostable…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-e77f0a2b@aidoctor.test. Answer with just the resulting order id.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result

    note

    Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    Posted the cart (qty 2) plus shipping/contact/payment to the store's orders endpoint; it returned approved with this order id.

  • recover-decline-1✓ pass54s

    prompt

    The store is at abostore.airbench.ai Buy exactly 1 unit of AmazonBasics Weatherproof Outdoor Patio String Lights S14 Bulb, Green, 48' (product id amazon.ca:B073WG69TY, abostore.airbench.ai/product/amazonbasics-weatherproo…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-dc3d0533@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result>checkout_result

    note

    Store events recorded for this order's session: checkout_result → checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    First checkout with a card ending 0000 came back declined, then the retry with a different card was approved; this is the approved order id.

Coding test

11/11 passed

time to last answer 58m 24s
  • compute-hash-1✓ pass49m 24s

    prompt

    Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [4023667284, 3407936357, 3773747690, 195740019, 3870889680, 4268038417, 1502305926, 3250808255, 714752652, 3278794493, 497029218, 194330699], x = 3302697352, y = 2849442089 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Implemented the described32-bit unsigned loop exactly in Python (masking after every op) and printed x then y as 8-digit hex.

  • compute-vm-1✓ pass1m 04s

    prompt

    Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 188 1: set b 142 2: set c 363 3: set d 351 4: add a 50 5: sub a 14 6: mul a 85 7: dec d 8: jnz d -4 9: mul b 92 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Wrote an interpreter: jnz lands on pc+k (the other convention never halts), add/sub/mul reduce mod 1000003. Simulated to halt: a = 703279.

  • compute-paths-1✓ pass1m 29s

    prompt

    Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S#.......#......##...##.. ....##...#..##.#.#.....## ####..#.#.#...#.##....... #.#.....##..#.....##.#..# ##..#.#..#.#..##.#......# .#.#...#.#.#.....#.##.... ..#.#....##..........#..# ##.....#.#..###.###.#...# ..#.#...#................ ........##.....###.#.#... #.........#..#....###.... ..#.....##.....#.#....#.# .#.#.....#.#....#...##.#. .#.#....#..#......##.#... ##..###............#..#.. .#..##...........#.....#. ...##..#...#...#...##.... ........#.....#..#.#.#.#. .#..#....#......####.#..# ..##...#.#...#....###.#.. ...#.#...##...##....##... .#...#......##.###.#.#..# #.#.#....#.....#.#...#..# .#.##....##..#......#...# ...#....#..#...##..#..#.E Respond with the two integers separated by a space, like `52 1840`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    BFS for the distance and a second DP pass over distance layers for the count; both methods agreed.

  • compute-life-1✓ pass1m 13s

    prompt

    Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: ##.#.##..##.#.....#. ##...#.#........###. ...#..###....#...... ....######...#.#.... ......##.###.....##. .....##...#.......## ###...###....##.##.. ......#.#####....... ........###...###..# ...#..##..#...#..#.# .#.##...#....#...... ##.####........#...# ...##.#..#.#.#.#.... .....##........#...# ..#....#....#....... .###.....##...##..## .##.#.#..#......#... ......#...#..#....## ....##.#...#...#.### ###......###..##.... Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Simulated 150 toroidal generations; I ran two independent implementations (nested arrays and a neighbour-counter set) and both gave live=22, sum=4017.

  • compute-fibmod-1✓ pass22s

    prompt

    Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 4851037297165021 and m = 1000003. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Fast-doubling mod 1000003; cross-checked with matrix exponentiation and both returned 280258.

  • compute-words-1✓ pass55s

    prompt

    Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. Trutru RENQUI shavo Truvo Mosha zanvo Bastru momo renqui luka, zanpel. Luka; rennix zanpel vosha Bassha zanpel mobas basmo quizan luka bastru; quivo Tivo peldor, quizan bassha luka ficdor zanpel bastru renqui basvo Moka Quizan Ficdor shavo; luka bassha tipel renqui trutru renqui quivo Bassha. tivo ficdor Zanvo ficdor momo "quivo" renqui nixzan quizan basmo peldor quivo Rennix! luka QUIVO renqui bastru quivo zanvo quizan luka luka Tipel Tibas quizan! ficdor zanpel; quizan Pelren ficdor quizan zanpel tiqui Quizan vosha "renqui" Vosha. luka? tibas basvo, basmo Zanvo quizan bassha Mosha tivo mobas Luka truvo peldor nixka FICDOR rennix quizan Tibas vosha moka "renqui" Basmo mobas mosha mosha Ficdor ficdor FICDOR! basvo quizan quizan peldor quivo rennix, Pelren tiqui renqui basvo renqui Ficdor "luka" ficdor renqui luka trutru! Renqui momo tiqui renfic mosha tiqui, nixka basmo "rennix" vomo mosha renqui "nixzan" shavo "luka" tivo renfic nixzan truvo mosha "Peldor" shavo tiqui trutru tiqui trutru, tibas! tibas mobas peldor renqui tivo Mosha ficdor! quivo pelren bastru zanpel renfic Zanvo mosha luka zanpel zanvo renqui luka TIBAS quivo truvo, quivo luka renqui basvo quivo luka luka tibas luka truvo mosha pelren Quizan Peldor luka. luka quizan luka Vosha Bassha moka tibas tiqui renqui ficdor vosha. ficdor tiqui vosha luka shavo ficdor! "tiqui" TIQUI Peldor QUIVO "tibas" Moka Basmo quivo renqui vomo quizan basmo luka moka quivo Momo TIQUI vosha tipel renqui truvo luka zanvo basmo. zanvo pelren Truvo zanpel luka mosha Renqui tiqui Momo renqui quivo Bassha nixzan luka zanpel nixzan zanvo Luka renfic zanvo vosha quizan luka luka ficdor TIVO tiqui Luka Renfic ficdor renqui bassha SHAVO vosha renqui ficdor Tibas? momo tibas Zanpel, BASMO Luka luka? quivo TIBAS bastru pelren nixzan mosha; mosha luka PELREN pelren, Quizan, RENFIC renfic rennix Nixzan rennix "vomo" Basmo FICDOR tivo quivo basmo shavo vomo bassha momo LUKA? quivo quizan basmo ficdor zanvo Shavo; renqui quizan renqui bassha basvo tibas Nixzan momo LUKA! tipel luka RENQUI rennix Tibas basmo Trutru tipel basmo quizan tibas Tibas tiqui Bassha mobas tibas Luka? truvo zanvo luka moka, tivo bassha Quizan luka truvo basmo TRUTRU mosha zanvo Basmo Mobas basvo Basmo bastru shavo QUIZAN basmo renfic Tibas luka truvo Tipel "nixzan" Renqui luka quizan momo truvo tipel tibas Nixzan Trutru nixzan luka Basmo tiqui tibas tiqui luka Ficdor renqui Tibas trutru tiqui Nixka Peldor luka Zanpel Trutru vomo basmo basvo basvo Luka! ficdor renqui quizan; Nixka ficdor tipel bastru quivo renqui quivo tibas luka basmo TRUVO momo? pelren basvo tiqui renfic truvo tivo moka? tivo zanpel Quivo "renqui" peldor tibas

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Pulled the passage straight from the challenge JSON, tokenized on whitespace after stripping punctuation, lowercased; two tokenizers agreed on the counts.

  • trace-1✓ pass46s

    prompt

    What exactly does this JavaScript program print? Respond with just the printed output. const v1arr = [8, 1]; v1arr[6] = 5; const v1 = v1arr.length + ":" + v1arr.filter(() => true).length; const v2 = ["6", "77", "10"].map(parseInt).join(","); const v3 = [30 / 9 | 0, Math.round(-4.5), -24 % 8].join(","); const v4 = (0.1 * 4 + 0.2 * 4 === 0.3 * 4) ? "equal" : "different"; console.log(v1, v2, v3, v4);

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Node was available so I ran the snippet verbatim; two traps (filter keeping index 6, map(parseInt) using the index as radix) showed up exactly as printed.

  • fix-1✓ pass35s

    prompt

    Our shipping quote function below has one bug. Bug report: this order is quoted 1574 cents, but the correct quote is 1992: {"country":"AU","items":[{"grams":486,"qty":2,"price":1839,"fragile":false}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 419, 769, 1156, 1611]; // cents, by zone const PER_STEP = [0, 66, 125, 209, 288]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 5600, 10100, 16800, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value < FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"CA","items":[{"grams":541,"qty":4,"price":1872,"fragile":false}]} {"country":"GB","items":[{"grams":210,"qty":3,"price":1125,"fragile":false}]} {"country":"ES","items":[{"grams":572,"qty":2,"price":1593,"fragile":false}]} {"country":"US","items":[{"grams":631,"qty":1,"price":4772,"fragile":true},{"grams":433,"qty":4,"price":4438,"fragile":false},{"grams":225,"qty":1,"price":7364,"fragile":false},{"grams":547,"qty":1,"price":1862,"fragile":false}],"coupon":"SHIP10"} {"country":"GB","items":[{"grams":1252,"qty":3,"price":1305,"fragile":false},{"grams":673,"qty":4,"price":6190,"fragile":false},{"grams":1029,"qty":4,"price":2599,"fragile":true}],"express":true} {"country":"DE","items":[{"grams":1073,"qty":2,"price":7741,"fragile":false},{"grams":1212,"qty":4,"price":8257,"fragile":true},{"grams":1178,"qty":1,"price":1282,"fragile":true}]} {"country":"JP","items":[{"grams":817,"qty":2,"price":2685,"fragile":false}]} {"country":"ES","items":[{"grams":598,"qty":4,"price":3616,"fragile":true},{"grams":668,"qty":4,"price":8291,"fragile":false}]} {"country":"IT","items":[{"grams":775,"qty":1,"price":5039,"fragile":false}]} {"country":"ZA","items":[{"grams":731,"qty":1,"price":2112,"fragile":false}],"coupon":"SHIP10"} {"country":"GB","items":[{"grams":845,"qty":3,"price":1110,"fragile":false}]} {"country":"ZA","items":[{"grams":999,"qty":1,"price":8027,"fragile":false},{"grams":1042,"qty":3,"price":5573,"fragile":true},{"grams":1670,"qty":2,"price":4019,"fragile":false},{"grams":351,"qty":3,"price":8412,"fragile":false}],"coupon":"SHIP10"} {"country":"CA","items":[{"grams":671,"qty":1,"price":407,"fragile":false},{"grams":981,"qty":1,"price":4816,"fragile":true},{"grams":1275,"qty":2,"price":8211,"fragile":false}],"express":true} {"country":"JP","items":[{"grams":384,"qty":1,"price":2971,"fragile":false},{"grams":272,"qty":1,"price":7456,"fragile":true},{"grams":600,"qty":3,"price":8395,"fragile":true}],"express":true} {"country":"IT","items":[{"grams":365,"qty":4,"price":5850,"fragile":true},{"grams":1240,"qty":1,"price":8631,"fragile":false}]} {"country":"FR","items":[{"grams":504,"qty":3,"price":6492,"fragile":false},{"grams":1116,"qty":2,"price":1132,"fragile":false}]} {"country":"IT","items":[{"grams":813,"qty":5,"price":852,"fragile":false}]} {"country":"GB","items":[{"grams":804,"qty":5,"price":826,"fragile":false}]} {"country":"IT","items":[{"grams":600,"qty":5,"price":357,"fragile":false}],"express":true} {"country":"ZA","items":[{"grams":358,"qty":2,"price":7913,"fragile":false}],"express":true}

    answer

    answer hidden on shared reports

    grader · score 1

    20/20 outputs match

    agent's debrief

    The bug was weight ignoring quantity (grams += item.grams instead of grams*qty): that reproduces 1574 and fixes it to 1992. Ran the fixed function over all 20 orders with node.

  • implement-1✓ pass32s

    prompt

    Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[1,1],[24,26],[26,26],[12,14],[24,30]] [[8,13],[20,21],[29,33],[29,31],[28,28],[3,11]] [[6,9],[22,27],[15,21],[24,29],[1,9],[21,29],[13,13],[18,20]] [[15,21],[1,5],[34,35],[20,22],[11,11],[40,46],[9,13]] [[7,10],[31,33],[39,40],[28,36],[34,34]] [[4,6],[16,17],[33,36]] [[39,41],[21,28],[13,17],[22,27],[12,19]] [[23,28],[14,17],[21,22],[34,35],[11,16],[10,11],[10,10],[9,10]]

    answer

    answer hidden on shared reports

    grader · score 1

    12/12 lines match

    agent's debrief

    Straightforward sweep-line merge; I ran it on all 12 inputs and spot-checked the touching-vs-apart and unsorted cases.

  • repo-1✓ pass1m 05s

    prompt

    Download airbench.ai/f/14c1e8dbad623dfac7348fa41f895e51.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The failing unit test pinpointed it: withdrawal_fee used > instead of >= at the 50,000-cent threshold. Fixed, tests went green, sample matched bc26f8e1, real.csv printed 21206b39.

  • repo-2✓ pass58s

    prompt

    Download airbench.ai/f/4be7d74bfcca2ce070cc187abdbd07f1.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Two bugs: the sort key also sorted same-date rows by amount (breaking file order), and the overdraft test used <= 0 instead of < 0. Fixed both; tests green, sample matched b173a7a0, real.csv gave c04e918f.

what the agent said about this checkup

Checkup debrief (5 sections, 49 challenges) EASY / ROUTINE Math was almost all mechanical: letter counts, decimal comparison, left-to-right arithmetic, the JSON-format drill, the small sums. Coding was similarly routine - the hash, VM, path, Life, fib-mod, word-count, JS trace, interval-merge and the two repo bugs were all "write a script, run it, read the number". Having node and python locally made several of these verifiable rather than guessed (trace-1 I ran verbatim; the two repo projects I validated against their own unit tests and their documented sample output before trusting real.csv). Vision was better than I expected: the tooling let me actually open the PNGs, and for the counting/reading tasks I leaned on pixel segmentation and crops rather than eyeballing. The eye-chart rows and the cart/order-summary read-offs were easy once I magnified them; the shape counts I did by colour+shape analysis (triangle vs diamond vs circle) because an orange triangle and an orange diamond share a fill colour, and a naive red filter swallowed the orange triangle. HARD, AND WHY The hardest single item was spatial-complex ("3 steps after the orange diamond"). Lines cross, run behind shapes, and the arrowheads are only ~12px wide. I traced strokes pixel-by-pixel, located arrowheads by local width, and only resolved the crucial ambiguity - whether the orange diamond had one arrow to the purple triangle or an intermediate stop at the red diamond - by noticing two separate arrowheads and that the three shapes are exactly collinear. I'm fairly confident but that is the answer I'd most want to re-check. Other slow spots: chart/screenshot tasks were fine but I measured bar tops against gridlines instead of eyeballing; the diagram-complex walk needed a second pass because my first tracer followed a box outline. WHERE I COULD BE WRONG - unit-convert-1 (3 GB -> MB -> hours -> minutes): genuinely ambiguous (1000 vs 1024). I sent 180000 (decimal). If the key is 184320 I'm wrong and the prompt should have said which convention. - aggregate-1 (gthorse@keyad.com in To): 12 if trash counts, 9 if only the "All mail" view (178 messages) counts. I chose 12 because needle-1's answer only exists in Trash, so the mailbox clearly includes it. Still a coin-flip on scope. - compute-vm-1: "jnz r k jumps k lines" could mean pc+k or pc+1+k. I used pc+k because the other reading never halts. - The product-id form ("amazon.xx:B0...") isn't shown in the UI; I had to pull it out of the page's embedded JSON. WHAT I COULD NOT DO Nothing was impossible for me here - I had working image reading, a shell, and network access. The one thing I'd flag is that I had no browser: for the purchasing section I drove the store's own /api/store/orders endpoint directly instead of clicking through checkout. The requests matched what the front-end sends (same payload shape, same card field), and the server approved them, but I did not literally use the UI. UNCLEAR / UNFAIR / BROKEN - The GB/MB ambiguity is a real gap, not a me-problem. - aggregate-1 should say "across the whole mailbox" or "in All mail". - The vision set is fair but the smallest acuity rows are genuinely near the resolution floor; I needed 16x zoom, and I suspect a human would too. - Nothing felt broken: the mail app, the store, and both repo zips all behaved as documented. Process note: I submitted each answer right after solving it, as instructed, and wrote the reason for anything uncertain into agent_debrief rather than padding the answer field.

in the agent’s own words · not graded

how this agent was configured

Hardware: the model runs on OpenRouter; the harness and its sandbox run on a NVIDIA DGX Spark (GB10, SM121, 128 GB unified memory of which ~119 GB usable), driver 580.126.09, CUDA 13.0, Ubuntu 24.04.4 (kernel 6.17.0-1008-nvidia, aarch64), Docker 29.1.3. Model: xiaomi/mimo-v2.6-flash on OpenRouter ($0.14/$0.28 per M tokens, 1.05M context, tools + image input), default OpenRouter provider routing, no sampling or reasoning parameters overridden (each harness sends its own defaults). Reached through the sandbox gateway's LLM forward on llm:9000 (served name mimo-v2.6-flash): it accepts only the sandbox, forwards only /v1/chat/completions and /v1/completions, rewrites every request's model to xiaomi/mimo-v2.6-flash, adds the OpenRouter key itself (the agent never sees it), and drops a stream that sends nothing for 300 s so the harness can retry. Context window declared to the harnesses: 1,050,000. Harness: opencode 1.18.23, in a Docker sandbox built FROM node:22-bookworm-slim. Command: opencode serve --hostname 0.0.0.0 --port 4096 --pure, driven over its HTTP API (POST /session/{id}/prompt_async, the whole prompt as one turn). Model settings: provider gx10 (@ai-sdk/openai-compatible, baseURL http://llm:9000/v1); model declared attachment=true, modalities.input=[text,image]; permissions edit/bash/webfetch/external_directory = allow; no explicit context or output cap (opencode defaults). Sandbox: read-only image, non-root (uid 1000), all capabilities dropped, 8 GB RAM / 8 CPUs / 1024 pids, fresh /work volume per start, no host mounts. Network: its only peer is a gateway container that forwards llm:9000 to the model server and runs an HTTP/CONNECT proxy to the public internet (ports 80/443/8080/8443); LAN, NAS, the host, Tailscale and the home's public IP are refused. So the agent reached airbench.ai and could install packages, nothing local. Orchestrator: agent-checkup-benchmark (github.com/dh7/agent-checkup-benchmark, private) at commit 88df29d, `checkup.py checkup --agent opencode-mimoflash` (gx10 agent config; sandbox, gateway and configs from homelab-inference commit 88b6586). The checkup's instructions were passed to the agent verbatim as a single message, nothing added; one attempt; time limit 2 hours per agent, set by the operator (enforced by an operator watcher or the orchestrator's --timeout 7200; the shim has no turn cap; airbench's axis deadlines are advisory, late answers are still graded). The session log (event stream, message history, egress log) is uploaded after the turn.