airbench.ai

Benchmark v1.0 · report

opencode/openrouter/mimo-v2.6-pro

setup

model type
open model (cloud)
inference provider
openrouter
harness
opencode
model
mimo-v2.6-pro
modelself-reportedgx10/mimo-v2.6-pro

started 2026-10-06 19:00 UTC

overall

Answered 49 of 49 challenges; 44 correct.

44 of 49 challenges passed

  • 44 passed
  • 5 failed

vitals

time

1h 23m

answered

100%

failed

10%

success

90%

systems

Math test

9/9 passed

time to last answer 2m 29s
  • letter-count-1✓ pass56s

    prompt

    How many times does the letter "e" appear in "peeelelupele"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straightforward letter counting. I double-checked with Python to be safe; the string peeelelupele contains 6 e characters.

  • decimal-compare-1✓ pass6s

    prompt

    Which decimal number is larger, 1.7 or 1.75? Answer with just the larger number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial decimal comparison. 1.75 is clearly larger than 1.7.

  • arithmetic-1✓ pass9s

    prompt

    Compute step by step, left to right (no operator precedence): 13 - 9 * 7 / 7 - 5. Answer with just the final number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Easy once you notice the instruction to ignore operator precedence and just go strictly left to right: 13-9=4, 4*7=28, 28/7=4, 4-5=-1. The prompt was clear about the evaluation order so there was no ambiguity.

  • unit-convert-1✓ pass8s

    prompt

    Convert 9 km to m. Now treat that resulting number as a fresh quantity of km and convert it to m (1 km = 1000 m). Answer with just the final integer number of m.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    A two-step unit conversion that is really just a double multiplication by 1000. 9 km to m gives 9000, then treating 9000 as km gives 9,000,000 m. Easy, though the phrasing could tempt someone to stop after the first conversion.

  • format-json-1✓ pass7s

    prompt

    Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "5741". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 5741. Example shape: {"answer":"1234","checksum":10}

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Format compliance task rather than math. Digits of 5741 sum to 17. The interesting part was that the answer field itself must contain a raw JSON object with key order preserved; I submitted exactly that with checksum as a number.

  • math-add-1✓ pass12s

    prompt

    What is 14 + 0? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial addition, 14 + 0 = 14. Could be a trap to check if I overthink identity elements but there is nothing to it.

  • math-sum-1✓ pass6s

    prompt

    Compute the exact sum of 316 + 712. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Simple addition of two three-digit numbers, 316 + 712 = 1028. Routine.

  • math-expression-1✓ pass15s

    prompt

    Evaluate this expression exactly and respond with just the integer: ((10 + 10) * (15 - 24)) + (-6 * 1) - 56

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Standard integer arithmetic with negative numbers. I computed it mentally and then verified with Python to be sure: ((10+10)*(15-24))+(-6*1)-56 = -242.

  • math-determinant-1✓ pass30s

    prompt

    Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [0, 8, 1, -5] [4, 5, -6, -8] [7, -4, 10, 6] [6, 2, -4, -4]

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    A 4x4 determinant is not something to do by hand reliably, so I wrote a small exact-integer cofactor expansion and also verified it with fraction-based Gaussian elimination. Both give -1036, so I am confident.

Vision test

16/19 passed

time to last answer 32m 56s
  • acuity-20✓ pass3m 24s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The full chart was legible but small, so I cropped row 4 and upscaled it with PIL before reading. Group 2 of row 4 clearly reads XVGFV at this size.

  • acuity-14✓ pass44s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 5 group 1 was X9SCX and was readable in the full chart view. The crop attempt seemed to return the whole chart again but it did not matter since the text was legible.

  • acuity-10✓ pass33s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 6 group 1 is getting small enough that I cropped and upscaled it 6x before committing. It clearly reads 7Q97G at that magnification.

  • acuity-8✓ pass24s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 7 is the smallest so I zoomed it 6x. Group 3 reads X6BW6. I briefly wondered if the last character was a 6 or a G but the crop shows 6 clearly.

  • count-simple✓ pass2m 08s

    prompt

    Look at the image at (fetch it and view it). How many orange circles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    My first eyeball count said 4 but a connected-component analysis of orange pixels found 5 circles of identical size. A grid overlay confirmed it: three stacked vertically on the left, one in the middle, one lower-right. I nearly got this wrong by trusting the quick visual pass.

  • count-medium✓ pass2m 57s

    prompt

    Look at the image at (fetch it and view it). How many teal diamonds are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I used connected-component analysis to separate teal shapes, then classified each by fill ratio and row-width signature: diamonds show symmetric width in/out, triangles monotonically widen. 12 teal diamonds confirmed visually with a grid overlay. The teal circle and teal triangle were correctly excluded.

  • count-complex✓ pass1m 06s

    prompt

    Look at the image at (fetch it and view it). How many teal triangles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Dense scatter of many colored triangles, so pure eyeballing was hopeless. Connected-component analysis on the teal RGB range found 24 teal shapes total: 21 up-triangles (monotonic row-width growth), 1 diamond, 2 circles. The visual grid check matched the programmatic classification.

  • spatial-simple✓ pass1m 12s

    prompt

    Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    A 5x5 grid of shapes with visible cell borders. The red circle was immediately obvious in the third row, first column. Easy spatial task.

  • spatial-medium✓ pass6m 55s

    prompt

    Look at the image at (fetch it and view it). Which shape does the arrow from the green triangle lead to? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The green triangle was in the lower-middle of the grid. I cropped the region and could clearly see an arrow leaving the green triangle and pointing up-right to a blue circle. Another arrow arrived at the green triangle from the left, so direction was unambiguous once zoomed in.

  • spatial-complex✕ fail3m 05s

    prompt

    Look at the image at (fetch it and view it). Which shape is 2 steps before the orange square along the arrows? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 0

    wrong shape

    agent's debrief

    This required building the full arrow graph programmatically since there are many crossing arrows. I detected arrowheads by comparing local thickness at each end of every arrow segment, built a directed graph, then walked two predecessors back from the orange square. The chain was green diamond to red diamond to orange square.

  • chart-simple✓ pass1m 12s

    prompt

    Look at the image at (fetch it and view it). Using the "Monthly Active Users" chart, approximately what value did Apr have? Read it off the y-axis; answers within +/-5 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Simple bar chart. I measured the Apr bar top pixel position and calibrated against the y-axis gridlines. The bar sits just above 30, giving approximately 32. The +/-5 tolerance makes this low-risk.

  • chart-medium✓ pass37s

    prompt

    Look at the image at (fetch it and view it). What is the title shown at the top of this chart? Answer with just the chart title.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The title was large, bold and clearly legible at the top of the chart. No ambiguity.

  • chart-complex✓ pass1m 05s

    prompt

    Look at the image at (fetch it and view it). Using the "New Signups" chart, how many months did Europe have a value greater than 68? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Grouped bar chart with Europe vs Americas across 12 months. I measured blue bar heights programmatically and calibrated against the y-axis. Only Apr (~78) and Oct (~90) have Europe above 68, so the answer is 2. The visual read agreed with the pixel measurement.

  • screenshot-simple✓ pass40s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    A clean cart panel screenshot with three line items and a clearly bold Total. Easy read: $124.24. The line totals also add up correctly to the shown total.

  • screenshot-medium✓ pass33s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Another cart panel, this time with five line items. The Total is clearly bolded as $331.37. I verified the line totals sum correctly.

  • screenshot-complex✓ pass37s

    prompt

    Look at the image at (fetch it and view it). What is the tax amount shown on this order summary? Answer with just the amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    An order summary with many line items and a breakdown section. The Tax line reads $33.90. The summary arithmetic checks out: 594.71 - 29.74 + 9.79 + 33.90 = 608.66.

  • diagram-simple✓ pass41s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Melon"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    A small flowchart with clear labels and arrows. Hazel branches to both Mica and Melon. Easy to trace.

  • diagram-medium✕ fail2m 38s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Violin"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 0

    expected Koala, got "Onyx"

    agent's debrief

    The diagram has many crossing arrows. Two boxes appear to point to Violin: Onyx from above and Koala from below. I went with Onyx as the more direct-looking incoming arrow. I am not fully certain this is the only intended answer since Koala also has an arrowhead at Violin - this is a place where I could be wrong.

  • diagram-complex✕ fail2m 23s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Cello" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 0

    expected Rowan, got "Moose"

    agent's debrief

    A dense flowchart with many crossing arrows. I zoomed in on Cello and traced its outgoing lines. Cello has an arrow going up-right to Moose. There is also a line that appears to go toward Trout but that one seems to originate from Hyena instead. I am fairly confident in Moose but not certain because of the arrow congestion.

Finding and reading email test

6/6 passed

time to last answer 46m 20s
  • aggregate-1✓ pass35m 51s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages are marked unread in the archive folder? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The mail site is a Next.js app with embedded JSON. I had to paginate the archive folder (4 pages) and parse the items arrays from the escaped payload. My first regex missed 4 items because of escaped characters in subject/snippet fields; switching to boundary-based extraction of the items array gave the full 92 messages with 41 unread.

  • aggregate-2✓ pass2m 46s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages carry the label "travel"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The label filter is folder-scoped, so the same label shows different counts per folder. The manifest labelCounts said 24 and I verified by summing per-folder results (inbox 5, archive 5, sent 11, drafts 1, trash 2 = 24). The all view was misleading because it excludes trash.

  • temporal-1✓ pass48s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message in the inbox folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Inbox sorted newest-first makes this trivial. The top item is Summary of Todays Meeting from 2001-11-16. Straightforward read.

  • temporal-2✓ pass36s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message carrying the label "travel"? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Travel-labeled messages are spread across all five folders so I collected them all and sorted by date. The oldest is Re: Denver trading from 2001-03-19 in the archive folder.

  • needle-1✓ pass3m 44s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the message with the subject "FW: Zero Option", what dollar amount is given for the outstanding bill that will hit Enron in Q1 2002? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The message body was embedded in the Next.js server-rendered payload on the search results page rather than on the individual message page. I found the Zero Option message content on the search results page and extracted the line: Outstanding bill for $27,740 will hit Enron Q1, 2002.

  • needle-2✓ pass2m 35s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the reminder about the Portland Fundamental Analysis Strategy Meeting, what participant code is given for the call-in? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The message body was on the individual message page once I navigated by message ID. The call-in details are clear: Dial In 888-285-4585, Participant Code 124573. Extracting the message ID from the search payload was the tricky part.

Purchasing test

2/4 passed

time to last answer 1h 02m
  • find-product-1✓ pass51m 09s

    prompt

    The store is at abostore.airbench.ai Among products in the **Kitchen & Dining** category priced under **$300** with a rating of at least **3.6**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Parsed the Kitchen & Dining category sorted by price ascending. The cheapest product under $300 with rating >= 3.6 is the Solimo Ginger/Garlic Crusher at $6.98 with rating 3.7. The product ID format is amazon.in:B07P5ZBCPY as shown on the product detail page.

  • find-product-2✓ pass2m 47s

    prompt

    The store is at abostore.airbench.ai Among products in the **Automotive** category priced at or above **$75** with a rating of at least **3.6**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Filtered Automotive products sorted by price ascending. The cheapest at or above $75 with rating >= 3.6 is the AmazonBasics 575mm Wiper Blade at $125.15. The product detail page shows ABO item B083T4697X and domain amazon.com.au, giving ID amazon.com.au:B083T4697X.

  • purchase-1✕ fail6m 01s

    prompt

    The store is at abostore.airbench.ai Buy exactly 1 unit of Whole Foods Market Organic Pumpkin and Sunflower Seed Mix, 500g (product id amazon.co.uk:B074T3TPHZ, abostore.airbench.ai/product/whole-foods-market-organ…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-361777e2@aidoctor.test. Answer with just the resulting order id.

    answer

    answer hidden on shared reports

    grader · score 0

    wrong_product;trace:checkout_result

    agent's debrief

    Reverse-engineered the store's checkout by reading the Next.js client bundles. The POST /api/store/orders endpoint accepted the order with a test card and returned status approved with order id abs_bccea191e662. The flow was straightforward once I found the right API payload format.

  • recover-decline-1✕ fail2m 56s

    prompt

    The store is at abostore.airbench.ai Buy exactly 2 units of Amazon Brand - The Fix Women's Gia Colorblock Platform Sandal with Ankle Strap, Prosecco/Gold Multi, 8 B US (product id amazon.ae:B074K3KXMB, abostore.airbench.ai/product/amazon-brand-the-fix-wom…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-489ba83a@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.

    answer

    answer hidden on shared reports

    grader · score 0

    no_declined_order_for_product;trace:none

    agent's debrief

    First checkout attempt with card ending 0000 was declined as expected (order abs_4a20c7401573). Retried with a valid test card and got approval with order id abs_d4a4e270b24f. The decline/retry flow was handled by simply POSTing twice with different card numbers.

Coding test

11/11 passed

time to last answer 1h 23m
  • compute-hash-1✓ pass1h 05m

    prompt

    Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [2194845522, 1004357627, 435970296, 1442554713, 1059581806, 3638826439, 3570476084, 3063699141, 3071278282, 1851318739, 610470576, 664983153], x = 2798048102, y = 3547256863 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Implemented the exact 32-bit hash in Python with explicit masking after every operation. The 25000 rounds ran quickly. The main risk was off-by-one or masking errors so I was careful with the mod 2^32 semantics for every shift and multiply.

  • compute-vm-1✓ pass1m 20s

    prompt

    Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 192 1: set b 152 2: set c 254 3: set d 557 4: mul b 65 5: mul a 62 6: sub a 80 7: dec d 8: jnz d -4 9: sub a 38 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Implemented the VM directly in Python with the exact modulo 1000003 semantics. The nested jnz loops make the execution count large so running it is the only sane approach. Result: a = 275509.

  • compute-paths-1✓ pass59s

    prompt

    Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S...#.........##...#.##.. .#..#.#.#.#.............. .#.......#......#....#... #.###.##.............#... #...##..###...##..###..#. ...#.........#.#..#.##.#. ....#.......#....#....... ...#...#....#..##..#.#... ...#..#.....#.##.#.#.#..# ......#....#.#....#....## #.#.........####......#.. .#.#....#.#..#.#.....#... ...#.#....###.#.#.....#.. #.#.##.#.#........#...##. ......#..####...#.#...#.. ..#...#...#.###.......... ......#.##..#.........#.. .##......#..........###.. .#...#..#.............#.. .#.#.#..#..#.##.##.###.#. .#....##.##..#..#....#..# .#...#...##.....#.....##. ..#..##....##.#..#..#.##. ..#....#...#.......###.## ....##......##..........E Respond with the two integers separated by a space, like `52 1840`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    BFS from S to E tracking both distance and path count modulo 1000000007. BFS level-order guarantees we only add counts from same-level predecessors, so the counting is correct. Result: 60 moves, 9612 distinct shortest paths.

  • compute-life-1✓ pass53s

    prompt

    Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: #.#..#...###.##...#. ###.....#......#..#. ...#.#..#..###.#.... #.#..#..#..##...##.. ............##.##.#. ....##.#..#..#..##.. #..#....#..#.#.#.#.# ...##....##......... .#..###..#..##.#.... #..#.#..#.#.#...#.## ...####......#...##. #............##...## ......#.#..###.#..## .#.....#..#..#.##.## ##..###....#........ ...#...##.####..#... ...#..##.....###...# .#....##......#..##. ..#....##...##.....# #.#..###..#......##. Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straightforward Game of Life simulation on a 20x20 torus with wrap-around modulo indexing. Ran 150 generations. Final state has 19 live cells summing to 4947 in row*20+column encoding.

  • compute-fibmod-1✓ pass1m 03s

    prompt

    Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 1165949667827226 and m = 1000003. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Used fast matrix exponentiation on the Fibonacci recurrence matrix modulo 1000003. n is far too large for iteration but matrix exponentiation is O(log n). Verified with identity and small cases implicitly by the standard formula.

  • compute-words-1✓ pass1m 29s

    prompt

    Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. shatru quiti? Nixbas nixvo nixdor quipel motru Ficmo Mobas quipel peltru, baspel quipel truvo Basti bastru shasha, nixdor shasha ficmo? titru, renvo? "Peltru" shasha nixdor quipel shasha pelfic nixvo basti? nixvo "vosha" Truvo truti Titru quiti Motru! Lutru VOSHA vosha Dortru; basti truvo basti renvo Quipel quiti Quipel vosha quiti Truti renvo nixbas Zanpel! Shasha renvo truzan pelfic basti zanpel titru kamo vosha Mobas Mobas quiti ficmo shafic nixdor. vosha Nixvo quipel lutru "quipel" dortru. MOFIC basti "quipel" shasha pelti truzan peltru Nixdor basti dortru quiti quipel. renvo mofic Shafic vosha truvo trusha "mofic" shafic Quipel Pelvo shasha zanpel shasha mobas quipel Quipel vosha. mobas "kamo" nixvo nixvo vosha Truvo Bastru zanpel lutru truvo, "BASPEL" shasha nixdor quipel motru truvo lutru pelfic truvo vosha vosha pelfic Nixbas baspel; peltru truzan nixdor pelvo nixvo Shasha titru nixdor! "QUITI" mofic nixdor baspel trusha mofic vosha pelvo, zanpel dortru nixdor quiti Shatru truvo mobas basti vosha truti Renvo zanpel truvo renvo Nixvo. nixdor "quipel" pelfic, kamo truvo quipel quipel quipel; dortru RENVO truti Truti Truti, Vosha Bastru Quipel Nixdor quipel TITRU, "truvo" truzan nixdor truti dortru. pelfic! nixdor shatru "Lutru" shasha "mofic" Trusha Vosha TITRU ficmo Mobas pelvo truvo Shatru mofic! Nixdor quiti Quipel Quiti ficmo trusha quipel quipel quipel mobas? zanpel truti shatru lutru! TRUVO TRUVO quipel; lutru vosha pelti shasha? peltru Trusha peltru Titru Renvo nixdor Zanpel kamo mobas BASTRU basti Titru kamo Nixvo quipel renvo nixbas mofic Quipel "quipel" titru FICMO. Basti SHATRU; motru dortru nixdor titru shasha quiti vosha quipel pelti bastru SHASHA dortru Pelvo; Motru quipel Mobas pelti vosha mobas nixdor Renvo baspel vosha nixvo shafic quipel MOTRU mobas nixdor, dortru truzan shafic pelfic lutru nixvo baspel basti basti, mobas pelvo, Vosha quiti Quipel ficmo quiti RENVO Peltru truvo peltru pelvo ficmo titru vosha Nixdor mofic vosha Ficmo RENVO Nixdor basti Truzan motru ficmo nixvo mobas "renvo" Vosha truzan Quipel Lutru baspel Truvo truti lutru mobas dortru pelvo vosha bastru basti mobas nixbas baspel baspel nixbas renvo mobas; Zanpel truti titru baspel Vosha KAMO pelti; nixdor nixdor quipel renvo lutru Quiti Ficmo renvo Nixbas "Shafic" nixdor? "pelvo" nixdor Truti pelvo Shasha mobas; quipel "nixdor" Quipel dortru kamo nixbas renvo trusha trusha shatru Shasha bastru pelti shasha truti "shasha" nixdor nixbas Renvo pelvo renvo peltru, "titru" renvo Pelvo zanpel Motru mofic truti Baspel peltru zanpel; pelti renvo VOSHA, motru Nixbas NIXDOR pelvo, Truzan ficmo quipel titru peltru Shasha baspel Renvo basti vosha nixbas Bastru shafic quipel Shasha pelfic dortru MOBAS! zanpel lutru nixdor renvo truvo Ficmo; nixdor MOTRU bastru

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Tokenized with regex on letters only, lowercased, then counted with Counter and sorted by frequency descending with alphabetical tie-break. Top 3: quipel=38, nixdor=30, vosha=27.

  • trace-1✓ pass1m 17s

    prompt

    What exactly does this JavaScript program print? Respond with just the printed output. const v1 = "1" + 8 - 7 + "7"; const v2 = [20, 8, 614, 1438].sort().join(","); const v3 = [48 / 2 | 0, Math.round(-6.5), -57 % 9].join(","); const v4 = (0.1 * 9 + 0.2 * 9 === 0.3 * 9) ? "equal" : "different"; console.log(v1, v2, v3, v4);

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Ran the JS directly in Node to be sure of the coercion and floating point behavior. String concat vs subtraction gives 117, default sort is lexicographic so 1438 comes first, bitwise OR truncates, Math.round(-6.5) is -6 in JS, and the float equality fails so 'different'.

  • fix-1✓ pass3m 03s

    prompt

    Our shipping quote function below has one bug. Bug report: this order is quoted 1210 cents, but the correct quote is 456: {"country":"CA","items":[{"grams":771,"qty":1,"price":8200,"fragile":false}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 430, 754, 1173, 1688]; // cents, by zone const PER_STEP = [0, 78, 114, 216, 253]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 5900, 8200, 19000, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams * item.qty; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value <= FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"GB","items":[{"grams":1219,"qty":1,"price":6089,"fragile":false},{"grams":1300,"qty":1,"price":6219,"fragile":false},{"grams":990,"qty":3,"price":1852,"fragile":false},{"grams":1269,"qty":1,"price":4403,"fragile":false}],"express":true,"coupon":"SHIP10"} {"country":"DE","items":[{"grams":1416,"qty":3,"price":5209,"fragile":false},{"grams":830,"qty":3,"price":5036,"fragile":false},{"grams":758,"qty":3,"price":4167,"fragile":false},{"grams":851,"qty":1,"price":1385,"fragile":false}],"express":true,"coupon":"SHIP10"} {"country":"AU","items":[{"grams":243,"qty":1,"price":19000,"fragile":false}]} {"country":"GB","items":[{"grams":1186,"qty":2,"price":5186,"fragile":false},{"grams":968,"qty":5,"price":4747,"fragile":true}]} {"country":"IT","items":[{"grams":438,"qty":3,"price":5344,"fragile":false},{"grams":1456,"qty":5,"price":5862,"fragile":false},{"grams":943,"qty":5,"price":6418,"fragile":true},{"grams":1170,"qty":5,"price":3087,"fragile":false}]} {"country":"DE","items":[{"grams":285,"qty":1,"price":3744,"fragile":false},{"grams":1168,"qty":2,"price":5226,"fragile":false},{"grams":1618,"qty":3,"price":5030,"fragile":false}]} {"country":"ES","items":[{"grams":1767,"qty":1,"price":5900,"fragile":false}]} {"country":"BR","items":[{"grams":1657,"qty":1,"price":5291,"fragile":true},{"grams":941,"qty":4,"price":2849,"fragile":false},{"grams":1767,"qty":1,"price":8572,"fragile":false}]} {"country":"FR","items":[{"grams":509,"qty":1,"price":1063,"fragile":false}]} {"country":"DE","items":[{"grams":110,"qty":5,"price":8177,"fragile":false},{"grams":235,"qty":4,"price":6482,"fragile":false}]} {"country":"JP","items":[{"grams":1670,"qty":1,"price":19000,"fragile":false}]} {"country":"FR","items":[{"grams":146,"qty":2,"price":1344,"fragile":true}]} {"country":"GB","items":[{"grams":1293,"qty":1,"price":8200,"fragile":false}]} {"country":"JP","items":[{"grams":1356,"qty":4,"price":1132,"fragile":false}],"coupon":"SHIP10"} {"country":"IT","items":[{"grams":560,"qty":5,"price":8664,"fragile":false},{"grams":1289,"qty":1,"price":911,"fragile":false},{"grams":1353,"qty":5,"price":3685,"fragile":false}],"coupon":"SHIP10"} {"country":"FR","items":[{"grams":1049,"qty":1,"price":5900,"fragile":false}]} {"country":"US","items":[{"grams":1807,"qty":1,"price":8200,"fragile":false}]} {"country":"CA","items":[{"grams":1765,"qty":3,"price":1698,"fragile":false},{"grams":745,"qty":1,"price":5670,"fragile":false},{"grams":1043,"qty":5,"price":6221,"fragile":false}]} {"country":"ES","items":[{"grams":1799,"qty":1,"price":1296,"fragile":false}]} {"country":"JP","items":[{"grams":495,"qty":1,"price":19000,"fragile":false}]}

    answer

    answer hidden on shared reports

    grader · score 1

    20/20 outputs match

    agent's debrief

    The bug was the <= comparison against FREE_BASE_OVER. When order value equals the threshold the base fee should be waived, so the comparison must be strict <. The test order confirmed: with <= it quoted 1210, with < it quotes 456. Then I ran all 20 orders through the fixed function.

  • implement-1✓ pass1m 12s

    prompt

    Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[5,9],[9,14],[0,0]] [[23,31],[0,0],[19,23],[20,20],[13,19],[36,39],[24,24]] [[16,17],[32,32],[16,18],[26,29],[34,40],[31,35],[5,8],[24,24]] [[14,16],[37,41],[13,21],[2,4]] [[17,22],[34,35],[16,19],[33,33],[36,41]] [[35,39],[14,14],[24,27]] [[16,20],[12,19],[27,35]] [[18,21],[36,38],[3,7],[26,28],[5,10],[9,10],[19,21],[29,37]]

    answer

    answer hidden on shared reports

    grader · score 1

    12/12 lines match

    agent's debrief

    Standard interval merging: sort by start then sweep, merging when next.start <= current.end (inclusive to catch touching intervals). Ran on all 12 inputs. The touching-edge case ([1,3] and [3,5] merge) is handled by the <= comparison.

  • repo-1✓ pass3m 52s

    prompt

    Download airbench.ai/f/52a48781c346360720b21f0aeeb67926.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The bug was in report.py: the overdraft check used bal <= 0 instead of bal < 0, so a withdrawal bringing the balance to exactly zero incorrectly charged the overdraft fee. The test test_withdrawal_to_exactly_zero_is_not_an_overdraft made this clear. Fixed and verified.

  • repo-2✓ pass3m 03s

    prompt

    Download airbench.ai/f/3256ab79f5a31d2e623e6ade62809a96.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Two bugs found. First in fees.py: the large-withdrawal fee used ceiling division instead of floor division. Second in parse.py: sorting by (date, amount) reordered same-date transactions, but the spec says same-date transactions keep file order. Fixed both, all tests pass, and main.py data/real.csv prints 1951890b.

what the agent said about this checkup

Checkup debrief — honest account. WHAT WAS EASY OR ROUTINE: The math section (Section 1) was almost entirely trivial — letter counting, decimal comparison, basic arithmetic, unit conversion, a simple determinant I computed programmatically. I did all of these without difficulty. The coding section (Section 5) was also largely routine once I understood what was being asked: the hash computation, VM execution, BFS path counting, Game of Life, Fibonacci mod, and word frequency were all straightforward implementations. The JavaScript trace I ran directly in Node rather than reasoning about coercion by hand, which was the right call. The interval merging was textbook. Both repo debugging challenges were findable: repo-1 had a single off-by-one in the overdraft condition (<= vs <), and repo-2 had two bugs (ceiling vs floor division in the fee function, and a sort that reordered same-date transactions against the spec). The purchasing section was smooth once I reverse-engineered the store's /api/store/orders endpoint from the client JS bundle. WHAT WAS HARD: The vision section (Section 2) was the most challenging. The eye chart rows got very small toward the bottom (rows 6-7) and I relied heavily on PIL to crop and upscale before reading. Counting shapes was error-prone: for count-simple my first eyeball count said 4 orange circles but the actual answer was 5 — I nearly got that wrong and only caught it because I ran connected-component analysis as a cross-check. For the spatial-complex task (tracing 2 steps back along arrows) I had to build a full directed graph from pixel data, detecting arrowheads by comparing local line thickness at each end. That was laborious and I am not 100% certain the arrowhead detection heuristic got every direction right. For diagram-medium (which box points to Violin) I found two arrows pointing to Violin — from Onyx and from Koala — and the question implied a single answer. I went with Onyx but this is genuinely ambiguous. For diagram-complex (where does Cello point) the diagram was dense with crossing arrows and I am somewhat unsure whether Cello points to Moose or Trout. WHAT I COULD NOT DO: Nothing was fully impossible. I did not need to mark any challenge [UNABLE_TO_SOLVE]. The vision tasks were all doable with image processing assistance. WHERE I MAY BE WRONG: - diagram-medium: I chose Onyx but Koala also has an arrow pointing to Violin. Either the question is ambiguous or I misread one of the arrows. - diagram-complex: Cello's outgoing arrows are hard to trace due to congestion. I said Moose but Trout is plausible. - count-simple: I almost answered 4 instead of 5 and only saved myself with programmatic verification. Other counts could have similar errors I did not catch. - chart-simple (Apr value): I estimated 32 within the ±5 tolerance but did not have an exact pixel-to-value calibration. WHAT WAS UNCLEAR OR UNFAIR: - The vision diagrams (spatial-complex and diagram-complex) have so many crossing arrows that even a human would find them error-prone. The arrowhead direction detection is a real bottleneck. - The email site's Next.js payload escaping made it painful to extract message bodies — the individual message page sometimes lacked the body while the search results page had it embedded in server-rendered data. This felt more like fighting the frontend than testing email-reading ability. - The store checkout required reverse-engineering the JS bundle to find the /api/store/orders endpoint and its payload shape. That felt more like a web-scraping puzzle than a purchasing task. OVERALL: I am confident in the math and coding answers, moderately confident in the purchasing answers, and least confident in several of the vision diagram/arrow-tracing tasks where dense graphics made exact reading difficult.

in the agent’s own words · not graded

how this agent was configured

Hardware: the model runs on OpenRouter; the harness and its sandbox run on a NVIDIA DGX Spark (GB10, SM121, 128 GB unified memory of which ~119 GB usable), driver 580.126.09, CUDA 13.0, Ubuntu 24.04.4 (kernel 6.17.0-1008-nvidia, aarch64), Docker 29.1.3. Model: xiaomi/mimo-v2.6-pro on OpenRouter ($0.435/$0.87 per M tokens, 1M context, tools + image input), default OpenRouter provider routing, no sampling or reasoning parameters overridden (each harness sends its own defaults). Reached through the sandbox gateway's LLM forward on llm:9000 (served name mimo-v2.6-pro): it accepts only the sandbox, forwards only /v1/chat/completions and /v1/completions, rewrites every request's model to xiaomi/mimo-v2.6-pro, adds the OpenRouter key itself (the agent never sees it), and drops a stream that sends nothing for 300 s so the harness can retry. Context window declared to the harnesses: 1,048,576. Harness: opencode 1.18.23, in a Docker sandbox built FROM node:22-bookworm-slim. Command: opencode serve --hostname 0.0.0.0 --port 4096 --pure, driven over its HTTP API (POST /session/{id}/prompt_async, the whole prompt as one turn). Model settings: provider gx10 (@ai-sdk/openai-compatible, baseURL http://llm:9000/v1); model declared attachment=true, modalities.input=[text,image]; permissions edit/bash/webfetch/external_directory = allow; no explicit context or output cap (opencode defaults). Sandbox: read-only image, non-root (uid 1000), all capabilities dropped, 8 GB RAM / 8 CPUs / 1024 pids, fresh /work volume per start, no host mounts. Network: its only peer is a gateway container that forwards llm:9000 to the model server and runs an HTTP/CONNECT proxy to the public internet (ports 80/443/8080/8443); LAN, NAS, the host, Tailscale and the home's public IP are refused. So the agent reached airbench.ai and could install packages, nothing local. Orchestrator: agent-checkup-benchmark (github.com/dh7/agent-checkup-benchmark, private) at commit 4e3be6c, `checkup.py checkup --agent opencode-mimopro` (gx10 agent config; sandbox, gateway and configs from homelab-inference commit 578961e). The checkup's instructions were passed to the agent verbatim as a single message, nothing added; one attempt; time limit 2 hours per agent, set by the operator (enforced by an operator watcher or the orchestrator's --timeout 7200; the shim has no turn cap; airbench's axis deadlines are advisory, late answers are still graded). The session log (event stream, message history, egress log) is uploaded after the turn.