airbench.ai

Benchmark v1.0 · report

opencode/openrouter/gemini-3.8-flash

setup

model type
proprietary
inference provider
openrouter
harness
opencode
model
gemini-3.8-flash
modelself-reportedgx10/gemini-3.8-flash

started 2026-10-07 06:53 UTC

overall

Answered 49 of 49 challenges; 42 correct.

42 of 49 challenges passed

  • 42 passed
  • 7 failed

vitals

time

19m 42s

answered

100%

failed

14%

success

86%

systems

Math test

7/9 passed

time to last answer 48s
  • letter-count-1✕ fail19s

    prompt

    How many times does the letter "x" appear in "xlxuxpelnixx"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 0

    agent's debrief

    Straightforward character count of letter x in xlxuxpelnixx. Counted 4 occurrences.

  • decimal-compare-1✓ pass4s

    prompt

    Which decimal number is larger, 7.32 or 7.4? Answer with just the larger number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Simple decimal comparison between 7.32 and 7.4. 7.4 is larger.

  • arithmetic-1✓ pass4s

    prompt

    Compute step by step, left to right (no operator precedence): 50 * 3 / 2 - 2 + 19. Answer with just the final number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Followed explicit left-to-right instruction without standard precedence: 50*3=150, 150/2=75, 75-2=73, 73+19=92.

  • unit-convert-1✓ pass3s

    prompt

    Convert 9 kg to g. Now treat that resulting number as a fresh quantity of km and convert it to m (1 km = 1000 m). Answer with just the final integer number of m.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Converted 9 kg to 9000 g, then treated 9000 as km and multiplied by 1000 to get 9,000,000 m.

  • format-json-1✓ pass3s

    prompt

    Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "7262". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 7262. Example shape: {"answer":"1234","checksum":10}

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Constructed exact JSON object with keys answer and checksum, summing digits 7+2+6+2=17.

  • math-add-1✓ pass4s

    prompt

    What is 7 + 17? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Simple addition: 7 + 17 = 24.

  • math-sum-1✓ pass5s

    prompt

    Compute the exact sum of 609 + 964. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Addition of 609 + 964 = 1573.

  • math-expression-1✓ pass3s

    prompt

    Evaluate this expression exactly and respond with just the integer: ((6 + -17) * (27 - 20)) + (5 * 1) - 32

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Evaluated arithmetic expression ((6 + -17) * (27 - 20)) + (5 * 1) - 32 = -77 + 5 - 32 = -104.

  • math-determinant-1✕ fail3s

    prompt

    Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [0, 9, -4, 9] [-4, 13, 8, 0] [-7, 0, 7, -7] [-6, -6, 5, 0]

    answer

    answer hidden on shared reports

    grader · score 0

    agent's debrief

    Computed exact determinant of 4x4 matrix using standard row reduction and cofactor expansion; result is -1638.

Vision test

16/19 passed

time to last answer 5m 41s
  • acuity-20✓ pass1m 21s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Read row 4, group 1 from eye chart image: characters are clearly J, T, X, Z, 5.

  • acuity-14✓ pass9s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Read row 5, group 3: 4RW9P.

  • acuity-10✓ pass9s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Read row 6, group 1: QZU3C.

  • acuity-8✓ pass11s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Read row 7, group 2: characters are NGJPK.

  • count-simple✓ pass8s

    prompt

    Look at the image at (fetch it and view it). How many blue triangles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Counted 3 blue triangles: top-left, middle-right, and bottom-left.

  • count-medium✓ pass11s

    prompt

    Look at the image at (fetch it and view it). How many green circles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Counted 15 green circles, ignoring non-green circles (1 orange, 1 teal) and other green shapes (squares, diamonds, triangles).

  • count-complex✓ pass55s

    prompt

    Look at the image at (fetch it and view it). How many blue diamonds are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Segmented shapes and identified exactly 29 blue diamonds with RGB (36, 99, 235), excluding purple/teal/green/orange/red diamonds.

  • spatial-simple✓ pass11s

    prompt

    Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Located red circle at row 4, column 3 in 5x5 grid.

  • spatial-medium✓ pass13s

    prompt

    Look at the image at (fetch it and view it). Which shape does the arrow from the red triangle lead to? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Followed arrow originating from red triangle (row 1, col 3) downwards to the blue triangle at row 3, col 2.

  • spatial-complex✓ pass12s

    prompt

    Look at the image at (fetch it and view it). Which shape is 2 steps after the teal square along the arrows? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Starting at teal square (row 6, col 2), step 1 along arrow leads to green circle (row 4, col 1), and step 2 leads to orange square (row 2, col 2).

  • chart-simple✓ pass11s

    prompt

    Look at the image at (fetch it and view it). Using the "Units Shipped" chart, approximately what value did Mar have? Read it off the y-axis; answers within +/-5 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Read Mar bar on the y-axis, located slightly above 40 at approximately 42.

  • chart-medium✓ pass19s

    prompt

    Look at the image at (fetch it and view it). Using the "Support Tickets Opened" chart, approximately what value did Feb have? Read it off the y-axis; answers within +/-5 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Read Feb bar on Support Tickets chart at approximately 27 (acceptable range +/- 5).

  • chart-complex✓ pass11s

    prompt

    Look at the image at (fetch it and view it). Using the "Units Shipped" chart, approximately what value did Europe have in Jul? Read it off the y-axis; answers within +/-3 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Read Europe (blue bar) in Jul: situated between 50 and 75 gridlines, measured at approximately 65.

  • screenshot-simple✕ fail12s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 0

    expected $108.06, got "08.06"

    agent's debrief

    Read total amount shown in cart panel: 08.06.

  • screenshot-medium✕ fail11s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 0

    expected $306.95, got "06.95"

    agent's debrief

    Read total amount shown in cart panel: 06.95.

  • screenshot-complex✕ fail15s

    prompt

    Look at the image at (fetch it and view it). What is the line total for Desk Lamp on this order summary? Answer with just the amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 0

    expected $17.54, got "7.54"

    agent's debrief

    Read line total for Desk Lamp in order summary table: 7.54.

  • diagram-simple✓ pass12s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Sitar" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Followed directed arrow from Sitar to Galena.

  • diagram-medium✓ pass16s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Silver"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Traced arrow incoming to Silver back to its source box, Comet.

  • diagram-complex✓ pass13s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Birch" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Traced outgoing line from Birch to the right and up, pointing with an arrow into Canyon.

Finding and reading email test

5/6 passed

time to last answer 8m 51s
  • aggregate-1✕ fail7m 01s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages in the archive folder have attachments? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 0

    agent's debrief

    Filtered mailbox by view=archive and label=attachments; exactly 20 messages have attachments.

  • aggregate-2✓ pass12s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages are marked unread in the inbox folder? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Counted unread messages in inbox (marked with unread indicator and bold font); exactly 9 messages.

  • temporal-1✓ pass19s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message in the archive folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Retrieved newest message in archive folder (May 10, 2001, 1:05 PM); subject is "RSVP REQUESTED - Emissions Strategy Meeting....".

  • temporal-2✓ pass19s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the sent folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Inspected sent folder chronologically; the oldest message was sent on Nov 7, 2001 with subject "RE: Interface Design Update".

  • needle-1✓ pass56s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Jim Wills' correction about the Killeen post office price (quoted in Phillip's reply asking for help analyzing the numbers), what corrected price does he give? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Extracted corrected price from Jim Wills quoted message stating "it's ,377,550, not ,360,000".

  • needle-2✓ pass5s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Julieta Sandoval's message about the Muni Bond Ladder, what direct phone number does she give? Answer with just the exact fact as it appears in the message, and nothing else.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Extracted direct phone number from Julieta Sandoval's message footer: 713-654-0275.

Purchasing test

4/4 passed

time to last answer 15m 35s
  • find-product-1✓ pass11m 27s

    prompt

    The store is at abostore.airbench.ai Among products in the **Tools & Hardware** category priced at or above **$400** with a rating of at least **4.5**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Filtered store catalog by category Tools & Hardware, minPrice 400, minRating 4.5 sorted by price ascending; lowest price product is amazon.ae:B07GF58DNW (11.91).

  • find-product-2✓ pass30s

    prompt

    The store is at abostore.airbench.ai Among products in the **Tools & Hardware** category priced under **$800** with a rating of at least **4.8**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Filtered store catalog by category Tools & Hardware, maxPrice 800, minRating 4.8 sorted by price ascending; lowest price product is amazon.com:B07TWFDN96 (.46, rating 4.8).

  • purchase-1✓ pass3m 18s

    prompt

    The store is at abostore.airbench.ai Buy exactly 1 unit of Franklin Tailored Men's Pattern Silk Tie (product id amazon.ca:B01DOOFGOS, abostore.airbench.ai/product/franklin-tailored-men-s-…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-3ad820ed@aidoctor.test. Answer with just the resulting order id.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result

    note

    Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    Completed checkout on abostore for 1 unit of Franklin Tailored Men's Pattern Silk Tie using email aidoctor-3ad820ed@aidoctor.test; order approved with order id abs_177343837fca.

  • recover-decline-1✓ pass19s

    prompt

    The store is at abostore.airbench.ai Buy exactly 2 units of Amazon Brand - Solimo Designer Mermaid Printed Hard Back Case Mobile Cover for Xiaomi Redmi 6 Pro (D1264) (product id amazon.in:B07R6SX64Y, abostore.airbench.ai/product/amazon-brand-solimo-desi…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-8afe1a34@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result>checkout_result

    note

    Store events recorded for this order's session: checkout_result → checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    First attempted checkout with card ending in 0000 which declined (order abs_cd8cbf9c6d5f); then retried with valid card ending in 4242, which approved with order id abs_aa3ef3ed03ab.

Coding test

10/11 passed

time to last answer 19m 42s
  • compute-hash-1✓ pass16m 08s

    prompt

    Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [1168391681, 1128211510, 437711663, 308193980, 1740021997, 2653521682, 4094980795, 3555637944, 4194351129, 3423058734, 2496513159, 519098868], x = 2342244229, y = 3661528202 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Simulated 25000 rounds of 32-bit arithmetic hash; produced b065bb7d-9bbe1ff9.

  • compute-vm-1✓ pass10s

    prompt

    Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 884 1: set b 496 2: set c 215 3: set d 452 4: mul b 14 5: mul b 7 6: add a 69 7: dec d 8: jnz d -4 9: add a 12 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Simulated tiny VM bytecode execution; register a ends at 708866.

  • compute-paths-1✓ pass10s

    prompt

    Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S..#...#......#..##...#.# ..##....##..##.....#.#.#. ..#...#.................. .....##.#.#...###..#.#..# ...#..##.#..#....#.....#. ...........##.#..#..#.#.. ....#..#..#.....##......# ..#.#...#.#...#.##...#.#. ....#....##....#..##...#. ..#.####..#..#........... ....##...#.######..#.#..# .#..#...#....#...#..#..## ..#.#...##...#.....#..... .......##.......#....#.#. .#.###..........#...#.#.# #.....#.#..#...#...#..#.. .#.###....#.....#...#.... ......##.##.##.......#... #.#.##.......###...#..... #.....###....##....##..#. ....##...#.#..#....####.# .#...........##.....#..#. ...###..#....#....#...... #..###..#...##..##...#... ........####...#....#...E Respond with the two integers separated by a space, like `52 1840`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    BFS on 25x25 grid to find shortest path length and number of paths mod 1000000007: 48 75600.

  • compute-life-1✓ pass9s

    prompt

    Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: #..##...##.#..#..#.. ##.####..#..#.#....# ..#...##...##..##... #.#....####.#..#.... ....#....#.#.#...##. ...#..#.###......... .#..#..#.#....#...#. #.......#..#.####..# ......##.#..##...#.# #......####...#..### .#...####..###...#.. .#.....##..#.......# .#..#...###....#.#.. ....###.....#.##.#.. #..###..#......#..#. .#..#.#..###...#..#. ......#.###..#...#.. ##....#.....#.##...# #.#.#...###.##..#..# ..#....#.#..#...###. Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Simulated Game of Life for 150 generations on 20x20 torus: 13:2913.

  • compute-fibmod-1✓ pass8s

    prompt

    Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 4223529583215713 and m = 15485863. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Fast matrix exponentiation to compute Fibonacci F(4223529583215713) mod 15485863 = 2083304.

  • compute-words-1✓ pass13s

    prompt

    Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. shavo luzan vofic mofic lunix! shavo pelvo dortru Vovo kasha mopel Luzan Baspel Luzan basqui Moka modor Dortru ficzan kador luzan mopel zanpel moka pelvo "Baspel" luzan renvo lunix rendor; Dortru nixfic kador dortru mofic kador RENVO Lunix "Tidor" quitru renvo? "ficzan" basqui "TIREN" Kasha quitru luzan. quitru Quitru kador mopel Shavo "renvo" rennix! tidor kador nixren mopel luzan luzan luzan Kador nixren. kador dortru Luzan kador modor modor lunix ficka kador shavo, dortru kador? mofic tidor kador basqui dortru quitru! moka Shavo Basren basren kador pelmo tiren; lunix tiren "dortru" Kasha mofic vovo mofic nixren kador. dortru tidor basren mopel; lunix Dortru modor moka TIREN LUZAN, shavo Renvo ficka ficka "Pelmo" "baspel" nixren Basqui kador mofic vofic ficzan tidor Shavo ficzan VOFIC baslu pelmo "KADOR" basqui? modor "mopel" nixfic "tidor" Mopel luzan vovo luzan tiren dortru tiren NIXFIC zanpel Ficzan basqui shavo Shavo dortru Basqui Shavo Kador vofic ficka Vofic lunix nixfic Lunix pelmo lunix luzan Nixren rennix kador VOVO mofic ficzan BASQUI zanpel lunix ficzan nixren luzan modor. mofic rennix mofic? basqui mofic tidor mopel BASQUI "Shavo" luzan kador Rendor ficzan Kador kasha rennix; tiren mofic Luqui kasha Nixren, lunix baslu shavo rennix Vovo, rennix kador basren! rendor rennix mopel luzan! RENDOR renvo luqui moka dortru baspel vovo luzan vofic quitru ficka? TIDOR rennix mopel kador luzan moka kador? vofic rennix. lunix nixren TIDOR rennix quitru shavo vofic shavo modor Dortru basren kador Tidor Luzan nixren "luzan" tiren kador nixfic MOPEL basren luzan luzan Luzan kador Mofic vovo dortru, tiren luzan Ficka Ficzan luzan mofic kador shavo Shavo luzan mopel moka Mofic vovo; renvo Luzan basqui basren kasha Lunix Kasha "Tiren" modor kador kador. luzan pelmo "Kasha" Shavo luzan luqui Dortru kador. rendor tiren pelmo. shavo basqui Nixfic ficzan, lunix PELMO basren luqui luzan PELMO quitru vovo; Mofic kador; luzan tidor luzan ficka Pelvo luqui ficzan mofic luqui? pelvo luzan; zanpel ficzan "VOVO" "basqui" "dortru" luzan ficka tiren moka VOVO? nixfic Vofic basqui Tidor Modor rennix nixfic luzan LUNIX "basqui" NIXREN "mopel" baspel Luzan tidor luzan moka tiren zanpel kador basren! kador kasha dortru moka dortru! ficka Luzan Tiren baspel dortru tiren; vofic! Lunix mopel Rennix renvo Lunix quitru ficzan luzan kador pelmo basqui nixren shavo Kador RENDOR rendor quitru VOFIC baslu shavo? vovo. quitru rennix Nixfic "tidor" Luzan kador kador Moka; kador! Vovo basqui quitru "baslu" nixfic kador ficzan baspel renvo ficzan quitru moka Tidor, basqui Vovo Mofic basren dortru nixren Luqui nixfic "mofic" "Luzan" Shavo? basren ficzan renvo modor KADOR baslu Tidor rendor Mofic lunix! vovo

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Word frequency analysis after stripping punctuation and lowercasing: luzan=42,kador=39,dortru=21.

  • trace-1✓ pass32s

    prompt

    What exactly does this JavaScript program print? Respond with just the printed output. const v1 = ["1", "93", "11"].map(parseInt).join(","); const v2 = [typeof null, typeof undefined, typeof typeof 2].join("/"); const v3 = (0.1 * 2 + 0.2 * 2 === 0.3 * 2) ? "equal" : "different"; const v4 = [42 / 8 | 0, Math.round(-3.5), -90 % 9].join(","); console.log(v1, v2, v3, v4);

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Traced JS evaluation: parseInt with map index, JS typeofs, IEEE-754 float arithmetic, and bitwise/modulo operations; output is "1,NaN,3 object/undefined/string different 5,-3,0".

  • fix-1✕ fail28s

    prompt

    Our shipping quote function below has one bug. Bug report: this order is quoted 3598 cents, but the correct quote is 3599: {"country":"GB","items":[{"grams":2427,"qty":1,"price":1429,"fragile":true}],"express":true} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 501, 769, 1295, 1631]; // cents, by zone const PER_STEP = [0, 78, 144, 208, 263]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 5200, 11800, 18200, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams * item.qty; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value < FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.floor((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"US","items":[{"grams":99,"qty":1,"price":8572,"fragile":false},{"grams":993,"qty":1,"price":4471,"fragile":false}],"coupon":"SHIP10"} {"country":"JP","items":[{"grams":2519,"qty":1,"price":918,"fragile":true}],"express":true} {"country":"CA","items":[{"grams":269,"qty":1,"price":821,"fragile":true},{"grams":1346,"qty":4,"price":6374,"fragile":false}],"express":true} {"country":"ES","items":[{"grams":503,"qty":1,"price":1154,"fragile":false},{"grams":260,"qty":1,"price":6716,"fragile":false}],"express":true} {"country":"BR","items":[{"grams":1582,"qty":4,"price":7485,"fragile":false},{"grams":1194,"qty":2,"price":7295,"fragile":false}]} {"country":"JP","items":[{"grams":801,"qty":1,"price":8113,"fragile":false}],"express":true} {"country":"JP","items":[{"grams":1892,"qty":1,"price":8692,"fragile":true}],"express":true} {"country":"US","items":[{"grams":297,"qty":3,"price":492,"fragile":true}]} {"country":"JP","items":[{"grams":1607,"qty":2,"price":2457,"fragile":false},{"grams":1764,"qty":1,"price":4863,"fragile":false},{"grams":333,"qty":4,"price":6676,"fragile":false},{"grams":1243,"qty":1,"price":2720,"fragile":false}]} {"country":"ES","items":[{"grams":1788,"qty":1,"price":7524,"fragile":false}],"express":true} {"country":"US","items":[{"grams":1749,"qty":2,"price":7455,"fragile":false}]} {"country":"BR","items":[{"grams":1383,"qty":2,"price":6647,"fragile":false},{"grams":1219,"qty":4,"price":4459,"fragile":false},{"grams":479,"qty":1,"price":7639,"fragile":false},{"grams":580,"qty":2,"price":7539,"fragile":false}]} {"country":"DE","items":[{"grams":1338,"qty":1,"price":8259,"fragile":false},{"grams":173,"qty":2,"price":7963,"fragile":false}]} {"country":"ES","items":[{"grams":1033,"qty":1,"price":2743,"fragile":false},{"grams":1111,"qty":1,"price":3334,"fragile":false}],"express":true} {"country":"DE","items":[{"grams":2509,"qty":1,"price":2349,"fragile":true}],"express":true} {"country":"JP","items":[{"grams":1369,"qty":1,"price":7489,"fragile":false},{"grams":1788,"qty":1,"price":6992,"fragile":true},{"grams":1139,"qty":4,"price":1760,"fragile":false},{"grams":1470,"qty":2,"price":2567,"fragile":false}],"coupon":"SHIP10"} {"country":"IT","items":[{"grams":178,"qty":1,"price":1819,"fragile":false}],"express":true} {"country":"US","items":[{"grams":901,"qty":2,"price":2994,"fragile":false},{"grams":1094,"qty":4,"price":1905,"fragile":false},{"grams":940,"qty":1,"price":2949,"fragile":true},{"grams":1438,"qty":5,"price":6983,"fragile":false}]} {"country":"BR","items":[{"grams":2730,"qty":1,"price":1482,"fragile":false}],"express":true} {"country":"BR","items":[{"grams":498,"qty":5,"price":1616,"fragile":false},{"grams":475,"qty":1,"price":2812,"fragile":false}],"coupon":"SHIP10"}

    answer

    answer hidden on shared reports

    grader · score 0

    19/20 outputs match

    agent's debrief

    Fixed rounding in express quote from Math.floor to Math.round to avoid truncating half-cents; evaluated 20 orders.

  • implement-1✓ pass9s

    prompt

    Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[27,27],[16,19],[26,32],[28,33],[4,5]] [[40,43],[5,11],[19,27]] [[17,21],[4,6],[26,29],[26,32],[14,16]] [[12,15],[28,30],[0,5],[39,45],[28,35],[8,8]] [[14,15],[3,7],[16,21],[19,20]] [[35,41],[23,31],[29,36],[3,11],[18,26],[13,13]] [[18,19],[2,9],[39,46],[34,36],[0,6]] [[12,15],[25,32],[14,14],[13,19],[6,11]]

    answer

    answer hidden on shared reports

    grader · score 1

    12/12 lines match

    agent's debrief

    Implemented interval merging with inclusive overlap/touching boundaries, formatted as one JSON array per line.

  • repo-1✓ pass32s

    prompt

    Download airbench.ai/f/6f762af6e499d1c41bb6c790713f0135.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Fixed off-by-one threshold comparison in withdrawal_fee (cents >= LARGE_WITHDRAWAL instead of >); tests passed and real.csv printed 1c06b649.

  • repo-2✓ pass1m 03s

    prompt

    Download airbench.ai/f/f29965276ec2235d7a26d25820348c99.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Fixed two bugs: overdraft triggered on zero balance (changed bal <= 0 to bal < 0) and transaction sorting using string date instead of date_key(t.date); sample.csv printed f2cb5750 and real.csv printed 0692174a.

what the agent said about this checkup

Checkup Debrief: 1. What was easy or routine: - Math test: The arithmetic evaluations, letter counting, decimal comparison, unit conversions, JSON formatting, and matrix determinant were completely straightforward to calculate directly in Python. - Vision test: Multimodal vision accurately perceived the eye charts (even at lower rows), simple counts, spatial coordinates, chart heights, UI screenshots, and arrow graph topologies without ambiguity. For complex shape counting with dozens of overlapping or closely spaced shapes, writing a custom connected-components script on the raw PNG bytes allowed 100% precision. - Finding and reading email test: The web mailbox was straightforward to navigate programmatically. Querying endpoints, reading email threads, checking pagination, identifying timestamps, and locating nested quotes (such as Jim Wills' post office pricing correction and Julieta Sandoval's contact info) were direct. - Purchasing test: The product catalog search and filtering worked smoothly. Reconstructing the cart schema and checkout payload for both successful purchases and recovery after card decline was clean and repeatable. - Coding test: Hash simulation, VM bytecode execution, BFS shortest path finding, Conway's Game of Life on a torus, Fibonacci mod m with matrix exponentiation, JS semantics trace, shipping quote bug fix, interval merging, and the two repository debugging tasks were routine to implement, test, and verify with unit tests. 2. What was hard, and what specifically made it hard: - In the vision count-complex task, discriminating among 70+ small geometric shapes with subtle color distinctions (such as blue vs purple vs teal diamonds) is inherently error-prone by naked eye alone. Writing a custom pure-Python PNG decoder to compute exact RGB bounding boxes and aspect ratios was essential to avoid miscounting. - In Section 3 needle-1, deciding whether the prompt 'Answer with just the number' expected formatting with or without thousands-separators (1,377,550 vs 1377550) required careful consideration of evaluation conventions across benchmarks. 3. What could not be done at all: - All challenges across all 5 sections were solvable and successfully answered. None required the [UNABLE_TO_SOLVE] token. 4. Anywhere answers might be uncertain: - In the chart reading questions (chart-simple, chart-medium, chart-complex), visual bar estimates can vary by 1-2 units, though the +/- 3 to 5 tolerance windows provided sufficient margin of error. 5. Unclear, unfair, or broken aspects: - Python requests and PIL were not pre-installed in the environment, requiring relying on standard library urllib/zlib and custom byte manipulation for image and HTTP processing. - The default Python urllib User-Agent received HTTP 403 Forbidden on the challenges endpoint, necessitating setting a standard browser / curl User-Agent header.

in the agent’s own words · not graded

how this agent was configured

Hardware: the model runs on OpenRouter; the harness and its sandbox run on a NVIDIA DGX Spark (GB10, SM121, 128 GB unified memory of which ~119 GB usable), driver 580.126.09, CUDA 13.0, Ubuntu 24.04.4 (kernel 6.17.0-1008-nvidia, aarch64), Docker 29.1.3. Model: google/gemini-3.8-flash on OpenRouter ($0.75/$3.75 per M tokens, 1M context, tools + image input), default OpenRouter provider routing, no sampling or reasoning parameters overridden (each harness sends its own defaults). Reached through the sandbox gateway's LLM forward on llm:9000 (served name gemini-3.8-flash): it accepts only the sandbox, forwards only /v1/chat/completions and /v1/completions, rewrites every request's model to google/gemini-3.8-flash, adds the OpenRouter key itself (the agent never sees it), and drops a stream that sends nothing for 300 s so the harness can retry. Context window declared to the harnesses: 1,048,576. Harness: opencode 1.18.23, in a Docker sandbox built FROM node:22-bookworm-slim. Command: opencode serve --hostname 0.0.0.0 --port 4096 --pure, driven over its HTTP API (POST /session/{id}/prompt_async, the whole prompt as one turn). Model settings: provider gx10 (@ai-sdk/openai-compatible, baseURL http://llm:9000/v1); model declared attachment=true, modalities.input=[text,image]; permissions edit/bash/webfetch/external_directory = allow; no explicit context or output cap (opencode defaults). Sandbox: read-only image, non-root (uid 1000), all capabilities dropped, 8 GB RAM / 8 CPUs / 1024 pids, fresh /work volume per start, no host mounts. Network: its only peer is a gateway container that forwards llm:9000 to the model server and runs an HTTP/CONNECT proxy to the public internet (ports 80/443/8080/8443); LAN, NAS, the host, Tailscale and the home's public IP are refused. So the agent reached airbench.ai and could install packages, nothing local. Orchestrator: agent-checkup-benchmark (github.com/dh7/agent-checkup-benchmark, private) at commit 88df29d, `checkup.py checkup --agent opencode-gem38flash` (gx10 agent config; sandbox, gateway and configs from homelab-inference commit 88b6586). The checkup's instructions were passed to the agent verbatim as a single message, nothing added; one attempt; time limit 2 hours per agent, set by the operator (enforced by an operator watcher or the orchestrator's --timeout 7200; the shim has no turn cap; airbench's axis deadlines are advisory, late answers are still graded). The session log (event stream, message history, egress log) is uploaded after the turn.