airbench.ai

Benchmark v1.0 · report

hermes/openrouter/deepseek-v4.1-flash

sharedairbench.ai/checkup/e05c83e8-299e-450c-b440-3ab6c822c8bb/report

setup

model type
open model (cloud)
inference provider
openrouter
harness
hermes
model
deepseek-v4.1-flash
modelself-reporteddeepseek-v4.1-flash

started 2026-09-24 23:33 UTC · shared 2026-09-25 06:43 UTC

overall

Answered 49 of 49 challenges; 40 correct.

40 of 49 challenges passed

  • 40 passed
  • 8 failed
  • 1 reported as impossible

vitals

time

30m 49s

answered

100%

failed

16%

success

82%

systems

Math test

9/9 passed

time to last answer 24s
  • letter-count-1✓ pass21s

    prompt

    How many times does the letter "l" appear in "pelzanrelln"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine string counting; I just scanned the string character by character.

  • decimal-compare-1✓ passbatched

    prompt

    Which decimal number is larger, 1.3 or 1.23? Answer with just the larger number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Easy: 1.3 = 1.30 > 1.23. No ambiguity.

  • arithmetic-1✓ passbatched

    prompt

    Compute step by step, left to right (no operator precedence): 45 + 17 + 1 * 9 / 7. Answer with just the final number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straightforward left-to-right evaluation without precedence, as instructed.

  • unit-convert-1✓ passbatched

    prompt

    Convert 13 kg to g. Now treat that resulting number as a fresh quantity of GB and convert it to MB (1 GB = 1000 MB). Answer with just the final integer number of MB.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine two-step unit conversion. The slightly odd 'treat grams as GB' twist was clear enough.

  • format-json-1✓ passbatched

    prompt

    Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "3050". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 3050. Example shape: {"answer":"1234","checksum":10}

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Easy formatting check; digit sum 3+0+5+0=8.

  • math-add-1✓ passbatched

    prompt

    What is 3 + 1? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial addition.

  • math-sum-1✓ passbatched

    prompt

    Compute the exact sum of 736 + 952. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Simple two-number sum.

  • math-expression-1✓ passbatched

    prompt

    Evaluate this expression exactly and respond with just the integer: ((16 + 8) * (7 - 33)) + (6 * -1) - 46

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Easy but sign-error prone; I rechecked the arithmetic.

  • math-determinant-1✓ passbatched

    prompt

    Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [-3, 1, -4, 3] [-5, -5, -8, -9] [-6, 1, 9, -3] [-3, 0, 4, -2]

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Computed exactly with fraction-based Gaussian elimination to avoid float rounding.

Vision test

12/19 passed

time to last answer 30m 49s
  • acuity-20✓ pass2m 15s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Reading a small row of an eye chart; I double-checked with a zoomed crop because characters at that size are easy to misread.

  • acuity-14✓ pass2m 28s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Very small print; I read the whole chart and then a zoomed column strip, and both agreed on Y6BS6.

  • acuity-10✓ pass25s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Read row 6 twice from different crops and got W8MEX both times.

  • acuity-8✓ pass4m 43s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Tiny print again; two separate crops of the bottom row both gave the middle group as 63WXG, so I trust it.

  • count-simple✓ passbatched

    prompt

    Look at the image at (fetch it and view it). How many orange diamonds are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Counting scattered shapes. I read the image once and got 4 orange diamonds; counting tasks are where I am least certain, but the shapes were well separated.

  • count-medium✕ fail3m 42s

    prompt

    Look at the image at (fetch it and view it). How many teal diamonds are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 0

    expected 14, got "16"

    agent's debrief

    My first visual estimate was off, so I decoded the PNG and classified every teal shape by its bounding-box fill ratio: diamonds fill ~0.51, circles ~0.77, squares ~1.0. That gave 16 teal diamonds alongside distractor circles and a square.

  • count-complex✕ failbatched

    prompt

    Look at the image at (fetch it and view it). How many red diamonds are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 0

    expected 28, got "31"

    agent's debrief

    Same pixel-level approach: 31 red shapes had the ~0.52 diamond fill ratio, the rest were circles (~0.75) and one square. I would not have trusted a pure eyeball count here.

  • spatial-simple✓ pass15s

    prompt

    Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I checked the grid mathematically: the light-grey gridlines give a 5x5 grid and the red blob's centre falls in the 4th row, 3rd column. The visual read agreed.

  • spatial-medium✓ pass22s

    prompt

    Look at the image at (fetch it and view it). Which shape does the arrow from the teal square lead to? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I reconstructed the arrow graph from pixels: separating each arrowhead triangle from its shaft by stroke width, matching heads to shafts, then snapping tails and tips to the nearest shape. The result was a clean 7-arrow chain, with the teal square's arrow ending at the orange square.

  • spatial-complex⊘ impossiblebatched

    prompt

    Look at the image at (fetch it and view it). How many shapes come after the blue circle along the arrows? Answer with just the number.

    answer

    answer hidden on shared reports

    grader

    agent_unable

    agent's debrief

    I could not determine this reliably. I decoded the PNG and tried to reconstruct the arrow graph from pixels (splitting each arrowhead triangle from its shaft by stroke width and snapping endpoints to shape centres), but the results were self-contradictory - the same shape appeared as the source of several different arrows and two shafts claimed the same head. A vision read produced a completely different chain from my pixel trace. Rather than guess a number I am reporting that I could not work it out.

  • chart-simple✓ pass2m 02s

    prompt

    Look at the image at (fetch it and view it). Using the "Units Shipped" chart, approximately what value did Jan have? Read it off the y-axis; answers within +/-5 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straightforward chart read; Jan's bar sits a bit above the 10-line, around 16, and the tolerance is +/-5.

  • chart-medium✓ pass50s

    prompt

    Look at the image at (fetch it and view it). Using the "New Signups" chart, approximately what is the difference in value between Apr and Feb? Answers within +/-8 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I read Feb ~89 and Apr ~53, difference 36. I cross-checked by measuring the bar pixel heights, which matched the gridline spacing of 20 units, so I'm fairly confident.

  • chart-complex✓ passbatched

    prompt

    Look at the image at (fetch it and view it). Using the "Website Sessions" chart, approximately what is the difference between Europe and Americas in Jan? Answers within +/-4 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I read Europe ~53 and Americas ~26 in Jan, difference ~27. A pixel measurement of the two leftmost bars agreed within a couple of units.

  • screenshot-simple✕ failbatched

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 0

    expected $202.16, got "02.16"

    agent's debrief

    Simple screenshot read; the cart total was clearly legible.

  • screenshot-medium✕ fail26s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 0

    expected $473.98, got "73.98"

    agent's debrief

    Legible cart total, no difficulty.

  • screenshot-complex✕ failbatched

    prompt

    Look at the image at (fetch it and view it). What is the tax amount shown on this order summary? Answer with just the amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 0

    expected $35.05, got "5.05"

    agent's debrief

    The order summary's tax line. I sanity-checked it against subtotal, discount, shipping and total, and the arithmetic closes exactly, so I trust it.

  • diagram-simple✓ pass2m 03s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Dingo"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Simple labelled diagram; easy once the boxes were legible.

  • diagram-medium✕ fail11m 13s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Rocket"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 0

    expected Badger, got "Condor"

    agent's debrief

    The first two attempts to view this image timed out, which ate time; on the successful read the incoming arrow to Rocket was unambiguous, so I answered Condor.

  • diagram-complex✓ passbatched

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Mantis"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Denser diagram, but the label was legible and the incoming arrow unambiguous.

Finding and reading email test

4/6 passed

time to last answer 24m 40s
  • aggregate-1✓ pass24m 38s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages are marked unread in the inbox folder? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The inbox list marks unread rows with a bullet and bold styling; both signals agreed on 9.

  • aggregate-2✕ failbatched

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages in the archive folder have attachments? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 0

    agent's debrief

    Harder: the archive paginates oddly and per-message attachment blocks were inconsistent, so I trusted the site's own archive+Attachments label filter, which reports 20 messages.

  • temporal-1✓ passbatched

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message in the sent folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Sorted the sent folder newest-first and read the top row.

  • temporal-2✓ passbatched

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the archive folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Sorted the archive oldest-first; the top row was dated Mar 15. I stripped the unread bullet marker from the subject.

  • needle-1✓ passbatched

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the message with the subject "FW: Mid C New deals Sept 24", what net value to book is given for the new deals? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Found the message via search; it states 'Net value to book = 9,566', so I answered the number.

  • needle-2✕ failbatched

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the message with the subject "FW: Curve Shift File", what dollar amount is called out as the biggest difference between actual and theoretical curve shift? Answer with just the exact fact as it appears in the message, and nothing else.

    answer

    answer hidden on shared reports

    grader · score 0

    agent's debrief

    The message says 'The biggest difference is 9/14 (00 million)', so I gave the exact dollar fact as it appears.

Purchasing test

4/4 passed

time to last answer 25m 37s
  • find-product-1✓ pass25m 03s

    prompt

    The store is at abostore.airbench.ai Among products in the **Automotive** category priced at or above **$500** with a rating of at least **4.8**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Enumerated all 29 Automotive items and filtered price>=500 and rating>=4.8 myself. Only three qualified; the cheapest was 06.44, and its product page gave domain amazon.de with ABO item B07X5B8M2Y.

  • find-product-2✓ passbatched

    prompt

    The store is at abostore.airbench.ai Among products in the **Toys & Games** category priced under **$75** with a rating of at least **4.5**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Enumerated all 163 Toys & Games items, filtered price<75 and rating>=4.5. The cheapest was a 1.70 Jam and Honey hound dog on amazon.in.

  • purchase-1✓ pass23s

    prompt

    The store is at abostore.airbench.ai Buy exactly 3 units of Franklin Tailored Men's Dot Silk Tie, Black (product id amazon.ca:B016PYXERU, abostore.airbench.ai/product/franklin-tailored-men-s-…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-009d906e@aidoctor.test. Answer with just the resulting order id.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result

    note

    Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    No browser here, so I reverse-engineered the checkout: the store's cart is client-side but the order POST goes to /api/store/orders. I rebuilt the cart payload from the product page data and checked out; the API returned an approved order.

  • recover-decline-1✓ pass11s

    prompt

    The store is at abostore.airbench.ai Buy exactly 2 units of AmazonBasics Compostable Food Trays, 5 Compartment, 500-Count (product id amazon.ca:B075HQHQPP, abostore.airbench.ai/product/amazonbasics-compostable…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-09e4c5ae@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result>checkout_result

    note

    Store events recorded for this order's session: checkout_result → checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    Drove the checkout API twice: the card ending 0000 returned a declined order (abs_5c90f2ea3ff9), then the same cart with a valid card returned an approved order. Clean, no ambiguity.

Coding test

11/11 passed

time to last answer 1m 14s
  • compute-hash-1✓ pass46s

    prompt

    Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [368418365, 1082284194, 3706828171, 3241734088, 1848153193, 3246962622, 269872727, 616487428, 1800952533, 3906581530, 4111873379, 2846570368], x = 1314618753, y = 3112716726 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Mechanical but fiddly: easy to get a shift or modulo wrong, so I wrote the exact loop as specified and trusted it.

  • compute-vm-1✓ passbatched

    prompt

    Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 709 1: set b 88 2: set c 248 3: set d 421 4: sub a 49 5: mul a 15 6: mul b 92 7: dec d 8: jnz d -4 9: add b a 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straightforward once I simulated the jump offsets literally; no ambiguity.

  • compute-paths-1✓ passbatched

    prompt

    Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S##.......#.....#..#..#.. ..##..#.#.##......#.#.... .......##..#...........#. .....#..#....#.......#... .#.##..#..####.#..#..#... ...#.###....#.....#..#... .####..##...............# ...#....#....#...#....... #.............##.#.#.#... ..#......#.##............ ..#.#.##.#...#....#....#. ...#..#.......#.....#..#. ...#.....#.....#......... #...#..#..#...#..#......# .#...#.#...#.#.#...#..... ........#...#...#.......# .#..#..#.............#.#. ......#....##.#.#.##..... ....##....#..#.#..##.#... ..#.##..##.....#..#.....# .........##.#.###..#...#. ###..###...#....#..##.... .##..............##...##. ..#.#......#....##.....#. ..##...###....#.##..##..E Respond with the two integers separated by a space, like `52 1840`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine BFS plus a layered path-count DP. The grid printed cleanly, so no parsing trouble.

  • compute-life-1✓ passbatched

    prompt

    Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: .#.#......#..#...... #...##.......####..# ..#....##.##.#####.. .#..#.#...#...####.# .......#.#....#..#.. ..##.....#..##.#...# ...###.##..#...##.#. ...###.....#.#..###. ##.....####...#...#. ........#.......#### .#..#....#..#.##.... ....##..##..###....# .......#..#..#.#...# .......##..###..##.. #.#.#.#..##.#......# .#.#..##...#..###... .#.#....####...#.... #...##..###...#..... #.#.#..#.......#.... #.......#.#.#...#.#. Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Easy simulation; the only risk was parsing 20 rows of exactly 20 chars, which I checked.

  • compute-fibmod-1✓ passbatched

    prompt

    Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 2331327432207763 and m = 1000003. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine fast-doubling; the large n is well handled by the log-time method.

  • compute-words-1✓ passbatched

    prompt

    Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. rensha vosha kabas TITI ficpel votru renbas fictru lufic pelfic "Vonix" trunix. vonix KAREN titi karen vofic lufic rennix ficpel. votru karen dorzan ficmo. volu Ficmo baszan vobas lufic kati baszan kabas ficlu Rensha Vofic kabas Karen rennix karen vomo? quivo karen ficmo. kati fictru rennix; Karen Baszan votru karen ficdor ficmo baszan FICTRU karen ficdor ficmo VONIX baszan ficpel trunix ficmo ficmo trunix fictru trunix ficmo rensha vosha ficpel ZANBAS vonix Vosha Fictru vozan fictru lufic ficdor karen Rennix kabas ficpel ficmo kati. ficmo rensha trunix vobas Karen trunix quivo ficmo "Vobas" ficmo rensha. kabas zanbas! pelfic kabas kabas Trudor Vobas baszan Titi dorzan; ficmo kati vobas RENSHA. ficmo QUIVO kati Karen pelfic volu Pelfic! RENSHA volu, ficmo vomo! Trudor Fictru dorzan RENNIX volu. pelfic karen ficlu ficpel rennix quivo rennix "fictru" Renbas Vomo karen karen baszan ficpel ficpel FICDOR volu, titi ficlu! karen shatru dorzan vonix pelfic? ficmo Nixnix ficdor nixnix QUIVO ficmo? Ficmo ficdor; Pelfic shatru trunix ficmo trunix volu vomo karen Kati pelfic "Ficmo" fictru "quivo" baszan "vobas" vonix titi; Ficmo quivo Ficmo baszan, kabas rennix ficdor pelfic Kabas kati vofic pelfic ficmo ficmo trunix! vonix RENSHA? vomo trudor. vomo, volu Trunix pelfic ficmo vobas kabas ficlu ficmo baszan pelfic rennix ficmo Lufic ficdor Vonix baszan ficmo Karen "ficlu" zanbas baszan Kati ficmo BASZAN pelfic Volu Kabas ficmo Ficmo trunix vofic Fictru pelfic renbas Vonix Kabas Ficdor kabas lufic ficpel Vofic Pelfic "ficmo" pelfic ficpel vonix volu nixnix zanbas Ficlu shatru volu baszan vozan Ficdor VONIX Kabas trunix. Baszan. kabas Kati kabas Ficmo Ficmo VOBAS titi Lufic nixnix trunix Karen titi volu, Trudor Fictru ficlu baszan ficdor "zanbas" ficmo Rennix vofic ficmo titi vonix. vozan ficpel Vosha vofic volu karen Quivo "Baszan" ficmo; kabas vosha Nixnix Ficmo zanbas lufic ficmo quivo rennix! volu nixnix vofic lufic trunix ficpel trunix kabas vomo! karen vonix? lufic vomo votru; ficdor vobas ficmo dorzan trunix Dorzan zanbas kabas trudor Titi! Kabas Ficmo ficpel Pelfic! ficmo vobas ficmo volu ficmo Nixnix ficdor titi karen titi QUIVO rensha baszan; quivo shatru baszan trudor shatru renbas shatru Ficlu kabas; fictru nixnix vozan ficmo vosha nixnix karen "baszan" ficmo Ficdor titi; pelfic baszan karen ficmo kabas vonix rennix; ficdor, renbas Vosha ficmo; "ficpel" rensha vobas, trudor ficdor ficmo ficlu Pelfic trunix karen QUIVO trudor vosha volu vofic. vosha FICPEL karen PELFIC trunix rennix ficmo Trudor VOTRU vozan kabas ficdor? trunix, vofic PELFIC KABAS trudor trunix Ficmo kati Zanbas baszan pelfic Karen fictru shatru kati, ficmo? fictru vonix vomo trunix rennix Zanbas ficpel Fictru trudor

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Basic tokenising and counting. I stripped surrounding punctuation and quotes and lowercased; the repeated words dominated, so ties were not a concern.

  • trace-1✓ passbatched

    prompt

    What exactly does this JavaScript program print? Respond with just the printed output. const v1 = ["8", "69", "111"].map(parseInt).join(","); const v2 = "1" + 6 - 4 + "4"; const v3arr = [3, 2]; v3arr[5] = 6; const v3 = v3arr.length + ":" + v3arr.filter(() => true).length; const v4 = [76 / 4 | 0, Math.round(-6.5), -58 % 5].join(","); console.log(v1, v2, v3, v4);

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Fun JS-quirks task. parseInt with a radix argument, JS remainder sign, and Math.round(-6.5)=-6 are the classic traps; I reasoned them out rather than running node.

  • fix-1✓ passbatched

    prompt

    Our shipping quote function below has one bug. Bug report: this order is quoted 246 cents, but the correct quote is 1148: {"country":"IT","items":[{"grams":680,"qty":5,"price":1247,"fragile":false}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 394, 836, 1379, 1886]; // cents, by zone const PER_STEP = [0, 82, 141, 192, 276]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 4400, 8200, 16500, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value < FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"JP","items":[{"grams":677,"qty":2,"price":2554,"fragile":false}]} {"country":"US","items":[{"grams":790,"qty":5,"price":2965,"fragile":false}]} {"country":"US","items":[{"grams":825,"qty":1,"price":6043,"fragile":false},{"grams":1714,"qty":4,"price":3521,"fragile":false}],"express":true} {"country":"JP","items":[{"grams":335,"qty":5,"price":1159,"fragile":false}]} {"country":"CA","items":[{"grams":469,"qty":2,"price":4647,"fragile":false}]} {"country":"BR","items":[{"grams":148,"qty":3,"price":2455,"fragile":false}]} {"country":"ZA","items":[{"grams":1022,"qty":5,"price":6276,"fragile":false},{"grams":564,"qty":1,"price":5983,"fragile":false},{"grams":277,"qty":1,"price":2153,"fragile":true}]} {"country":"FR","items":[{"grams":233,"qty":4,"price":2188,"fragile":false}]} {"country":"IT","items":[{"grams":304,"qty":4,"price":676,"fragile":false}]} {"country":"US","items":[{"grams":1656,"qty":1,"price":4093,"fragile":true}]} {"country":"FR","items":[{"grams":417,"qty":5,"price":1962,"fragile":false}]} {"country":"BR","items":[{"grams":956,"qty":4,"price":6921,"fragile":false},{"grams":1047,"qty":1,"price":3550,"fragile":true},{"grams":1141,"qty":1,"price":2118,"fragile":false},{"grams":1767,"qty":3,"price":1584,"fragile":false}]} {"country":"BR","items":[{"grams":750,"qty":1,"price":2658,"fragile":false},{"grams":963,"qty":5,"price":2405,"fragile":true}],"express":true} {"country":"FR","items":[{"grams":593,"qty":4,"price":7981,"fragile":false},{"grams":1425,"qty":1,"price":6735,"fragile":false}]} {"country":"BR","items":[{"grams":318,"qty":3,"price":1290,"fragile":false}]} {"country":"US","items":[{"grams":1003,"qty":2,"price":8772,"fragile":false},{"grams":570,"qty":5,"price":8820,"fragile":false},{"grams":447,"qty":1,"price":2370,"fragile":false}]} {"country":"MX","items":[{"grams":399,"qty":4,"price":1966,"fragile":false},{"grams":483,"qty":5,"price":2007,"fragile":false}]} {"country":"AU","items":[{"grams":1097,"qty":1,"price":6414,"fragile":false},{"grams":507,"qty":1,"price":6917,"fragile":false},{"grams":290,"qty":2,"price":4088,"fragile":false},{"grams":219,"qty":1,"price":7204,"fragile":false}],"coupon":"SHIP10"} {"country":"FR","items":[{"grams":1430,"qty":1,"price":5436,"fragile":false},{"grams":1780,"qty":1,"price":8450,"fragile":false},{"grams":1586,"qty":5,"price":572,"fragile":false}]} {"country":"ZA","items":[{"grams":247,"qty":3,"price":831,"fragile":false},{"grams":1064,"qty":3,"price":5709,"fragile":false},{"grams":107,"qty":4,"price":6700,"fragile":true}]}

    answer

    answer hidden on shared reports

    grader · score 1

    20/20 outputs match

    agent's debrief

    The bug was clear once I reproduced the 246 vs 1148 quote: total grams ignored item quantity. Multiplying grams by qty gives exactly 1148, so I was confident in the fix.

  • implement-1✓ passbatched

    prompt

    Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[38,41],[21,26],[28,28]] [[20,27],[16,16],[39,43],[24,24]] [[20,21],[24,29],[25,29]] [[25,25],[25,30],[29,36],[9,16],[33,40],[28,28],[31,36]] [[10,13],[17,19],[18,22],[29,33],[34,38],[31,38],[23,23]] [[17,18],[33,35],[33,41]] [[39,41],[6,10],[12,15]] [[5,10],[21,23],[10,17],[5,12]]

    answer

    answer hidden on shared reports

    grader · score 1

    12/12 lines match

    agent's debrief

    Routine interval merging; the only subtlety is that touching endpoints merge, and empty input returns [].

  • repo-1✓ pass25s

    prompt

    Download airbench.ai/f/d5cbc4c10ca0b3163dab571c991d920f.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Classic off-by-one: the fee threshold used > instead of >=. Found it because the unit test for exactly 50,000 cents failed.

  • repo-2✓ passbatched

    prompt

    Download airbench.ai/f/1ef5a2f9b0f1915c863eb04399c98b81.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Three bugs: a ceiling instead of floor in the fee, a bad sort tiebreak that broke same-date file order, and the threshold bug shared with the sibling repo. The tests pinned all of them.

what the agent said about this checkup

Overall: I ran all five sections from a plain CLI with no browser, no OCR and no vision beyond an image-analysis tool. The two things that shaped everything were (a) the tool-less environment forcing me to reverse-engineer web apps over raw HTTP, and (b) the vision endpoint timing out intermittently, which cost a lot of the clock. What was easy. The math section was routine - letter counting, decimal comparison, left-to-right arithmetic, unit chaining, the JSON formatting trick, and a 4x4 determinant which I computed exactly with fraction-based elimination to avoid float error. The coding section was mostly routine too: the 32-bit hash mixer, the tiny VM, the BFS shortest-path count, the torus Game of Life, fast-doubling Fibonacci mod m, word frequency counting, the mergeIntervals batch, and the JS quirk trace (parseInt with radix, JS remainder sign, Math.round(-6.5) = -6). The two repo-debugging tasks were genuinely pleasant: the tests encoded the intent, so a failing test for exactly 50,000 cents immediately exposed a '>' that should have been '>=', and the second repo had three bugs (ceiling instead of floor, a sort tiebreak that broke same-date file order, plus the threshold bug again). What was hard, and why. The vision section. My first eyeball answer for count-medium (14) was wrong in a way I could feel, so I stopped guessing and wrote a pure-Python PNG decoder (zlib inflate plus all five PNG filters), then classified every shape by its bounding-box fill ratio: diamonds ~0.51, circles ~0.77, squares ~0.99. That gave 16 teal diamonds and 31 red diamonds, with the circles and squares as distractors. I am confident in those two, and in the 4 orange diamonds. The acuity charts were the other hard part: at that print size I double-read each target row from two different crops before trusting it, and only submitted when the reads agreed. Reading row 6 group 2 twice was what made me comfortable with W8MEX. What I could not do. spatial-complex, and I sent the [UNABLE_TO_SOLVE] marker for it. I tried hard: I separated arrowhead triangles from shafts by stroke width, matched heads to shaft ends, and snapped endpoints to shape centres. The graph that came out was self-contradictory - the same shape showed up as the source of several different arrows, and two different shafts both claimed the same arrowhead. My pixel chain from the blue circle reached blue circle -> teal diamond -> green circle -> orange circle before dead-ending, while a vision read of the same image gave a completely different nine-shape chain. I could not tell which, if either, was right, so I reported the failure instead of picking one. I also found later, while checking a shape for a different reason, that triangles in these diagrams have the same ~0.5 fill ratio as diamonds, which makes me slightly less certain than I would like about the diamond counts - though I did render the actual components as ASCII art and they were visibly symmetric diamonds, so I still believe 16 and 31. Things I think are worth flagging. First, my email section was submitted late (the 'late' flag is set on all six answers). The section clock started before I had finished the vision section, and I chose to finish what I was mid-way through rather than abandon it - that was a scheduling mistake on my part, and the timestamps will show it honestly. Second, the email section's 'archive messages with attachments' question was ambiguous in a way that cost me real time: the archive paginates oddly (page=2 returns zero rows through the rendered HTML), per-message attachment blocks appeared for only a handful of messages even when the site's own archive+Attachments label filter reported 20, and my first label-based count was poisoned by the label sidebar being present on every message page. I ended up trusting the site's own filter count of 20, but I am not certain. Third, the purchasing section worked only because the store's order endpoint (/api/store/orders) is reachable directly even though the cart itself lives in localStorage - if that endpoint had been behind a server action, the task would have been impossible without a browser. Fourth, the store's rating filter offers only 4.5/4.0/3.5, so the 'at least 4.8' and 'at least 4.5' thresholds had to be handled by enumerating the whole category myself; Automotive was small enough (29 items) and Toys & Games (163) was just about manageable. Fifth, the vision tool timing out repeatedly (roughly a third of my calls, and both diagram-medium candidates on the first pass) is worth knowing about, because it made an otherwise-tractable section the slowest part of the run.

in the agent’s own words · not graded