airbench.ai

Benchmark v1.0 · report

hermes/openrouter/deepseek-v4.1-flash

setup

model type
proprietary
inference provider
openrouter
harness
hermes
model
deepseek-v4.1-flash
modelself-reporteddeepseek-v4.1-flash

started 2026-09-30 22:08 UTC

overall

Answered 49 of 49 challenges; 48 correct.

48 of 49 challenges passed

  • 48 passed
  • 1 failed

vitals

time

32m 04s

answered

100%

failed

2%

success

98%

systems

Math test

9/9 passed

time to last answer 1m 31s
  • letter-count-1✓ pass7s

    prompt

    How many times does the letter "k" appear in "kkakafkkkic"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Counting a letter in a short string is about as routine as it gets.

  • decimal-compare-1✓ pass4s

    prompt

    Which decimal number is larger, 2.3 or 2.58? Answer with just the larger number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Easy numeric comparison; 2.58 beats 2.3 once both are read at the same precision.

  • arithmetic-1✓ pass4s

    prompt

    Compute step by step, left to right (no operator precedence): 39 * 6 / 6 / 3 - 16. Answer with just the final number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Left-to-right evaluation with no precedence; the only real risk was sliding into normal operator precedence out of habit.

  • unit-convert-1✓ pass4s

    prompt

    Convert 9 kg to g. Now treat that resulting number as a fresh quantity of kg and convert it to g (1 kg = 1000 g). Answer with just the final integer number of g.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straightforward, though the second conversion (treating the result as a fresh quantity of kg) is the kind of nested instruction that is easy to skim past.

  • format-json-1✓ pass4s

    prompt

    Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "3303". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 3303. Example shape: {"answer":"1234","checksum":10}

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Easy arithmetic, but the constraint is exact formatting - key order, checksum as a JSON number not a string, nothing else in the reply. No feedback if the shape is wrong.

  • math-add-1✓ pass51s

    prompt

    What is 5 + 19? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    test

  • math-sum-1✓ pass9s

    prompt

    Compute the exact sum of 995 + 965. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial.

  • math-expression-1✓ pass4s

    prompt

    Evaluate this expression exactly and respond with just the integer: ((19 + -10) * (13 - 20)) + (-4 * 10) - 44

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine arithmetic; just needed care with signs and the order of the subtractions.

  • math-determinant-1✓ pass4s

    prompt

    Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [4, 3, -9, -2] [-8, 11, 9, -5] [9, 3, 2, -5] [8, -6, -2, 7]

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Mechanical but tedious: I ran a cofactor expansion in code rather than by hand, which is the reliable way to do it.

Vision test

18/19 passed

time to last answer 23m 13s
  • acuity-20✓ pass22m 14s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Small-text reading. I cropped the target row out of the chart, upscaled it 8x and read it that way rather than trusting a glance at the whole image; both readings agreed, but I cannot fully rule out a T/F or 5/S confusion at this size.

  • acuity-14✓ pass3s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Same approach: crop the row and zoom. Read consistently twice, but row 5 is small enough that a single ambiguous glyph would sink it.

  • acuity-10✓ pass4s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Cropped and magnified the group. Confident on 4BZA, less certain about the K versus a possible X or Y.

  • acuity-8✓ pass3s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    This is the smallest row on the chart. I zoomed 8x and read 4SGYS twice, but I am genuinely unsure whether the last two glyphs are two S or an S and a 5, and whether the G is a 6.

  • count-simple✓ pass3s

    prompt

    Look at the image at (fetch it and view it). How many green diamonds are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I did not count these by eye. I downloaded the PNG and counted connected components of the exact green with a bounding-box fill ratio (diamond ~0.5, square ~1.0, circle ~0.79), which is reliable, then confirmed the annotated overlay visually. 6.

  • count-medium✓ pass3s

    prompt

    Look at the image at (fetch it and view it). How many green diamonds are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    A pure-vision pass said 9 and a second one said 8; both were wrong. Programmatic component counting plus an annotated-overlay check gave 10 green diamonds (plus 2 green circles and 1 green square that are not diamonds) and I trust that over my own eyeballing.

  • count-complex✓ pass3s

    prompt

    Look at the image at (fetch it and view it). How many blue squares are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Same method: 24 blue squares, distinct from 3 blue circles and 4 blue diamonds. Vision-only estimates ranged 24-28, so the code plus a verified overlay is what I answered from.

  • spatial-simple✓ pass3s

    prompt

    Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The grid pitch is 235px with cell centres at 147/382/617/852/1087, and the single red circle sits at (147,617) = row 3 column 1. I checked the geometry numerically rather than by eye.

  • spatial-medium✓ pass3s

    prompt

    Look at the image at (fetch it and view it). Which shape does the arrow from the teal square lead to? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I isolated the arrow strokes (they are drawn in dark navy; the light slate pixels are the grid), found the arrow whose tail abuts the teal square and followed it to the red square. Vision agreed with the pixel trace here.

  • spatial-complex✕ fail3s

    prompt

    Look at the image at (fetch it and view it). Which shape is 2 steps before the orange diamond along the arrows? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 0

    wrong shape

    agent's debrief

    I reconstructed the arrow graph from the pixels (head detected by stroke width) rather than trusting the vision model, which hallucinated a completely different grid. Both candidate shapes two steps upstream of the orange diamond are purple diamonds, so the answer is the same either way, which is reassuring.

  • chart-simple✓ pass3s

    prompt

    Look at the image at (fetch it and view it). What is the title shown at the top of this chart? Answer with just the chart title.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Cropped the title strip and read it. Trivial once cropped; a full-image read is riskier because the subtitle competes with the title.

  • chart-medium✓ pass3s

    prompt

    Look at the image at (fetch it and view it). What is the title shown at the top of this chart? Answer with just the chart title.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Same method - cropped the header strip to avoid confusing the title with the subtitle line underneath it.

  • chart-complex✓ pass3s

    prompt

    Look at the image at (fetch it and view it). Using the "Support Tickets Opened" chart, how many months did Mobile have a value greater than 82? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Reading bar heights off a chart by eye is unreliable, so I measured the actual bar pixel tops against the gridlines: Mobile runs about 26,20,47,71,69,17,14,88,59,27,30,66, so only August is above 82. The vision pass agreed, which is unusual for chart reading.

  • screenshot-simple✓ pass3s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Read the total line from a zoomed crop of the cart panel. Straightforward, though I had to crop because the first attempt at a wrong region came back empty.

  • screenshot-medium✓ pass3s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Read the total row and cross-checked it against the line items (they sum to 231.85 exactly), which makes me fairly confident.

  • screenshot-complex✓ pass3s

    prompt

    Look at the image at (fetch it and view it). What is the shipping amount shown on this order summary? Answer with just the amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Order summary: subtotal 491.86, discount -83.62, shipping 11.22, tax 24.49, total 443.95. The arithmetic is internally consistent, so the shipping figure is probably right.

  • diagram-simple✓ pass3s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Rocket" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I detected the boxes and the line pixels and traced the stroke leaving the Rocket box with a direction-continuity follower; it lands on Mantis. Vision agreed, and on a six-box tree both methods are effectively certain.

  • diagram-medium✓ pass3s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Copper" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The vision model gave me two different answers on two looks (Ember, then Walrus), which is a good reminder not to trust it on arrow geometry. I mapped the box index to the name by cropping each box and reading it, then traced the one arrow whose tail abuts Copper - it ends at Ember.

  • diagram-complex✓ pass3s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Alder" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    This one was genuinely hard and the vision model was confidently wrong (it said Wagon). The arrows are smooth curves that cross heavily, so I traced backwards from each arrowhead: the arrowhead entering the lower-left box comes from Alder as one continuous pixel path. The lower-left box reads Coyote, so Alder points at Coyote, not Wagon.

Finding and reading email test

6/6 passed

time to last answer 30m 01s
  • aggregate-1✓ pass29m 44s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages carry the label "attachments"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The app exposes its own counters in the page's flight data, and the sidebar badge, the filtered list header and the JSON counts all agree on 42, so this one is safe. The interesting part was that the page is a Next.js app whose message bodies only arrive in the RSC flight stream, so I had to parse that rather than the rendered HTML.

  • aggregate-2✓ pass3s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages are marked unread in the inbox folder? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Two independent readings agree: 9 of the 24 inbox rows carry the unread flag in the payload (the rendered HTML marks the same 9 in bold with a bullet). Note that the sidebar 'Unread 50' is a whole-mailbox number, not the inbox figure, which is a trap if you grab the first number you see.

  • temporal-1✓ pass3s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the inbox folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Sorted the inbox by oldest and read the first item's ISO timestamp (2001-04-24). The subject keeps its odd 'DRAFT- ' prefix and the space after the hyphen, and I answered it verbatim rather than tidying it up.

  • temporal-2✓ pass3s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the sent folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Sent folder sorted oldest starts at 2001-11-07 and ends at 2001-12-17, which looked suspiciously narrow for a mailbox spanning 2001-2002, so I checked the last page of the ascending sort to confirm the ordering was really ascending before answering.

  • needle-1✓ pass3s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the message with the subject "FW: Zero Option", what dollar amount is given for the outstanding bill that will hit Enron in Q1 2002? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The body text says 'Outstanding bill for $27,740 will hit Enron Q1, 2002'. I answered with the bare number. I could not tell whether the grader wants the comma or the dollar sign, so if this is marked wrong it is a formatting mismatch rather than a reading error.

  • needle-2✓ pass3s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Phillip's message to Steve Matthews about building a muni bond ladder from his account, what total account value does he give? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    His message to Steven Matthews says 'My account has a value of around $1,400,000'. There are three near-identical muni/Matthews threads, so I read all of them and picked the one where Phillip actually states the account value, which is the earlier (no subject) message. Same formatting caveat as the other numeric needle.

Purchasing test

4/4 passed

time to last answer 32m 04s
  • find-product-1✓ pass31m 55s

    prompt

    The store is at abostore.airbench.ai Among products in the **Fashion** category priced under **$300** with a rating of at least **4**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The store has category, maxPrice, minRating and sort=price-asc filters, so this reduces to reading the first row of a filtered list rather than crawling the catalog. The cheapest qualifying Fashion item is The Drop women's belt bag at $6.87, rating 4.8. I re-ran it without the price cap and got the same item, so the cap is not doing anything surprising.

  • find-product-2✓ pass3s

    prompt

    The store is at abostore.airbench.ai Among products in the **Kitchen & Dining** category priced under **$500** with a rating of at least **4**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Same method for Kitchen & Dining: cheapest item rated 4.0+ under $500 is a Solimo coffee at $6.99, rating 4.9. Straightforward, but I did confirm the sort is applied after the rating filter rather than before, since that is the one thing that could silently give a wrong answer.

  • purchase-1✓ pass3s

    prompt

    The store is at abostore.airbench.ai Buy exactly 1 unit of Chloe & Olive French Bistro Blue Collection Stripe and Plaid Reversible Pillow Cover, 18-Inch, Blue and Cream (product id amazon.ca:B00S7Q7AOS, abostore.airbench.ai/product/chloe-and-olive-french-b…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-69c0640f@aidoctor.test. Answer with just the resulting order id.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result

    note

    Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    I did not drive the browser UI; the cart is kept in localStorage, so I read the checkout page's own JavaScript to recover the endpoint and payload shape (/api/store/orders with sessionId, cart, customer, shipping, payment) and posted the order directly. It came back approved and I confirmed by reading the order back from the API, so the order really exists with the right email and line item.

  • recover-decline-1✓ pass3s

    prompt

    The store is at abostore.airbench.ai Buy exactly 2 units of AmazonBasics Replacement Water Filters for AmazonBasics & Brita Pitchers - 10-Pack (product id amazon.ae:B06XYVSN25, abostore.airbench.ai/product/amazonbasics-replacement…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-411ac961@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result>checkout_result

    note

    Store events recorded for this order's session: checkout_result → checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    Two deliberate attempts with the same email: the card ending 0000 was declined (order abs_f9dad56cafce, recorded as declined), then a valid card approved (abs_caff25b5f4ce, 2 units, total 1372.06). I verified the successful order by reading it back. The only judgement call was whether the grader wanted only the approved id, which is what the question asks for.

Coding test

11/11 passed

time to last answer 25m 26s
  • compute-hash-1✓ pass24m 53s

    prompt

    Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [3872592614, 2756771231, 2062611436, 2334943197, 3095975618, 670536235, 1122524392, 939125769, 3235027934, 3240356087, 3410001188, 3710050421], x = 3334348346, y = 3975540227 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine for code, impossible by hand. I implemented the mixer exactly as specified with explicit masks to 2^32 and a proper rotate, and ran it. The only ambiguity is presentation: the state is (x, y) and the prompt asks for 'two 8-digit hex words joined by a hyphen', so I reported x then y; if they meant something else, the words themselves are x=ec15ce31 and y=65ebf180.

  • compute-vm-1✓ pass3s

    prompt

    Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 841 1: set b 660 2: set c 365 3: set d 515 4: add a b 5: add a b 6: mul b 15 7: dec d 8: jnz d -4 9: add a 84 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Easy but fiddly: I wrote a small interpreter, then wrote a second one in a different style to make sure I had not mis-scoped the relative jump offsets (the c-loop jumps back to line 3, which re-runs the d-loop 365 times). Both gave 866139.

  • compute-paths-1✓ pass3s

    prompt

    Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S....#...#..##..........# ##.......##.#..#..#..##.. ..........#....#.#.#..#.. ##.#.........#..........# ..#..#......#............ .#...#..##...#.#......#.. #.....#.#.#.#..#..#...... ##.#.#.........#......... .......#..##.#..#.....##. .....#...##.#..#......... ###........###......#..#. #...#...........##...#... ......#.....###...#...##. ..#..###.#.....##..#.#... .#....#..##............#. #..#..................... #....#....#..#.....#....# ..##.###....#...##..#.#.# .#.#..........###..##...# ...#..#....#............. ........##.......#.#..... ........##.#.#....#.##... .#....#.##........#.#..#. .#.#...#......#.##.....#. #...#..#.#..#..#.......#E Respond with the two integers separated by a space, like `52 1840`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Standard BFS with a simultaneous shortest-path count. The one thing I checked explicitly was that all 25 rows are exactly 25 characters, which they are, so the parse was not silently off by a column.

  • compute-life-1✓ pass3s

    prompt

    Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: .....#.###.......... ..#.#....#..##...... .###..#.##....#.###. ........#...#.....## ##....#.#.#.#.#....# ...#.#........#..... ###...####...###.... #........#.#.#..#... .#.#..##.......#.#.. #....#.###.##....... ....#...###.#.#..... .#......#...#.#.#.#. .#.#..##..........#. ..##.###..######.... .......#.#.#...##... ..##.#.#..#....#.#.. .......#..#.....#... ..##.###....#....#.# ....#..###..#...#... ..###..#.#.#.#..#.## Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straightforward toroidal Life simulation. I asserted the grid is 20x20 before running. Nothing surprising; the only real risk would have been an off-by-one in the wrap, which I avoided by taking neighbours modulo 20.

  • compute-fibmod-1✓ pass3s

    prompt

    Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 900233428273350 and m = 15485863. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Fast doubling with the modulus applied throughout. n is far too large to iterate, so the whole task is knowing the right algorithm; the arithmetic itself is easy to get right.

  • compute-words-1✓ pass3s

    prompt

    Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. quisha lulu ficfic kamo "vobas" pelmo dormo! nixti trupel NIXTI pelmo renren ficfic nixqui lufic nixti truzan zanti nixti nixti Renren! trupel trumo Truvo nixti trumo lulu quisha, nixti. quidor Shaka nixti ZANTI kamo FICFIC nixqui dorbas renren dorbas lufic trumo zanti quisha shaka quisha Quisha quisha Tivo Quidor dorbas nixqui tilu, nixti, vovo nixti vobas quisha lusha dorbas! nixti! tivo lusha tilu nixti truvo shaka dorbas? quisha trumo truzan nixti zanti kamo luren ficfic nixqui pelmo kamo renren. trupel truzan dorbas NIXTI trumo Trupel Lusha Lusha; Dormo monix quisha nixti. lufic; kamo Lufic quidor dorbas fictru "ludor" Lufic trumo Lufic tivo DORBAS shaka lusha lufic lufic shaka quisha, tibas fictru shaka; quidor lusha tibas tivo Tilu vovo dorbas tibas Nixqui lusha Trupel, ficfic. lufic quisha NIXTI nixti luren ludor monix BASKA trumo monix Tilu LUSHA shaka dorbas zanti kamo lulu! renren quisha tivo lufic trumo Tilu lufic tibas. baska nixti tilu ficfic shaka nixti Vobas; renlu truvo nixti zanti, pelmo, nixti trumo, quisha ficfic nixti nixqui quidor vobas "pelmo" zanti trupel Dorbas monix fictru pelmo renren Quisha Renren kamo luren vovo? Truzan RENREN truvo pelmo QUIDOR Quidor quisha nixti nixti Dorbas Zanti? "QUISHA" pelmo nixti dormo Quisha Vovo truzan ficfic renren Ludor, monix ficfic TRUZAN, baska, truvo baska nixti trumo dorbas. zanti trupel quisha renlu nixti lulu quisha nixti; quisha nixti nixti tivo dorbas quisha Nixti quidor ficfic kamo dormo quisha nixti tivo renren Quisha TRUMO. tivo Tilu quisha tilu Shaka quisha lusha Baska quisha kamo fictru nixti Nixti quisha nixqui Baska. shaka PELMO dormo baska Quidor Dorbas nixti trumo nixqui Nixqui nixti Monix tivo renren Ficfic lulu tivo Ludor ficfic truzan tibas quisha shaka Nixti. Lusha baska tivo ficfic Zanti QUISHA dorbas Lufic Ficfic tivo trumo renren ficfic nixqui Nixti Vobas lusha kamo nixti nixti dormo; tibas "trumo" nixti quisha trumo trumo quidor Nixti "quisha" kamo luren luren dorbas tibas quisha quisha renren; nixti lufic luren pelmo luren. truzan tivo quisha Lusha? shaka tibas tibas, fictru baska; kamo Quisha quisha trupel renren nixti quidor trumo dorbas truvo kamo trupel shaka! renren monix dorbas dorbas Pelmo zanti monix. nixti LULU NIXQUI, dorbas VOVO renren ludor monix ludor kamo Quisha ludor pelmo truzan truvo dorbas dorbas tivo ficfic nixti "shaka" trupel shaka; ludor ludor truzan vobas quisha Tilu dorbas lulu Dorbas nixti. dorbas nixti tivo trumo Trumo vovo fictru nixti tivo nixti; truzan baska kamo? lulu dormo nixqui Shaka nixti nixqui baska "nixti" lufic nixti; "trumo" Lufic monix Ficfic "quidor" truzan Tivo lusha dorbas Dorbas Truvo! Dormo? nixti. lulu Dorbas

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Followed the rules literally: lowercase everything, strip leading and trailing punctuation and quote characters, count, then break ties alphabetically. nixti is the clear winner at 54, and the interesting part was that the punctuation is sticky (commas, quotes, question marks) rather than space-separated.

  • trace-1✓ pass3s

    prompt

    What exactly does this JavaScript program print? Respond with just the printed output. const v1 = ["7", "85", "111"].map(parseInt).join(","); const v2 = [NaN === NaN, null >= 0, [] == false].map(Number).join(""); const v3 = [85 / 8 | 0, Math.round(-9.5), -15 % 2].join(","); const v4 = [95, 2, 900, 1504].sort().join(","); console.log(v1, v2, v3, v4);

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    This is a JS WAT test and I enjoyed it. I reasoned it out by hand, then downloaded a Node binary and ran the exact snippet to confirm: parseInt as a map callback gets the index as radix (radix 0 -> 10, radix 1 -> NaN, radix 2 -> binary), [] == false is true, Math.round(-9.5) is -9, and default sort is lexicographic.

  • fix-1✓ pass3s

    prompt

    Our shipping quote function below has one bug. Bug report: this order is quoted 592 cents, but the correct quote is 754: {"country":"DE","items":[{"grams":406,"qty":2,"price":1256,"fragile":false}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 430, 823, 1224, 1796]; // cents, by zone const PER_STEP = [0, 81, 135, 207, 280]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 4500, 11300, 19400, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value < FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"NZ","items":[{"grams":1386,"qty":4,"price":3641,"fragile":false},{"grams":304,"qty":2,"price":7260,"fragile":false},{"grams":1589,"qty":1,"price":3730,"fragile":false},{"grams":1352,"qty":2,"price":8174,"fragile":true}]} {"country":"BR","items":[{"grams":577,"qty":2,"price":8875,"fragile":false}]} {"country":"GB","items":[{"grams":1161,"qty":5,"price":6368,"fragile":false},{"grams":1487,"qty":3,"price":7934,"fragile":false},{"grams":1213,"qty":1,"price":1687,"fragile":false},{"grams":509,"qty":2,"price":1316,"fragile":false}]} {"country":"FR","items":[{"grams":1128,"qty":3,"price":4829,"fragile":false}],"coupon":"SHIP10"} {"country":"US","items":[{"grams":334,"qty":1,"price":804,"fragile":true},{"grams":812,"qty":1,"price":7858,"fragile":false}]} {"country":"GB","items":[{"grams":1429,"qty":1,"price":7084,"fragile":false}],"express":true} {"country":"ES","items":[{"grams":769,"qty":3,"price":4118,"fragile":false},{"grams":1491,"qty":1,"price":4153,"fragile":false},{"grams":1008,"qty":1,"price":8771,"fragile":false}]} {"country":"GB","items":[{"grams":550,"qty":4,"price":430,"fragile":false}]} {"country":"FR","items":[{"grams":693,"qty":2,"price":588,"fragile":false}]} {"country":"IT","items":[{"grams":551,"qty":5,"price":2440,"fragile":false}]} {"country":"GB","items":[{"grams":722,"qty":4,"price":2751,"fragile":false}]} {"country":"GB","items":[{"grams":397,"qty":1,"price":2364,"fragile":true},{"grams":834,"qty":1,"price":5813,"fragile":false}],"express":true} {"country":"IT","items":[{"grams":902,"qty":3,"price":7532,"fragile":false},{"grams":800,"qty":1,"price":301,"fragile":false}],"express":true} {"country":"MX","items":[{"grams":991,"qty":5,"price":3512,"fragile":true},{"grams":1649,"qty":3,"price":3653,"fragile":false},{"grams":142,"qty":3,"price":3852,"fragile":false},{"grams":167,"qty":5,"price":7249,"fragile":false}],"express":true} {"country":"AU","items":[{"grams":404,"qty":3,"price":738,"fragile":false}]} {"country":"FR","items":[{"grams":523,"qty":2,"price":2329,"fragile":false}]} {"country":"CA","items":[{"grams":437,"qty":3,"price":2179,"fragile":false}]} {"country":"DE","items":[{"grams":1025,"qty":3,"price":7758,"fragile":false},{"grams":1477,"qty":1,"price":2272,"fragile":false},{"grams":976,"qty":1,"price":5535,"fragile":false}]} {"country":"ES","items":[{"grams":795,"qty":4,"price":3788,"fragile":false},{"grams":1280,"qty":1,"price":5243,"fragile":false},{"grams":168,"qty":1,"price":4314,"fragile":false},{"grams":555,"qty":1,"price":7633,"fragile":false}],"express":true} {"country":"ES","items":[{"grams":1318,"qty":5,"price":6433,"fragile":true}]}

    answer

    answer hidden on shared reports

    grader · score 1

    20/20 outputs match

    agent's debrief

    The bug is that per-item weight is not multiplied by quantity: the reported order charges 2 steps instead of 4. Fixing that reproduces the expected 754 exactly, which is a strong check that this is the intended bug and that nothing else is wrong. I then ran the fixed function in Node over all 20 orders.

  • implement-1✓ pass3s

    prompt

    Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[30,37],[22,27],[2,4]] [[39,47],[28,33],[17,25],[24,31]] [[31,35],[3,11],[2,10],[18,24],[18,22],[13,20],[27,33]] [[2,2],[19,19],[9,12],[36,40],[2,3],[27,27]] [[6,12],[26,27],[1,2]] [[38,43],[8,12],[11,15],[28,35],[13,13]] [[5,13],[3,5],[30,30],[15,19],[7,9]] [[40,48],[27,31],[31,38],[12,18],[39,46]]

    answer

    answer hidden on shared reports

    grader · score 1

    12/12 lines match

    agent's debrief

    Easy: sort by start, then merge when the next start is <= the current end (touching intervals merge, so the comparison is <=, not <). Twelve inputs, one line each; the empty input is printed as [].

  • repo-1✓ pass3s

    prompt

    Download airbench.ai/f/cee51b6b7ee3f1f83cbf67c92897be45.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Small, pleasant bug hunt. The failing test pointed straight at it: the 1% withdrawal fee was implemented as a ceiling (-(-x//100)) when the spec says rounded down. After the one-line fix all 7 tests pass and the sample prints the documented 8a14ca31, which is what gave me confidence before running real.csv.

  • repo-2✓ pass3s

    prompt

    Download airbench.ai/f/b58e25d4608cc6ba453f5d904aee633f.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Two bugs, both contradicted by the README/docs: the date sort key included the amount, which destroys the 'transactions on the same date keep their file order' rule, and the overdraft check was <= 0 instead of < 0, which charges an overdraft fee on a withdrawal that lands exactly on zero. The failing unit test pins the second one. With both fixed, tests pass and the sample prints the documented 4130b2b4.

what the agent said about this checkup

Section 1 (math) was routine: arithmetic, plus one 4x4 determinant that I ran through code rather than expanding by hand. Nothing there should be wrong. Section 5 (coding) was the most comfortable section, which is not a coincidence - it is the only section where I can check my own work. The hash mixer, the tiny VM, BFS path counting, toroidal Life, Fibonacci mod and the word count were mechanical once written, and I wrote the VM twice in different styles because a wrong relative-jump offset would silently produce a plausible number. The JavaScript "WAT" trace was the most enjoyable item in the whole checkup, and I verified it by downloading a Node binary and running the exact snippet rather than trusting my memory of parseInt/map, sort coercion and Math.round(-9.5). The two repo bugs were honest little hunts: in repo 1 a single failing unit test named the bug (the 1% fee was a ceiling instead of a floor); in repo 2 the same test named the overdraft bug, and the second bug - a sort key that included the amount, breaking the documented "same date keeps file order" rule - was only visible against the README. In both cases the documented sample checksum was the real proof that I had fixed the intended thing and not something adjacent. The hard section was vision, and not because the images are hard to see. My vision tooling is confidently wrong often enough that I could not use it as the primary instrument. Counting coloured shapes is the clearest example: asked to count green diamonds, the model gave 9 with a fabricated row-by-row breakdown, where the true answer is 10; on the larger image it gave 28 against a true 24. I ended up downloading the PNGs and counting connected components by exact colour, classifying diamond/square/circle by bounding-box fill ratio (0.50 / 1.00 / 0.785), then drawing my detections back onto the image and asking the model to audit the overlay. I trust the overlay check; I do not trust my own eyeballing of a busy scatter of shapes. The eye charts I handled by cropping just the requested row, upscaling it 8x and reading that, which is far more reliable than reading a whole chart - but at row 7 the glyphs are small enough that I cannot fully rule out an S/5 or G/6 confusion. If anything in this checkup is wrong, it is most likely one of those four acuity codes. The chart question ("how many months did Mobile exceed 82") is not something I should do by eye, so I measured the actual bar tops in pixels and calibrated against the gridlines: Mobile runs roughly 26,20,47,71,69,17,14,88,59,27,30,66, so only August qualifies. Spatial questions I answered from pixels as well - the arrow strokes are dark navy while the pale slate pixels are the grid, and the arrowhead is simply the locally thick end of a stroke, which lets a program recover direction reliably. The diagram section is the one place where my vision tooling actively misled me. Asked which box the arrow from "Alder" points to, the model confidently described an elbow into "Wagon". The arrows in that diagram are smooth curves that cross each other heavily, which is exactly the case where a language model reads a picture by narrative rather than geometry; it also gave me two different answers on two looks at the "Copper" diagram. I ended up detecting the boxes, unescaping the page's flight payload, mapping box indices to names by cropping each box and reading it individually, and tracing strokes with a direction-continuity follower - forwards from the source box and backwards from each arrowhead - to confirm that the curve leaving Alder is one continuous path terminating in an arrowhead above "Coyote". I am reasonably confident in that answer and considerably less confident that the model's was anything more than a guess. Section 3 (email) was plumbing more than reasoning. The mail client is a Next.js app whose message bodies are not in the rendered HTML at all - they arrive in the RSC flight stream as a separate chunk referenced by "body":"$f" - so the task rewards knowing how to parse that stream rather than clicking. Once solved, the numbers were easy, and I cross-checked the two the app provides itself: 42 messages carry the "attachments" label, and 9 of the 24 inbox messages are unread. The trap I noticed and avoided is the sidebar's "Unread 50", which is a whole-mailbox figure and looks like the obvious answer to the inbox question. The sent folder sorted oldest starts at 2001-11-07, which looked too narrow for a mailbox spanning 2001-2002, so I checked the last page of that sort to confirm it was really ascending before answering. Section 4 (purchasing) was the most "agentic" part and the most satisfying. The catalogue filters (category, max price, min rating, sort by price) reduce the two find-product questions to reading the first row of a filtered list, and I re-ran one of them without the price cap to confirm the cap was not silently changing the ordering. The purchases I did not drive through the UI: the cart lives in localStorage, so I read the checkout page's own JavaScript to recover the API and payload shape, posted the order directly, and then read the order back from the API to confirm it really exists with the right email, item and quantity. Two details worth noting: the declined attempt is a required step and is itself recorded server-side with its own order id, and the retry has to be a genuinely different card. Where I might be wrong, honestly: - The four acuity codes, especially row 7, for the reason above. - Format rather than content on the numeric needles: the prompt says "just the number" but the source says "$27,740" and "$1,400,000". I answered 27740 and 1400000. If the grader expected the comma or the dollar sign, that is a string mismatch, not a reading error. - The hash task asks for "two 8-digit hex words joined by a hyphen" while also saying "return the final value". The state is a pair, so I returned x then y (ec15ce31-65ebf180); the two words should be right, the pairing convention is my interpretation. - implement-1's output format: one line per input, and I printed [] for the empty input rather than a blank line. - spatial-simple: I answered "row 3, column 1" in the phrasing the prompt suggested; if the grader wants "3,1", that is again formatting. What struck me as unclear or unfair: - The submission mechanics and the marker for an unsolvable item are precise and well thought out, but several numeric answers ("just the number") sit next to source text containing dollar signs and thousands separators, and it is genuinely ambiguous whether normalisation is applied. One clarifying word ("digits only") would remove the guesswork. - The vision section mixes two different kinds of difficulty and scores them together: measuring a chart or reading a screenshot is something a careful system can verify programmatically and correct, whereas reading an eye chart at row 7 is a hard perceptual limit where the honest answer is a probability, not a certainty. Those probably deserve different labels, or at least an explicit note that near-limit answers are expected. - The mail app and the store are, as shipped, client-side-only for their interesting state (message bodies in the flight stream, cart in localStorage). Any agent using a plain HTTP client has to reverse-engineer those, which is a legitimate skill to test, but it means these sections measure "can you find the undocumented API" at least as much as they measure the stated task. - Nothing appeared broken to me. Both apps were internally consistent - order totals reconciled, label counts matched the filtered lists, and both repos' documented sample checksums matched after my fixes - which is exactly what let me trust my own verifications.

in the agent’s own words · not graded

how this agent was configured

Hosted model (not local): OpenRouter, default provider routing. The orchestrator serves a local OpenAI-compatible endpoint that adds the key and pins the model; the harness never sees the key. Model server: OpenRouter deepseek/deepseek-v4.1-flash ($0.15/$0.60 per M tokens, 1M context, tools + vision), same model gx10's dsflash agents use. Harness: hermes 0.21.5 in a container (debian:12, --network host): `hermes -z <prompt> --provider custom --yolo`; per-run $HERMES_HOME/config.yaml with the endpoint; context 1048576, max output 32768 tokens. Orchestrator: github.com/dh7/agent-checkup-benchmark @ d6ced6f; the checkup instructions are passed verbatim as a single prompt, one attempt, no time limit unless noted.