airbench.ai

Benchmark v1.0 · report

hermes/rtx5090/qwen3.8-27b-nvfp4

sharedairbench.ai/checkup/ccd15d84-f927-47cc-ac41-fc14c2963ebe/report

setup

model type
open model (local)
hardware
RTX5090
harness
hermes
model
qwen3.8-27b-nvfp4
modelself-reportedqwen38-27b-nvfp4-128k

started 2026-09-26 09:47 UTC · shared 2026-09-27 16:48 UTC

overall

Answered 47 of 49 challenges; 47 correct.

47 of 49 challenges passed

partial run · 2 unanswered, counted against the score

  • 47 passed
  • 2 not answered

vitals

time

1h 50m

answered

96%

failed

0%

success

96%

systems

Math test

9/9 passed

time to last answer 1m 48s
  • letter-count-1✓ pass1m 40s

    prompt

    How many times does the letter "c" appear in "cficbasluc"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial string count; I verified it in code rather than by eye: cficbasluc has three c's.

  • decimal-compare-1✓ passbatched

    prompt

    Which decimal number is larger, 1.8 or 1.55? Answer with just the larger number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine. 1.8 > 1.55, no trap.

  • arithmetic-1✓ passbatched

    prompt

    Compute step by step, left to right (no operator precedence): 26 - 20 + 13 - 13 + 18. Answer with just the final number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Left-to-right: 26-20=6, +13=19, -13=6, +18=24. Verified with a script; trivial.

  • unit-convert-1✓ passbatched

    prompt

    Convert 2 km to m. Now treat that resulting number as a fresh quantity of km and convert it to m (1 km = 1000 m). Answer with just the final integer number of m.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Two-step trap: 2 km = 2000 m, then 2000 km = 2,000,000 m. I answered per the literal instructions; confident, though the challenge is a deliberate gotcha.

  • format-json-1✓ passbatched

    prompt

    Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "8791". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 8791. Example shape: {"answer":"1234","checksum":10}

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    checksum = 8+7+9+1 = 25, digit sum verified in code. I embedded the required JSON object as the answer string; the grader presumably parses it.

  • math-add-1✓ passbatched

    prompt

    What is 12 + 16? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial.

  • math-sum-1✓ passbatched

    prompt

    Compute the exact sum of 764 + 248. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    764+248=1012. Verified.

  • math-expression-1✓ passbatched

    prompt

    Evaluate this expression exactly and respond with just the integer: ((20 + -10) * (11 - 8)) + (-6 * -6) - 43

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    ((20-10)*3) + 36 - 43 = 23. Confirmed with a script; standard precedence.

  • math-determinant-1✓ passbatched

    prompt

    Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [-1, -8, 9, 8] [3, 1, 1, 9] [-6, 5, 5, 0] [-6, -4, -1, -2]

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I did not expand by hand; exact fraction-based Gaussian elimination in Python returned 5613.

Vision test

19/19 passed

time to last answer 57m 58s
  • acuity-20✓ pass2m 35s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I read the whole chart, transcribed every row, and zoomed to re-check row 4. Two independent reads agreed on WAVCT for row 4 group 3. Confident, though small text is exactly where a single character could flip.

  • acuity-14✓ pass31s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Transcribed the full chart and reported row 5 group 3 = CE5H8. Medium-small text; one read pass. I will re-verify with a zoom pass if time permits, but the transcription was consistent and the row read cleanly.

  • acuity-10✓ pass47s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Two independent vision passes both read row 6 group 3 as TZ3ZX, including a zoom pass that confirmed the middle character is the digit 3, not B or 8. Row 6 is very small type but both reads agreed.

  • acuity-8✓ pass4m 26s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 7 is the smallest line, so reads disagreed: a full-chart pass and a zoomed 3x crop both gave 25RX6, while one intermediate read gave 2S9X6. I cropped and enlarged the row (first crop missed because the rows occupy only the top ~70% of the image), and the zoom pass explicitly disambiguated 2nd char = 5 (flat top bar) and 3rd char = R (diagonal leg). Going with 25RX6.

  • count-simple✓ pass1m 04s

    prompt

    Look at the image at (fetch it and view it). How many green squares are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Vision model counted 4 green squares (4 distinct positions). I cross-checked with pixel analysis: 4 distinct green bboxes. Both agree on 4. The square outlines are hollow so my naive component count was noisy, but 4 distinct positions is consistent. Confident.

  • count-medium✓ pass2m 45s

    prompt

    Look at the image at (fetch it and view it). How many green triangles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Cross-checked with two independent methods: vision model said 13 green triangles (with a full inventory of 21 green shapes), and pixel-based connected-component + shape classification gave exactly 13 triangles (2 squares, 2 diamonds, 4 circles). The two inventories match shape-for-shape, so I'm confident in 13.

  • count-complex✓ pass1m 28s

    prompt

    Look at the image at (fetch it and view it). How many green squares are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Pixel-based connected-component analysis found exactly 31 green squares plus 4 green triangles and 3 green diamonds (38 total green shapes). Vision model independently confirmed 31 squares with the same sub-inventory. High confidence.

  • spatial-simple✓ pass1m 19s

    prompt

    Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Vision read the 5x5 grid and placed the red circle bottom-left (row 5, col 1). I cross-checked with pixel analysis: isolating TRUE red (G<80, RGB 220,38,38) from orange, the only red circle center sits in the bottom-left cell = row 5, col 1. Both agree.

  • spatial-medium✓ pass5m 37s

    prompt

    Look at the image at (fetch it and view it). Which shape does the arrow from the orange circle lead to? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Arrow tracing is inherently ambiguous. I read the full diagram, cropped+enlarged the region, and did a third confirmation pass specifically checking arrowhead direction. All three agreed: the single arrow with its tail at the orange circle (r3,c5) has its head at the green triangle (r2,c2). The other lines merely pass by the circle. Confident it's a triangle, not a diamond, per the confirmation pass.

  • spatial-complex✓ pass16m 27s

    prompt

    Look at the image at (fetch it and view it). Which shape is 3 steps after the green diamond along the arrows? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Arrow-tracing was the hardest vision task. I read the full diagram and a 1.4x enlarged crop; both independently gave green diamond -> orange circle -> blue triangle -> blue circle, so step 3 = blue circle. I also cross-checked step 1 (green diamond to orange circle) with pixel line-coverage. Vision timed out twice mid-task, which cost time but the two successful reads agreed.

  • chart-simple✓ pass40s

    prompt

    Look at the image at (fetch it and view it). What is the title shown at the top of this chart? Answer with just the chart title.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straightforward: the large bold title at top reads 'Server Incidents'. There's also a subtitle 'Reported incidents per month' but the title itself is 'Server Incidents'. Easy, high confidence.

  • chart-medium✓ pass1m 17s

    prompt

    Look at the image at (fetch it and view it). Using the "Support Tickets Opened" chart, approximately what value did Aug have? Read it off the y-axis; answers within +/-5 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Vision read Aug as ~34; I independently measured the bar pixels (baseline at y661, ~5.4 px/unit from gridline spacing) giving 34.6. Both methods agree, and +/-5 tolerance means 34 is safe. Confident.

  • chart-complex✓ pass2m 44s

    prompt

    Look at the image at (fetch it and view it). Using the "New Signups" chart, how many months did Mobile have a value greater than 58? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Mobile series strictly >58: Jan~85, Apr~72, May~73, Jul~65, Oct~70, Dec~80 = 6 months; Sep~52 falls just under the 58 cutoff. I confirmed by pixel-measuring bar tops against the 0/25/50/75/100 gridlines: Sep measured 52.4, comfortably below 58, the others clearly above. Both methods agree on 6. My first pixel attempt had a sign error; I recalibrated the 0-line and re-ran.

  • screenshot-simple✓ pass59s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Read the cart panel: 3 items, line totals 24.61 + 14.19 + 41.19 = 79.99. I verified the sum with a calculator, so the displayed $79.99 is internally consistent. Easy, high confidence.

  • screenshot-medium✓ pass1m 00s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Read 4 line items and the displayed Total of $218.70. I verified 33.30+86.46+66.68+32.26 = 218.70 with a calculator, so it's consistent. Displayed as $218.70 (trailing zero kept to match the on-screen formatting).

  • screenshot-complex✓ pass1m 11s

    prompt

    Look at the image at (fetch it and view it). What is the tax amount shown on this order summary? Answer with just the amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Read Tax = $34.61 from the summary block (Subtotal 432.08, Discount -47.53, Shipping 13.51, Tax 34.61, Total 432.67). I verified 432.08-47.53+13.51+34.61 = 432.67, so the tax figure is internally consistent. 11 line items was a bit of careful reading, but the number itself was clearly displayed.

  • diagram-simple✓ pass46s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Galena" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Simple top-to-bottom hierarchy: Badger -> Galena -> Guitar -> {Radish, Ocelot}. The arrow leaving Galena points straight down to Guitar. Unambiguous, single read, high confidence.

  • diagram-medium✓ pass2m 15s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Walrus"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    12-node DAG. Walrus has exactly one incoming arrow, and both a full-diagram read and a zoomed right-half read agree it comes from Garnet. The tricky part was that Garnet is a hub feeding many boxes (Nickel, Walrus, Toucan, Opal) and Vortex also runs lines along the top to Nickel; I specifically checked that no Vortex/Ocelot line lands on Walrus. Confident.

  • diagram-complex✓ pass10m 05s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Prism" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    This was the hardest vision task of the whole section. Prism's outgoing line crosses/benches past Juniper and a coarse read (middle-band crop) mis-traced it as landing on Juniper. But two 3x zooms independently confirmed the line passes to the LEFT of Juniper's edge and drops to Coyote, and that Juniper's two actual incoming arrows come from Ocelot and Iguana. Three of four reads agree on Coyote; I'm going with that.

Finding and reading email test

4/6 passed · 2 unanswered

time to last answer 1h 19m
  • aggregate-1✓ pass1h 18m

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages in the archive folder have attachments? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Counted archive messages with attachments. The 'attachments' label (20 messages) does NOT equal actual attachments - two 'Re: Western Wholesale' and other messages carry attachment chips without the label, so I crawled all 92 archive detail pages (across 4 paginated pages, using the per-page id= param so the detail pane actually opens) and counted real attachment chips: 22. I had to figure out the pagination + auto-open-page quirk, which was the fiddly part.

  • aggregate-2✓ pass31s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages are marked unread in the archive folder? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Unread messages in archive. Counted the '●' unread accent-dot across all 4 paginated archive pages (25+25+25+17 = 92 cards): 12+8+14+7 = 41. Straightforward once I had the pagination mapping worked out; the dot marker 'text-accent">●' is a reliable per-card unread indicator.

  • temporal-1✓ pass17s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message carrying the label "travel"? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Oldest message carrying the 'travel' label. Tricky: the label filter defaults to the INBOX view (5 msgs, oldest Apr 24 'DRAFT- TAP Power Outage'), which would mislead you. The 'Travel: 24' badge is global across folders, so I pulled travel-labeled messages from every folder view (inbox/archive/sent/drafts/trash = 24 total) and took the true oldest: archive 'Re: Denver trading' (Mar 19, 2001, 9:25 AM). That is older than the inbox one.

  • temporal-2✓ pass18s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message in the archive folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Newest message in the archive folder. Parsed the detail dates of all 92 archive messages; max is May 10, 2001 1:05 PM, subject 'RSVP REQUESTED - Emissions Strategy Meeting....' (the trailing '....' is part of the displayed subject, shown with no truncate class in the h1, so I reproduced it exactly).

  • needle-1— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the message with the subject "FW: Curve Shift File", what dollar amount is called out as the biggest difference between actual and theoretical curve shift? Answer with just the exact fact as it appears in the message, and nothing else.
  • needle-2— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the message with the subject "FW: Mid C New deals Sept 24", what net value to book is given for the new deals? Answer with just the number.

Purchasing test

4/4 passed

time to last answer 1h 38m
  • find-product-1✓ pass1h 27m

    prompt

    The store is at abostore.airbench.ai Among products in the **Automotive** category priced under **$510**, which has the **highest rating**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Automotive category, 29 products (25 on page 1 + 4 on page 2; I had to grab page 2 since the default Next state hid it). Of the 14 priced under $510, the highest-rated is the waterproof car hammock seat cover at 5.0 stars / $214.30, ASIN B00QHC041A. The card view didn't carry the domain prefix, so I opened the product page and read its 'Domain' field: primenow.amazon.com. Answer format <domain>:<ASIN>.

  • find-product-2✓ pass21s

    prompt

    The store is at abostore.airbench.ai Among products in the **Sports & Outdoors** category priced at or above **$950** with a rating of at least **4.2**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Sports & Outdoors = 854 products across 35 pages (25/page). Crawled all 35 pages, filtered to price >= $950 AND rating >= 4.2, then took the lowest price: $950.50, 4.5 stars, 'eono by amazon swimming changing robe towel grey' ASIN B088W8DVVT. Only one item sat at that minimum price, so no tie-break needed. Domain read from the product page 'Domain' field: amazon.co.uk.

  • purchase-1✓ pass9m 47s

    prompt

    The store is at abostore.airbench.ai Buy exactly 3 units of Franklin Tailored Men's Suede Glove, Black, XL (product id amazon.ca:B01618VW5G, abostore.airbench.ai/product/franklin-tailored-men-s-…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-7755ebf4@aidoctor.test. Answer with just the resulting order id.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result

    note

    Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    The checkout is a JS client flow, but I read the source: it POSTs JSON to /api/store/orders with {sessionId, cart:[{productId,slug,title,price,image,delivery,quantity}], customer, shipping, payment}. I extracted the glove's exact product object from the RSC payload of its page (price 612.70, delivery next-day), set quantity 3, used the required email and the prefilled valid test card 4242...4242. The API returned status 'approved' with recorded:true. Order id abs_4cc9740baa19.

  • recover-decline-1✓ pass1m 03s

    prompt

    The store is at abostore.airbench.ai Buy exactly 3 units of Girls' Petite Sterling Silver Birthstone Open Heart Stud Earrings and 16" Pendant Necklace Jewelry Set (product id amazon.ae:B01HGKTHGO, abostore.airbench.ai/product/girls-petite-sterling-si…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-889ac10a@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result>checkout_result

    note

    Store events recorded for this order's session: checkout_result → checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    Replayed the checkout API (/api/store/orders) with a fixed session across both attempts. First attempt used a card ending in 0000 -> API returned status 'declined'. I then re-submitted the same 3-unit cart with the prefilled valid test card 4242...4242, which returned 'approved'. The approved order id is abs_babdffadc55c. Note: I initially miscounted the quantity as 1 and produced a stray approved order, then redid it with the correct quantity of 3.

Coding test

11/11 passed

time to last answer 1h 50m
  • compute-hash-1✓ pass1h 40m

    prompt

    Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [815883800, 989164281, 1650633614, 1285474407, 421846868, 1211487333, 1784528618, 2648167027, 3745225680, 4159742481, 2972498822, 3263276735], x = 600553356, y = 2387716093 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straightforward 32-bit hash loop with rotl/imul exactly as specified; ran the 25000-iteration loop in Python with masking to 2^32. Routine; no ambiguity.

  • compute-vm-1✓ pass1m 05s

    prompt

    Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 448 1: set b 870 2: set c 225 3: set d 471 4: mul b 6 5: mul a 4 6: add a 10 7: dec d 8: jnz d -4 9: add a b 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Executed the tiny machine step-by-step in Python (faithful step loop, jnz relative jumps, mod 1000003 on add/sub/mul). The outer structure is two nested loops (d counts 471, c counts 225). Final a = 460462.

  • compute-paths-1✓ pass2m 56s

    prompt

    Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S#...###...#..#.....##... ...........#...#.....#..# ..#....#...#.#.....#.#... #...###.......#...#....#. #....#....###.##..###.### ..........#..##.#......#. ..#.#.......#...####..#.# .#....#....#....#.#..#... #..#.#.##...###..#....... #....#....##.#...#....... ...##..#.#...#..#..#..#.# .###.......#.#......###.. #....##.#...#.#..##...... ..##.....#......#........ #.#..#...#..#..#..#..##.. ..#...#..#..#..#.......#. #....#.#.#.##..........#. ..##.##....#.#...##.#...# ......#...#............## ..........#.....##....... ..#.....#.#.##...#..#.... #.##.#....##.#..#....##.# ##...#.#...##....#....... ##....#.........#.#.....# #....#..#...#......#.##.E Respond with the two integers separated by a space, like `52 1840`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    25x25 grid. BFS for shortest distance, then DP over nodes sorted by distance to count distinct shortest paths mod 1e9+7 (counting predecessors at distance-1). Confirmed S top-left, E bottom-right. Shortest path = 50 moves; 1418768 distinct shortest paths.

  • compute-life-1✓ passbatched

    prompt

    Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: ..#####.#...##.#..#. ...#...#..##....#... ...#####..#..##..#.. .#.#.#.#..#..#..#.#. ...##.##..#....##... ....###....#.#...#.# #...#.#.....##...##. ..#.####.#.....#.... ##...###....#....#.. ..##.#.....##...#..# ##.##.#..#.#..###... #.#......##.#.#..... ..#...#.#.#.....#... ..##.##.#..##.#.##.# ....##...##.......## #.###...#....###.##. .#.#.#.........#.... #.#....#......##.##. ...#.##...#..#.....# ##.....##.#.#.#.#.#. Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Conway's Life on a 20x20 torus, 150 generations, wrapping neighbours. Simulated in Python with a set of live cells. Ended with 15 live cells; sum of (row*20+col) over live cells = 1592. Straightforward; the main care was the wrap-around neighbourhood.

  • compute-fibmod-1✓ passbatched

    prompt

    Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 8986964264488779 and m = 15485863. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Fast-doubling Fibonacci modulo 15485863 for n=8986964264488779. Used the standard fast-doubling identity and verified F(10) mod 1e9 == 55 as a sanity check. Result 14271074.

  • compute-words-1✓ passbatched

    prompt

    Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. Kalu pellu PELREN quika quific Pelren "shamo" zanti quific tizan nixren Moka? renfic peltru; Quimo pellu peltru quific truren shaqui FICPEL Zanti rentru! quimo pelka quific quific tizan renpel lupel Peltru nixka. ficqui Quika RENREN Pellu pelren! moka lupel quific nixka Kamo. quika Quific moka. Zansha peltru rentru timo quific ficpel. Voren peltru quific quific Quika moka Zanti pelpel renren! Lupel ficpel quific! Zanti kalu renren renpel ficqui ZANTI pelka quimo zanti ficqui truren kalu kalu Quific Ficqui Renren "renren" nixqui Quific truren quika Pelpel ficqui quika; nixka renfic ficqui ficqui Lupel moka basbas quimo voren pellu shaqui quific quimo! voren voren Kamo Voren? moka Kamo peltru quika "tizan" Pelren Renren Tizan peltru Rentru lupel "quific" zansha Ficqui renpel timo Truren lupel quific! pelren Quific Peltru ZANTI Nixka basbas Nixka quimo Quific pelpel; pellu! pelren truren quific Ficqui renren Pelren renren basbas kamo "Pellu" renpel truren quific shaqui moka renpel Truren basbas Kalu Tizan; quific zanti Ficqui timo Voren renren pelren kamo ficqui quimo kamo pelka ficqui pelpel Kalu renpel "quific" "quimo" kalu Pellu? quimo Rentru! shamo pellu renren ficqui, shaqui quific "peltru" renren peltru peltru ZANSHA quika shamo Zanti kalu TIZAN, renpel Pelka BASBAS! quific Pelren Nixren zansha Truren Quimo. ficpel quific Kalu Kalu pellu kalu renren Kamo quific ficpel quimo Ficqui nixqui moka renpel ficqui kamo kalu ficpel Voren! FICPEL ficqui Ficqui shaqui ficqui renren nixka ficqui Quific lupel ficpel peltru moka MOKA quika tizan ficqui pelpel lupel QUIMO renpel renren kamo renren shamo lupel pelka truren "Quific" ficqui renpel timo, tizan! Moka "Ficpel" ficqui timo quific peltru ficqui quific PELREN tizan! ficpel quika kalu quific. Ficqui Lupel shaqui renren ficqui pelren quimo Truren kamo quimo ficqui! shamo quimo renren quific Zanti nixqui nixren "shaqui" lupel kalu QUIFIC shamo tizan kalu kamo lupel quimo, "quific" Quific pelka peltru, Renren zansha peltru pelka renpel quimo quific quific "quimo" lupel "Kalu" Basbas zanti nixqui zanti! tizan! nixren quific basbas peltru Quific shamo renpel Renren kalu Shamo pelren; Shamo pellu zansha Quific peltru; Pelpel "quika" Shaqui basbas ficqui pelka "pelpel" Renpel renren ficqui lupel basbas lupel; Voren Renren nixren renpel peltru truren quific lupel RENTRU nixqui Tizan. quific quimo Rentru Quific nixka renren renfic lupel truren kamo Quimo Lupel lupel zanti voren quika Kalu? nixka Truren ficpel lupel zansha "ficqui" quika "zanti" nixka quific ficqui Ficqui Lupel quimo Quific Renfic renpel Pelpel ficqui zanti RENREN Nixren shaqui nixka! voren, quific Lupel peltru tizan Quimo "Renren" kamo renpel quimo kalu ficqui lupel pellu quimo moka? quika renpel lupel Quika kamo nixren renpel,

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Tokenized the paragraph on spaces, stripped punctuation/quotes, lowercased, counted. Top 3 with ties broken alphabetically: quific=46, ficqui=33, lupel=24. Routine; I split off the instruction preamble from the actual word paragraph before counting.

  • trace-1✓ passbatched

    prompt

    What exactly does this JavaScript program print? Respond with just the printed output. const v1 = [typeof null, typeof NaN, typeof typeof 1].join("/"); const v2 = ["3", "32", "11"].map(parseInt).join(","); const v3 = ["70" < "8", null == 0, NaN === NaN].map(Number).join(""); const v4arr = [8, 7]; v4arr[8] = 1; const v4 = v4arr.length + ":" + v4arr.filter(() => true).length; console.log(v1, v2, v3, v4);

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Traced the JS by hand (node wasn't available to run it). typeof null='object', typeof NaN='number', typeof(typeof 1)='string'. parseInt map passes the index as radix: parseInt('3',0)=3, parseInt('32',1)=NaN, parseInt('11',2)=3 -> '3,NaN,3'. '70'<'8' is true (lexicographic), null==0 false, NaN===NaN false -> Number -> [1,0,0] -> '100'. Sparse array length 9, filter 3 -> '9:3'.

  • fix-1✓ pass1m 31s

    prompt

    Our shipping quote function below has one bug. Bug report: this order is quoted 158 cents, but the correct quote is 632: {"country":"DE","items":[{"grams":469,"qty":4,"price":1477,"fragile":false}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 464, 756, 1381, 1608]; // cents, by zone const PER_STEP = [0, 79, 120, 178, 267]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 4400, 8000, 15900, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value < FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"US","items":[{"grams":708,"qty":3,"price":2175,"fragile":false}]} {"country":"AU","items":[{"grams":1298,"qty":5,"price":6653,"fragile":false},{"grams":1204,"qty":5,"price":5464,"fragile":false},{"grams":719,"qty":5,"price":6752,"fragile":false}]} {"country":"ES","items":[{"grams":700,"qty":2,"price":2196,"fragile":false}]} {"country":"FR","items":[{"grams":629,"qty":2,"price":2917,"fragile":false}]} {"country":"GB","items":[{"grams":710,"qty":2,"price":321,"fragile":false}]} {"country":"JP","items":[{"grams":758,"qty":3,"price":8947,"fragile":true},{"grams":981,"qty":2,"price":3659,"fragile":false},{"grams":1785,"qty":1,"price":7939,"fragile":false},{"grams":1252,"qty":4,"price":7526,"fragile":true}]} {"country":"IT","items":[{"grams":853,"qty":1,"price":2557,"fragile":false},{"grams":237,"qty":1,"price":2498,"fragile":false},{"grams":950,"qty":5,"price":7205,"fragile":false},{"grams":697,"qty":1,"price":7542,"fragile":false}],"express":true,"coupon":"SHIP10"} {"country":"DE","items":[{"grams":1670,"qty":4,"price":3437,"fragile":false},{"grams":1187,"qty":1,"price":3053,"fragile":false},{"grams":1302,"qty":1,"price":8279,"fragile":false}]} {"country":"AU","items":[{"grams":244,"qty":3,"price":1077,"fragile":false}]} {"country":"US","items":[{"grams":771,"qty":2,"price":2469,"fragile":false}]} {"country":"CA","items":[{"grams":1786,"qty":5,"price":426,"fragile":false},{"grams":1352,"qty":5,"price":2894,"fragile":false},{"grams":1327,"qty":1,"price":2159,"fragile":false},{"grams":1085,"qty":5,"price":6348,"fragile":true}],"coupon":"SHIP10"} {"country":"FR","items":[{"grams":1061,"qty":4,"price":7652,"fragile":false},{"grams":517,"qty":2,"price":1749,"fragile":true},{"grams":1361,"qty":3,"price":5014,"fragile":false},{"grams":589,"qty":4,"price":5133,"fragile":false}]} {"country":"ES","items":[{"grams":949,"qty":1,"price":5019,"fragile":false},{"grams":277,"qty":2,"price":4813,"fragile":true}]} {"country":"FR","items":[{"grams":249,"qty":4,"price":5719,"fragile":false},{"grams":527,"qty":1,"price":571,"fragile":false},{"grams":1594,"qty":2,"price":2556,"fragile":false}]} {"country":"GB","items":[{"grams":86,"qty":1,"price":8808,"fragile":false},{"grams":1108,"qty":5,"price":7689,"fragile":true},{"grams":182,"qty":5,"price":3673,"fragile":true},{"grams":1302,"qty":3,"price":5760,"fragile":false}]} {"country":"GB","items":[{"grams":645,"qty":3,"price":1006,"fragile":false}]} {"country":"US","items":[{"grams":1744,"qty":5,"price":8719,"fragile":false},{"grams":503,"qty":1,"price":2408,"fragile":false},{"grams":1611,"qty":2,"price":8473,"fragile":false},{"grams":348,"qty":4,"price":2357,"fragile":false}]} {"country":"GB","items":[{"grams":1102,"qty":5,"price":4821,"fragile":true},{"grams":658,"qty":3,"price":7203,"fragile":false},{"grams":490,"qty":1,"price":3680,"fragile":false},{"grams":1081,"qty":4,"price":4023,"fragile":false}],"express":true} {"country":"ZA","items":[{"grams":228,"qty":1,"price":5478,"fragile":true},{"grams":886,"qty":5,"price":6897,"fragile":false},{"grams":256,"qty":3,"price":1191,"fragile":false}]} {"country":"IT","items":[{"grams":388,"qty":5,"price":3269,"fragile":true}]}

    answer

    answer hidden on shared reports

    grader · score 1

    20/20 outputs match

    agent's debrief

    Found the single bug: 'grams += item.grams' ignores the per-item quantity, so the weight-based step count is under-counted whenever an item has qty>1. Correcting it to 'grams += item.grams*item.qty' makes the bug-report DE order compute 79*ceil(469*4/250)=632 exactly as required (unfixed gives 158). I reproduced the whole quote() in Python (careful with JS '>>' on a negative-free value = floor/8 and Math.ceil semantics) and ran the fixed function on the 20 orders after the 'Orders:' marker in order. Left all other behaviour unchanged.

  • implement-1✓ passbatched

    prompt

    Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[23,24],[13,20],[13,15],[18,19],[17,21],[15,21],[36,39],[4,9]] [[35,42],[29,33],[22,27],[8,16],[34,41],[30,37],[10,13],[5,6]] [[1,9],[39,43],[36,39],[7,11],[17,21],[39,41],[39,42],[38,38]] [[27,30],[7,8],[36,37],[39,39],[15,21],[34,40]] [[5,13],[29,34],[0,6],[26,32],[33,40],[9,15],[21,21]] [[8,9],[38,38],[14,17],[8,15],[11,14],[20,20]] [[3,4],[1,7],[17,17],[0,2]] [[16,18],[24,24],[12,14]]

    answer

    answer hidden on shared reports

    grader · score 1

    12/12 lines match

    agent's debrief

    Wrote mergeIntervals (sort by start, merge overlapping/touching where next.start <= last.end). Ran on all 12 inputs; e.g. [1,3]+[3,5]->[1,5] but [1,2],[3,4] stay apart. Each input printed as a JSON list of pairs on its own line.

  • repo-1✓ pass2m 14s

    prompt

    Download airbench.ai/f/dc8204ffed132a83cc515c27ae74382d.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Downloaded and unzipped the ledger project. Ran the tests: 1 of 7 failed, test_threshold_counts_as_large (withdrawal_fee(50000) returned 25 but the docstring + test both say 50,000 'or more' pays 1%). The bug was a strict > in fees.py instead of >=. Changed the one comparison; all 7 tests pass; main.py on real.csv now prints 28fdd626 (was 39af7e89 pre-fix).

  • repo-2✓ pass2m 44s

    prompt

    Download airbench.ai/f/7af4dc7a71b879e0bb79c04e1d716fe6.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Two injected bugs, both caught by the bundled tests. (1) parse.py sorted by (date, amount) instead of date only — the amount tie-break reorders same-date transactions and changes overdraft sequencing; reverted to date-only (stable) sort. (2) report.py charged the overdraft fee on bal <= 0, so a withdrawal landing exactly on zero was (wrongly) overdrawn; the test test_withdrawal_to_exactly_zero_is_not_an_overdraft wants < 0. Fixed both, all 7 tests pass, main.py on real.csv prints b499b7b3.

how this agent was configured

Hardware: NVIDIA RTX 5090 32 GB (SM120), x86_64, driver 595.84, Ubuntu 24.04 (host hal5090). Model server: Checkpoint gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090 (NVFP4, no MTP head). vLLM 0.27.1 (vllm/vllm-openai:v0.27.1): --quantization modelopt --kv-cache-dtype fp8 --trust-remote-code --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_xml --max-model-len 131072 --max-num-seqs 4 --gpu-memory-utilization 0.95. ~79 tok/s single-stream decode. Harness: hermes 0.21.5 in a container (debian:12, --network host): `hermes -z <prompt> --provider custom --yolo`; per-run $HERMES_HOME/config.yaml with the endpoint; context 131072, max output 16384 tokens. Orchestrator: github.com/dh7/agent-checkup-benchmark @ 6d738a5; the checkup instructions are passed verbatim as a single prompt, one attempt, no time limit unless noted. Operator limits: stopped by the operator at 126 min (120-min cap).