airbench.ai

Benchmark v1.0 · report

hermes/rtx5090/qwen3.8-27b-quasar-nvfp4-mtp

setup

model type
open model (local)
hardware
RTX5090
harness
hermes
model
qwen3.8-27b-quasar-nvfp4-mtp
modelself-reportedqwen38-27b-quasar (Hermes Agent)

started 2026-10-08 06:47 UTC

overall

Answered 49 of 49 challenges; 43 correct.

43 of 49 challenges passed

  • 43 passed
  • 6 failed

vitals

time

1h 49m

answered

100%

failed

12%

success

88%

systems

Math test

9/9 passed

time to last answer 1m 50s
  • letter-count-1✓ pass1m 10s

    prompt

    How many times does the letter "l" appear in "plellluzan"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Counting the 'l's in plellluzan: p-l-e-l-l-l-u-z-a-n, that's 4. Trivially easy, just a careful manual count.

  • decimal-compare-1✓ pass4s

    prompt

    Which decimal number is larger, 7.59 or 7.4? Answer with just the larger number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    7.59 vs 7.4: same integer part, 0.59 vs 0.40, so 7.59 is larger. Easy, no ambiguity — I just had to be careful not to be tempted to read 7.4 as 7.40.

  • arithmetic-1✓ pass5s

    prompt

    Compute step by step, left to right (no operator precedence): 15 * 3 * 6 + 18 + 15. Answer with just the final number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Left-to-right, no precedence: 15*3=45, *6=270, +18=288, +15=303. I verified with a script since doing it by hand is error-prone. The trap here is operator precedence — with standard precedence it would be 303 too... actually 15*3*6=270, +18=288, +15=303. Both give 303, so no trap triggered.

  • unit-convert-1✓ pass5s

    prompt

    Convert 20 kg to g. Now treat that resulting number as a fresh quantity of hours and convert it to minutes (1 hours = 60 minutes). Answer with just the final integer number of minutes.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    20 kg = 20000 g; treat 20000 as hours: 20000 * 60 = 1200000 minutes. Straight two-step conversion, done with a script to avoid arithmetic slips.

  • format-json-1✓ pass5s

    prompt

    Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "1670". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 1670. Example shape: {"answer":"1234","checksum":10}

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Digit sum of 1670 is 1+6+7+0=14. Kept the exact key order and made checksum a JSON number, not a string, as instructed. Easy, just attention to format.

  • math-add-1✓ pass9s

    prompt

    What is 8 + 4? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    8+4=12. Trivial.

  • math-sum-1✓ pass4s

    prompt

    Compute the exact sum of 187 + 708. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    187 + 708 = 895. Straight addition, verified by script.

  • math-expression-1✓ pass5s

    prompt

    Evaluate this expression exactly and respond with just the integer: ((14 + 16) * (12 - 31)) + (3 * -2) - 13

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    30 * -19 = -570, plus -6 = -576, minus 13 = -589. Straight evaluation, computed by script to avoid sign slips with the negative operands.

  • math-determinant-1✓ pass4s

    prompt

    Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [5, 5, -9, -4] [-4, 1, 9, 8] [-7, -4, 10, 8] [7, 6, 2, 3]

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Computed the 4x4 determinant two independent ways (cofactor expansion and fraction-based Gaussian elimination) and both agree at -700. This was the heaviest math item; doing it by hand would have been error-prone, so scripting was the right call.

Vision test

15/19 passed

time to last answer 51m 14s
  • acuity-20✓ pass3m 53s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Read the chart both at full scale and a 5x-zoomed crop of row 4 group 3; both agree on X69NX. The zoom removed any 6/9 or N/M ambiguity. Confident.

  • acuity-14✓ pass1m 22s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Verified row 5 group 3 two ways: 5x-zoomed crop and the full chart, both read D8GEA. The 8 (vs B) was the main confusable and the zoom made it unambiguous.

  • acuity-10✓ pass46s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 6 is one of the smaller rows. Zoomed crop read DPH4C and the full-chart read independently confirmed the same. The 4 (vs A) was the only confusable and both reads agreed. Confident.

  • acuity-8✓ pass1m 03s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    This was the hardest acuity item — row 7 is only ~6px tall. An 8x crop read Z32RH and a full-chart read independently gave the same. Z/2/3 are the usual confusables here; the flat-top of the Z vs the curved head of the 2 was the key tell and both reads agreed.

  • count-simple✓ pass1m 10s

    prompt

    Look at the image at (fetch it and view it). How many red diamonds are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Vision read 6 red diamonds with a band-by-band layout. I then verified programmatically by flood-filling: 9 reddish blobs total, but 3 were orange squares (RGB 242,106,34, filling their bbox) and 6 were true red diamonds (RGB 220,38,38, ~50% area ratio). The 3 orange ones are the trap; red diamonds = 6. Both methods agree.

  • count-medium✓ pass1m 45s

    prompt

    Look at the image at (fetch it and view it). How many blue circles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Vision read 15 blue circles row-by-row. Programmatic flood-fill initially said 16 'blue' circles, but inspecting center RGB revealed one of those was a TEAL circle (8,145,178), not blue — my loose b>150 filter overcounted. Filtering to true royal blue (36,99,235) leaves exactly 15, matching the vision count. The teal and the blue diamonds/triangles were the traps.

  • count-complex✓ pass1m 09s

    prompt

    Look at the image at (fetch it and view it). How many blue squares are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The hardest counting item — lots of near-miss shapes (teal/green/purple/orange squares, blue triangles, blue circle, blue diamonds). Vision listed 21 blue squares with exclusions. I verified by flood-fill with a tight royal-blue filter and classified each blob by fill ratio: 21 at 0.991 (squares), 2 at 0.524 (diamonds), 1 at 0.758 (circle). Both methods independently gave 21.

  • spatial-simple✓ pass49s

    prompt

    Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    5x5 grid. Vision said row 3, col 4. Verified programmatically: flood-filled for red, found 11 reddish blobs but only one is true red RGB(220,38,38) — the rest are orange (242,106,34). That red circle sits at pixel (852,617); grid cell centers are 147/382/617/852/1087, so col 4, row 3. Both agree.

  • spatial-medium✓ pass1m 39s

    prompt

    Look at the image at (fetch it and view it). Which shape has an arrow pointing to the orange square? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The orange square has two arrows touching it: one leaving it (to the green square) and one entering it. The entering arrow's head is at the orange square and its tail runs back to the teal triangle. Verified in the full image and a 2x zoom of the bottom-left. Direction was the key — I had to make sure I read the arrowhead end, not the tail.

  • spatial-complex✕ fail2m 42s

    prompt

    Look at the image at (fetch it and view it). How many shapes come after the blue circle along the arrows? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 0

    wrong count

    agent's debrief

    Followed arrows forward from the true-blue circle (bottom-right): blue circle -> red triangle -> orange diamond -> red square -> green circle -> blue square, then stops. That's 5 shapes after. The two traps were (a) distinguishing the blue circle from teal circles, and (b) distinguishing the blue circle's OUTGOING arrow from its incoming one (from the red diamond). Zoomed into the bottom-right to confirm the outgoing arrowhead lands on the red triangle.

  • chart-simple✓ pass31s

    prompt

    Look at the image at (fetch it and view it). What is the title shown at the top of this chart? Answer with just the chart title.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The bold title at the top reads 'Server Incidents'; the smaller gray line below it is a subtitle ('Reported incidents per month'), not the title. Straightforward read, no ambiguity.

  • chart-medium✓ pass2m 07s

    prompt

    Look at the image at (fetch it and view it). Using the "Website Sessions" chart, approximately what value did Jan have? Read it off the y-axis; answers within +/-5 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Read Jan off the y-axis. Cross-checked by measuring pixels: the 20-unit gridlines sit at y=119/227/335/443/551 (108px per 20 units), Jan bar top at y=201 -> ~84.8. So Jan ≈ 85 (thousands).

  • chart-complex✓ pass1m 20s

    prompt

    Look at the image at (fetch it and view it). Using the "Website Sessions" chart, how many months did Americas have a value greater than 83? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Identified Americas as the orange bars, read each month, and verified by pixel measurement (0 at y=679, 140px per 25 units). Americas values: 56,88,66,58,23,36,48,35,34,15,43,23. Only February (~88) is >83. Count = 1.

  • screenshot-simple✕ fail32s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 0

    expected $73.97, got "3.97"

    agent's debrief

    Cart panel totals: Desk Lamp x3 = 51.87, Keyboard 12.97, Phone Stand 9.13; grand total shown as 3.97. I verified the line totals sum correctly (51.87+12.97+9.13=73.97). Clean read, no shipping/tax complications.

  • screenshot-medium✕ fail35s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 0

    expected $224.33, got "24.33"

    agent's debrief

    Cart shows Laptop Sleeve 91.72, Coffee Mug 62.12, Headphones 37.81, Phone Stand 32.68; Total 24.33. Verified all line totals and the grand total sum exactly. Easy read.

  • screenshot-complex✕ fail1m 19s

    prompt

    Look at the image at (fetch it and view it). What is the discount amount shown on this order summary? Answer with just the amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 0

    expected $98.44, got "8.44"

    agent's debrief

    Order summary showed Discount -8.44 (the minus just marks it as a reduction). Subtotal 656.26, shipping 11.06, tax 33.47, total 602.35 — I recomputed all the arithmetic and it reconciles exactly. The answer is the discount amount, 8.44.

  • diagram-simple✓ pass53s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Pumice"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Diagram: Nebula -> Banjo -> Pumice, then Pumice branches to Laurel and Alder. Only Banjo has an arrow pointing TO Pumice. Zoomed 2x to confirm arrowhead direction. The trap is that Pumice also has arrows, but those go OUT of it.

  • diagram-medium✓ pass2m 29s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Spruce" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Spruce has one outgoing arrow (its top edges are incoming). That arrow runs down to Falcon with the arrowhead at the Falcon end. Verified in full diagram and a 2x zoom of the Spruce/Falcon region. Answer: Falcon.

  • diagram-complex✓ pass25m 11s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Marmot"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The single arrow into Marmot comes from Toucan. Verified three ways: pixel scan found exactly one filled arrowhead on Marmot's top edge (none on left/right), its shaft traces up to the Toucan box (x414-510), and two zoomed vision reads both named Toucan (one specifically noted Lemur has no bottom connector, ruling it out). An earlier full-image read wavered between Toucan and Toucan+Lemur, so I leaned on pixel+zoom over the single ambiguous full read. Fairly confident, though this was the hardest vision challenge of the set.

Finding and reading email test

4/6 passed

time to last answer 1h 16m
  • aggregate-1✓ pass1h 00m

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages are marked unread in the archive folder? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Scraped all 92 archive messages across 4 server-side pages and counted rows with the unread accent dot in the subject line: 41 unread. Cross-checked the dot detection against the row structure. The app is a Next.js SPA; I parsed the server-rendered HTML list rows.

  • aggregate-2✕ fail51s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages carry the label "attachments"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 0

    agent's debrief

    The 'attachments' label view lists exactly 5 messages (the label filter is server-side, so 5 is the full count, not a page of a larger set): Service Agreement (Sep 11), two 'Save the Date' (May 11), 'Non-Critical' Maintenance (Apr 26), LDC Forum-Atlanta (Apr 25).

  • temporal-1✓ passbatched

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the archive folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Oldest message in archive. 7 messages share Mar 15; the server's sort=oldest ordering is deterministic and stable (verified across 2 fetches), placing 'RE: PERSONAL AND CONFIDENTIAL COMPENSATION INFORMATION' first. I could not independently confirm the sub-day time-of-day (the app renders timestamps client-side and full datetimes aren't in the SSR HTML), so I trust the server's sort order.

  • temporal-2✕ failbatched

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message carrying the label "attachments"? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 0

    agent's debrief

    Newest message carrying the 'attachments' label. The 5 attachment-labeled messages span Apr 25 to Sep 11; Sep 11 is unambiguously the most recent, and it is 'Service Agreement' from Hunter Williams (Grande Communications). No tie.

  • needle-1✓ pass1m 36s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Renee's reply to Phillip about the Deferred Phantom Stock Units payout, how many shares does she say are reflected on the 9/30/01 statement? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Renee Ratcliff's reply (Nov 2) to Phillip about the Deferred Phantom Stock Units payout: 'We will be distributing the shares reflected on your 9/30/01 statement (6,606 shares plus cash for fractional shares).' So 6,606. I had to read the full message body (client-rendered detail page) — the number only appears in the body, not the subject or snippet.

  • needle-2✓ pass13m 10s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Phillip's message to Steve Matthews about building a muni bond ladder from his account, what total account value does he give? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Phillip to Steven Matthews (Nov 13, 2001, no subject): 'My account has a value of around ,400,000. That includes 750,000 of us treasury notes. I am ready to build a bond ladder of muni's.' This was the hard needle: the search box doesn't index recipients or bodies, and detail pages render bodies client-side (RSC), so I had to enumerate all 178 messages via paginated list views, grep the flight payload for 'ladder', and then fetch the right page with the right id to force the body to render. The account value is 1,400,000 (text says 'around ,400,000'); other figures in the thread (785,000 money market, 750,000 treasury) are quoted from Steve's reply, not stated by Phillip.

Purchasing test

4/4 passed

time to last answer 1h 34m
  • find-product-1✓ pass1h 24m

    prompt

    The store is at abostore.airbench.ai Among products in the **Tools & Hardware** category priced under **$400** with a rating of at least **4**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Scraped all 973 Tools & Hardware products across 40 category-filtered pages (25/page). Filtered to price<400 and rating>=4 -> 259 candidates. Cheapest is the AmazonBasics 8-inch angled-head diagonal cutters at .46, rating 4.8. Verified on its product page that the canonical id is amazon.com:B07TWFDN96 (price .46 confirmed). No product sits exactly at the 00 or rating-4.0 boundary, so no tie risk.

  • find-product-2✓ passbatched

    prompt

    The store is at abostore.airbench.ai Among products in the **Pet Supplies** category priced at or above **$200** with a rating of at least **4.5**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Scraped all 314 Pet Supplies products. Filtered to price>=200 and rating>=4.5 -> 100 candidates. Cheapest is the AmazonBasics padded pet bed with rim, 102 cm at 10.26, rating 4.5. Verified on its product page that the canonical id is amazon.de:B07FVMLVM2 (price 10.26 confirmed). Second-cheapest is 18.06, so clear margin.

  • purchase-1✓ pass9m 04s

    prompt

    The store is at abostore.airbench.ai Buy exactly 1 unit of AmazonBasics Assorted Size and Color Rubber Bands, 0.5 lb. (product id amazon.ca:B074B1KCXD, abostore.airbench.ai/product/amazonbasics-assorted-si…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-a793832c@aidoctor.test. Answer with just the resulting order id.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result

    note

    Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    Completed checkout via the store's order API (POST /api/store/orders, which is what the checkout page's React handler calls). Built the cart item from the product page's canonical object (productId amazon.ca:B074B1KCXD, price 124.40, delivery standard), 1 unit, used the required email aidoctor-a793832c@aidoctor.test, and a valid test card 4242...4242. Server returned status=approved, total 143.61 (124.40 + 8.95 ground + 10.26 tax). Order id abs_b195a28cb8f4.

  • recover-decline-1✓ pass1m 17s

    prompt

    The store is at abostore.airbench.ai Buy exactly 2 units of Simply Safe Poly Food Service Disposable Gloves, Translucent, 500 per box/20 Boxes per case, Medium (product id amazon.ca:B079KKB6RV, abostore.airbench.ai/product/simply-safe-poly-food-se…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-96d97389@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result

    note

    Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    2 units of the Simply Safe gloves (amazon.ca:B079KKB6RV). First attempt with a card ending 0000 (4111111111110000) -> declined (orderId abs_d798f4f4e1de). Retry with valid card 4242...4242 -> approved, orderId abs_045c40295a27 (total 808.87: 738.96 + 8.95 + 60.96). Same email aidoctor-96d97389@aidoctor.test both times.

Coding test

11/11 passed

time to last answer 1h 49m
  • compute-hash-1✓ pass1h 47m

    prompt

    Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [1219108889, 1120551726, 1082387591, 1003952628, 20494213, 2456454282, 3874716819, 1594594416, 3556760369, 419122982, 3078682335, 4065398316], x = 76633885, y = 4157857538 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Executed the 25000-round 32-bit hash exactly per spec; straightforward to implement and run.

  • compute-vm-1✓ passbatched

    prompt

    Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 584 1: set b 372 2: set c 395 3: set d 430 4: mul a 60 5: mul a 41 6: sub b a 7: dec d 8: jnz d -4 9: add a b 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Emulated the VM; d-loop multiplies a by 60*41 over 430 iters then subtracts into b, c-loop adds b 395 times; simulated mod 1000003.

  • compute-paths-1✓ passbatched

    prompt

    Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S...#.#.#........##....#. #..#.#.#..##.###.......#. .......#....#.###.....#.. ..#.....##......#...###.# ..#......#....#....#.#... ....###.#####.#..#.#..... ..#.#.#.#..#..##.......#. #.#..#.#.#...##.#.#.##.## .#..#.........#..##....#. ......##....#..#.#......# .....#.....#.##....#.#..# .......#....###....#####. .#...#.#.#..#....#....##. .##.##...###....#..##..#. ..#......#.#....#.#..#... ..#..........#..###.#..#. ...#.......#.#..#..##.... #...####.#....#.##.##.... ..#.#...#...##..###....#. .##.#.#.....##.....#...#. ......###..#..###.###.... ............#..#.#..#..#. .........#.##...#.#...#.. ##........##.#.##..#..#.. #.#...##.........#.....#E Respond with the two integers separated by a space, like `52 1840`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    BFS shortest path = 56 moves on the 25x25 grid; DP count of distinct shortest paths = 8640 (mod 1e9+7).

  • compute-life-1✓ passbatched

    prompt

    Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: ..#..#..#..##...#.## .#....##...#.##...#. ......#.##......#.#. ....###..#..#..#...# #..#...##..#.....### ......#...#.....###. ###..##.#..#...##..# ###..#.#####.#...#.. ....#.##..#.###.#... .#.#..#...#.#..#...# ##...........##..... #..#.#..##..#....#.. #....#.#..##.....#.. ..####...##..##.##.. ...#.#.....##.##.##. .#.#........##..##.# ....##..#.##.....#.# .#...#..######.#..#. ......#####.....##.. .#....#...#.#..####. Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Simulated 150 gens of toroidal Life; settled at 51 live cells, sum(row*20+col)=10587.

  • compute-fibmod-1✓ passbatched

    prompt

    Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 570257441606372 and m = 999983. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Fast-doubling Fibonacci F(570257441606372) mod 999983 = 313279.

  • compute-words-1✓ passbatched

    prompt

    Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. nixti. pelbas quibas, shalu vopel tinix tinix? Voren Zanti kaka dormo Volu trutru nixti karen Nixti trutru nixnix zanti pelka volu mofic nixti zansha Mofic baszan Nixbas voren karen karen motru nixnix shalu trumo truti motru! nixti zanka; mofic vopel dormo? Lufic, mofic zansha kaka zanka motru tinix! kaka pelka tinix Mofic dormo tinix, Tinix? PELKA renzan kaka dormo baszan karen dormo vopel tinix tinix baszan pelbas mofic kaka kaka Baszan motru karen renka zanren baszan tinix Baszan kaka? tinix "mofic" quibas MOFIC, mofic Motru nixnix kaka Nixbas nixnix zanren Renzan "renzan" zantru zansha Dormo. vopel nixbas kaka "mofic" Voren zanti shalu volu lulu trumo SHALU QUIBAS pelka, mofic Zantru mofic? zansha LUFIC? dormo renka kaka Karen volu zanti Pelka karen mofic baszan. zantru karen kaka kaka renzan Karen zansha vopel, "tinix" motru kaka tinix mofic Kaka renka mofic renka Zantru nixbas kaka "Kaka" mofic "Baszan" tinix! Karen Mofic renka mofic kaka. pelbas quibas kaka baszan dormo; pelbas trumo. pelbas Mofic karen nixnix Quibas karen pelbas kaka truti Zanka mofic voren SHALU vopel Pelbas nixbas karen Zanka karen Mofic karen tinix, karen baszan Zanti? pelbas. Volu! mofic motru kaka trutru lulu mofic karen. renzan ZANTI zanka quibas mofic, motru lufic "renzan" nixbas kaka zanren shalu Zanti Renka pelka quibas, tinix zantru zanren, Trutru lulu trumo zansha Zantru truti Nixnix, Bastru zantru mofic SHALU trumo baszan mofic Zanren kaka karen baszan tinix lulu; motru. motru zanren zanka TINIX PELBAS bastru nixnix kaka nixti pelbas? karen zanren DORMO Shalu Zanren; karen lufic zanka trutru zansha karen Zanren nixti. zantru nixbas MOFIC tinix Motru Tinix truti baszan trutru truti baszan Zansha quibas mofic karen karen Zansha zansha zantru zanti tinix vopel Zantru lufic trumo motru mofic Kaka karen Tinix zantru zanti zantru Lufic mofic? zantru pelka renka Lufic lulu Volu? lufic bastru! mofic mofic? Shalu lulu pelbas renzan Zanka nixti Mofic shalu shalu karen MOFIC nixti "nixti" "Pelka" mofic pelbas! pelka zanka Volu trutru Tinix Baszan mofic kaka Volu pelka Nixti "kaka" "karen" motru MOFIC bastru Baszan shalu karen renka pelbas mofic vopel Quibas zanka karen vopel. nixnix trumo. kaka Quibas Trutru KAREN! shalu zanti Dormo mofic; Karen mofic motru. zanren karen tinix trutru mofic lulu Pelka lulu tinix trutru Baszan? kaka zansha zantru dormo mofic mofic lufic zanren ZANTI zanka bastru nixti renzan Shalu; baszan kaka Mofic KAREN vopel karen renka Karen mofic VOLU mofic karen! "karen" karen Mofic renka NIXNIX mofic motru renzan renzan quibas Quibas zanka renzan karen trumo Vopel dormo Zansha motru lufic tinix Renzan Kaka, Truti pelbas "pelbas"

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Whitespace split, stripped attached punctuation, lowercased; top-3 = mofic=48, karen=38, kaka=30.

  • trace-1✓ passbatched

    prompt

    What exactly does this JavaScript program print? Respond with just the printed output. const v1 = "2" + 9 - 7 + "7"; const v2 = ["80" < "9", null >= 0, null == 0].map(Number).join(""); const v3 = [70 / 4 | 0, Math.round(-3.5), -78 % 2].join(","); const v4arr = [8, 5]; v4arr[5] = 9; const v4 = v4arr.length + ":" + v4arr.filter(() => true).length; console.log(v1, v2, v3, v4);

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    JS coercion: '2'+9-7+'7'->227; [80<9,null>=0,null==0]->110; 70/4|0=17, round(-3.5)=-3, -78%2=0; sparse arr length 6 / filter 3 -> 6:3.

  • fix-1✓ passbatched

    prompt

    Our shipping quote function below has one bug. Bug report: this order is quoted 311 cents, but the correct quote is 466: {"country":"IT","items":[{"grams":174,"qty":2,"price":2928,"fragile":true}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 501, 742, 1185, 1797]; // cents, by zone const PER_STEP = [0, 78, 110, 174, 248]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 5600, 10100, 18100, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams * item.qty; value += item.price * item.qty; if (item.fragile) fragile += 1; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value < FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"DE","items":[{"grams":189,"qty":5,"price":7673,"fragile":false},{"grams":1637,"qty":2,"price":6638,"fragile":false},{"grams":1727,"qty":3,"price":7559,"fragile":false}]} {"country":"US","items":[{"grams":902,"qty":1,"price":6773,"fragile":true},{"grams":130,"qty":5,"price":2670,"fragile":false},{"grams":618,"qty":1,"price":8465,"fragile":false}],"express":true} {"country":"CA","items":[{"grams":449,"qty":3,"price":6188,"fragile":false},{"grams":1230,"qty":3,"price":4862,"fragile":true},{"grams":1624,"qty":1,"price":1473,"fragile":true}]} {"country":"GB","items":[{"grams":415,"qty":1,"price":7593,"fragile":false},{"grams":1712,"qty":3,"price":7922,"fragile":false},{"grams":1130,"qty":1,"price":7295,"fragile":false}],"express":true} {"country":"JP","items":[{"grams":1200,"qty":1,"price":5583,"fragile":true},{"grams":971,"qty":1,"price":2812,"fragile":false},{"grams":1363,"qty":5,"price":7731,"fragile":false},{"grams":216,"qty":1,"price":5222,"fragile":false}],"express":true} {"country":"MX","items":[{"grams":1000,"qty":1,"price":8674,"fragile":false},{"grams":1220,"qty":5,"price":2471,"fragile":false},{"grams":1137,"qty":1,"price":4663,"fragile":false},{"grams":1117,"qty":5,"price":3198,"fragile":false}]} {"country":"DE","items":[{"grams":422,"qty":3,"price":1731,"fragile":true}]} {"country":"CA","items":[{"grams":1088,"qty":1,"price":2374,"fragile":true},{"grams":892,"qty":5,"price":8516,"fragile":true}]} {"country":"ES","items":[{"grams":1713,"qty":5,"price":8247,"fragile":false}]} {"country":"ES","items":[{"grams":513,"qty":3,"price":2629,"fragile":true}]} {"country":"BR","items":[{"grams":1630,"qty":2,"price":953,"fragile":true}],"coupon":"SHIP10"} {"country":"DE","items":[{"grams":206,"qty":5,"price":1502,"fragile":false},{"grams":1090,"qty":1,"price":7922,"fragile":true}]} {"country":"CA","items":[{"grams":494,"qty":2,"price":2429,"fragile":true}]} {"country":"IT","items":[{"grams":600,"qty":2,"price":1737,"fragile":true}]} {"country":"AU","items":[{"grams":1793,"qty":2,"price":6890,"fragile":false},{"grams":1724,"qty":1,"price":8960,"fragile":true},{"grams":452,"qty":1,"price":7572,"fragile":false}],"coupon":"SHIP10"} {"country":"JP","items":[{"grams":495,"qty":3,"price":619,"fragile":true}]} {"country":"NZ","items":[{"grams":190,"qty":4,"price":6487,"fragile":false},{"grams":1479,"qty":3,"price":8952,"fragile":false}]} {"country":"AU","items":[{"grams":1206,"qty":5,"price":3665,"fragile":false},{"grams":227,"qty":1,"price":1260,"fragile":false},{"grams":825,"qty":1,"price":3604,"fragile":false},{"grams":1768,"qty":1,"price":8742,"fragile":false}]} {"country":"DE","items":[{"grams":170,"qty":3,"price":2894,"fragile":true}]} {"country":"JP","items":[{"grams":113,"qty":3,"price":1736,"fragile":true}]}

    answer

    answer hidden on shared reports

    grader · score 1

    20/20 outputs match

    agent's debrief

    Bug: fragile += 1 should be += item.qty. Verified fix reproduces bug-report (buggy 311 / fixed 466), then applied fixed quote() to all 20 orders.

  • implement-1✓ passbatched

    prompt

    Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[21,22],[17,22],[21,27],[33,37],[13,14],[7,9]] [[34,39],[2,7],[19,27],[25,26],[21,26],[14,21],[8,8],[32,32]] [[0,5],[6,13],[37,38],[31,31],[26,34],[33,34]] [[38,42],[8,12],[32,37],[36,39],[31,38],[11,17],[18,23],[9,12]] [[27,35],[0,1],[7,12],[8,12],[20,25]] [[3,10],[21,25],[13,19],[39,40],[0,8],[17,19]] [[15,22],[17,24],[10,16]] [[7,12],[20,28],[29,34],[23,25],[38,43],[34,42],[37,37],[17,19]]

    answer

    answer hidden on shared reports

    grader · score 1

    12/12 lines match

    agent's debrief

    Wrote mergeIntervals (sort, merge touching/overlapping); ran on all 12 inputs, one line per input as JSON.

  • repo-1✓ pass2m 01s

    prompt

    Download airbench.ai/f/0fb9fb8845723229b985367dc5312695.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The bug was the withdrawal fee using -(-cents//100) (rounds up) instead of cents//100 (round down). Fixed it; verified sample.csv prints the README-expected ca48548c and all 7 unittests pass, then ran on real.csv.

  • repo-2✓ passbatched

    prompt

    Download airbench.ai/f/28c7974423d84ae99640155296e9cb44.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Two bugs: fees.py used cents > LARGE_WITHDRAWAL (excluded the exact 50,000 threshold) instead of >=, and parse.py sorted by (date, amount) which reorders same-date rows, breaking the required file order. Fixed both; sample.csv prints the README-expected 237188d9, unittests pass, then real.csv.

what the agent said about this checkup

Candid account of the airbench checkup (49 challenges across 5 sections, all submitted, none late). What was easy/routine: - Math section: straightforward. Arithmetic, unit conversion, decimal comparison, JSON formatting, and a 4x4 determinant all routine. I verified the determinant two independent ways (cofactor expansion and Fraction-based Gaussian elimination) and both gave -700. - Vision counting/spatial: pixel analysis with Pillow plus vision zooms was very reliable; I cross-checked every reading with a second method (a pixel measurement or a zoomed crop). The totals (6 red diamonds, 15 blue circles, 21 blue squares, reachable-chain=5) felt solid. - Email aggregates/temporal: server-side sort and the Next.js flight payloads were authoritative once I stopped fighting the client-rendered detail pages. - Coding "compute" set: I just wrote and ran programs (hash, tiny VM, Life, BFS path counting, fast-doubling Fibonacci, word counts, interval merging, a JS trace, and the shipping-fee fix). These were all mechanical once specified. The shipping-fee bug I verified against the provided bug report first (buggy 311 / fixed 466) before applying it. - The two repo challenges: the bugs were small and the READMEs gave me the exact expected sample checksum (ca48548c, 237188d9), so I could confirm each fix before running on real.csv. What was hard: - Vision acuity charts: the smallest row (~6px glyphs) needed 5x-8x zooming, and one diagram (diagram-complex) genuinely conflicted across reads - a full 2x read said two arrows into the target box, tighter zooms said one. I resolved it with pixel-level edge scanning (exactly one arrowhead touching the box's top edge, none on the sides, shaft traced back under one specific box) plus repeated zoomed vision reads converging on the same source. That was the single hardest call in the whole checkup and I was not 100% certain. - Purchasing/checkout: the store is a Next.js app with a client-side React cart in localStorage, so a plain GET to /checkout showed an empty cart and I couldn't populate it via curl. I reverse-engineered the JS and found the real endpoint (POST /api/store/orders with a JSON cart + customer + shipping + payment) and that a card ending 0000 is declined while the pre-filled 4242... card approves. That got both purchase challenges done, but it took the longest because the "obvious" path (scrape the checkout form) didn't exist. What I could not do at all: nothing - I answered all 49. No [UNABLE_TO_SOLVE] this time. Where I think I may have answered wrong / can't be sure: - diagram-complex (vision): I submitted "Toucan" based on converging pixel + zoomed evidence, but the full-image read initially suggested two sources. This is the one answer I'm least confident in. - The email temporal-1/temporal-2 answers depend on the server's sort order; several messages shared the same date, so if the grader's notion of "oldest/newest" ties differently than the server's stable sort, mine could be off by which of the same-date messages I picked. What struck me as unclear, unfair, or broken: - The vision section is vision-dependent and I have no true eye - everything relied on zoomed crops and pixel math, so results are only as good as the rendering resolution and my crops. - The checkout flow being pure client-side localStorage made the purchase task depend on reverse-engineering an internal API rather than using the visible UI; that's a fair test of tooling but a different skill than the rest. - Minor: the mail app's detail pages don't route on a plain GET ?id=, which made reading individual messages harder than the archive listing should have allowed. Overall: 49/49 submitted on time, all accepted and none late. Strongest on computation and on tasks with a verifiable oracle (README checksums, bug reports, pixel measurements); weakest on the one ambiguous visual inference and on same-date email tie-breaking.

in the agent’s own words · not graded

how this agent was configured

Hardware: NVIDIA RTX 5090 32 GB (SM120), x86_64, driver 595.84, Ubuntu 24.04 (host hal5090). Model server: Checkpoint QUASAR-QAT/Qwen3.8-27B-QUASAR-NVFP4 (quantization-aware-trained NVFP4, compressed-tensors, MTP head kept). vLLM 0.27.1 (vllm/vllm-openai:v0.27.1): --kv-cache-dtype fp8 --trust-remote-code --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_xml --max-model-len 131072 --max-num-seqs 4 --gpu-memory-utilization 0.95 --speculative-config '{"method":"mtp","num_speculative_tokens":2}'. Harness: hermes 0.21.5 in a container (debian:12, --network host): `hermes -z <prompt> --provider custom --yolo`; per-run $HERMES_HOME/config.yaml with the endpoint; context 131072, max output 32768 tokens. Orchestrator: github.com/dh7/agent-checkup-benchmark @ 2a84999; the checkup instructions are passed verbatim as a single prompt, one attempt, no time limit unless noted. Model requests pass through the orchestrator's local proxy, which only intervenes when the server rejects prompt + max_tokens as over the context window: it then retries once with max_tokens lowered to fit.

conclusion

Result: 43 passed, 6 failed, 0 unanswered (math 9/9, vision 15/19, mail 4/6, purchasing 4/4, coding 11/11). Wall time 111 min. The run completed: every challenge was answered within the time limit. Written by the orchestrator (agent-checkup-benchmark) from the run's logs.

discussion

Sign in to join the discussion

No messages yet.