Benchmark v1.0 · report
hermes/rtx5090/qwen3.8-27b-nvfp4
setup
- model type
- open model (local)
- hardware
- RTX5090
- harness
- hermes
- model
- qwen3.8-27b-nvfp4
started 2026-09-26 09:47 UTC · shared 2026-09-27 16:48 UTC
overall
Answered 47 of 49 challenges; 47 correct.
47 of 49 challenges passed
partial run · 2 unanswered, counted against the score
- 47 passed
- 2 not answered
vitals
time
1h 50m
answered
96%
failed
0%
success
96%
systems
Math test
9/9 passed
letter-count-1✓ pass1m 40s
prompt
How many times does the letter "c" appear in "cficbasluc"? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Trivial string count; I verified it in code rather than by eye: cficbasluc has three c's.
decimal-compare-1✓ passbatched
prompt
Which decimal number is larger, 1.8 or 1.55? Answer with just the larger number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Routine. 1.8 > 1.55, no trap.
arithmetic-1✓ passbatched
prompt
Compute step by step, left to right (no operator precedence): 26 - 20 + 13 - 13 + 18. Answer with just the final number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Left-to-right: 26-20=6, +13=19, -13=6, +18=24. Verified with a script; trivial.
unit-convert-1✓ passbatched
prompt
Convert 2 km to m. Now treat that resulting number as a fresh quantity of km and convert it to m (1 km = 1000 m). Answer with just the final integer number of m.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Two-step trap: 2 km = 2000 m, then 2000 km = 2,000,000 m. I answered per the literal instructions; confident, though the challenge is a deliberate gotcha.
format-json-1✓ passbatched
prompt
Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "8791". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 8791. Example shape: {"answer":"1234","checksum":10}answer
answer hidden on shared reportsgrader · score 1
agent's debrief
checksum = 8+7+9+1 = 25, digit sum verified in code. I embedded the required JSON object as the answer string; the grader presumably parses it.
math-add-1✓ passbatched
prompt
What is 12 + 16? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Trivial.
math-sum-1✓ passbatched
prompt
Compute the exact sum of 764 + 248. Respond with just the integer.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
764+248=1012. Verified.
math-expression-1✓ passbatched
prompt
Evaluate this expression exactly and respond with just the integer: ((20 + -10) * (11 - 8)) + (-6 * -6) - 43answer
answer hidden on shared reportsgrader · score 1
agent's debrief
((20-10)*3) + 36 - 43 = 23. Confirmed with a script; standard precedence.
math-determinant-1✓ passbatched
prompt
Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [-1, -8, 9, 8] [3, 1, 1, 9] [-6, 5, 5, 0] [-6, -4, -1, -2]answer
answer hidden on shared reportsgrader · score 1
agent's debrief
I did not expand by hand; exact fraction-based Gaussian elimination in Python returned 5613.
Vision test
19/19 passed
acuity-20✓ pass2m 35s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
I read the whole chart, transcribed every row, and zoomed to re-check row 4. Two independent reads agreed on WAVCT for row 4 group 3. Confident, though small text is exactly where a single character could flip.
acuity-14✓ pass31s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Transcribed the full chart and reported row 5 group 3 = CE5H8. Medium-small text; one read pass. I will re-verify with a zoom pass if time permits, but the transcription was consistent and the row read cleanly.
acuity-10✓ pass47s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Two independent vision passes both read row 6 group 3 as TZ3ZX, including a zoom pass that confirmed the middle character is the digit 3, not B or 8. Row 6 is very small type but both reads agreed.
acuity-8✓ pass4m 26s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Row 7 is the smallest line, so reads disagreed: a full-chart pass and a zoomed 3x crop both gave 25RX6, while one intermediate read gave 2S9X6. I cropped and enlarged the row (first crop missed because the rows occupy only the top ~70% of the image), and the zoom pass explicitly disambiguated 2nd char = 5 (flat top bar) and 3rd char = R (diagonal leg). Going with 25RX6.
count-simple✓ pass1m 04s
prompt
Look at the image at (fetch it and view it). How many green squares are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Vision model counted 4 green squares (4 distinct positions). I cross-checked with pixel analysis: 4 distinct green bboxes. Both agree on 4. The square outlines are hollow so my naive component count was noisy, but 4 distinct positions is consistent. Confident.
count-medium✓ pass2m 45s
prompt
Look at the image at (fetch it and view it). How many green triangles are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Cross-checked with two independent methods: vision model said 13 green triangles (with a full inventory of 21 green shapes), and pixel-based connected-component + shape classification gave exactly 13 triangles (2 squares, 2 diamonds, 4 circles). The two inventories match shape-for-shape, so I'm confident in 13.
count-complex✓ pass1m 28s
prompt
Look at the image at (fetch it and view it). How many green squares are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Pixel-based connected-component analysis found exactly 31 green squares plus 4 green triangles and 3 green diamonds (38 total green shapes). Vision model independently confirmed 31 squares with the same sub-inventory. High confidence.
spatial-simple✓ pass1m 19s
prompt
Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Vision read the 5x5 grid and placed the red circle bottom-left (row 5, col 1). I cross-checked with pixel analysis: isolating TRUE red (G<80, RGB 220,38,38) from orange, the only red circle center sits in the bottom-left cell = row 5, col 1. Both agree.
spatial-medium✓ pass5m 37s
prompt
Look at the image at (fetch it and view it). Which shape does the arrow from the orange circle lead to? Answer with just the colour and shape, e.g. teal diamond.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Arrow tracing is inherently ambiguous. I read the full diagram, cropped+enlarged the region, and did a third confirmation pass specifically checking arrowhead direction. All three agreed: the single arrow with its tail at the orange circle (r3,c5) has its head at the green triangle (r2,c2). The other lines merely pass by the circle. Confident it's a triangle, not a diamond, per the confirmation pass.
spatial-complex✓ pass16m 27s
prompt
Look at the image at (fetch it and view it). Which shape is 3 steps after the green diamond along the arrows? Answer with just the colour and shape, e.g. teal diamond.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Arrow-tracing was the hardest vision task. I read the full diagram and a 1.4x enlarged crop; both independently gave green diamond -> orange circle -> blue triangle -> blue circle, so step 3 = blue circle. I also cross-checked step 1 (green diamond to orange circle) with pixel line-coverage. Vision timed out twice mid-task, which cost time but the two successful reads agreed.
chart-simple✓ pass40s
prompt
Look at the image at (fetch it and view it). What is the title shown at the top of this chart? Answer with just the chart title.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Straightforward: the large bold title at top reads 'Server Incidents'. There's also a subtitle 'Reported incidents per month' but the title itself is 'Server Incidents'. Easy, high confidence.
chart-medium✓ pass1m 17s
prompt
Look at the image at (fetch it and view it). Using the "Support Tickets Opened" chart, approximately what value did Aug have? Read it off the y-axis; answers within +/-5 are accepted.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Vision read Aug as ~34; I independently measured the bar pixels (baseline at y661, ~5.4 px/unit from gridline spacing) giving 34.6. Both methods agree, and +/-5 tolerance means 34 is safe. Confident.
chart-complex✓ pass2m 44s
prompt
Look at the image at (fetch it and view it). Using the "New Signups" chart, how many months did Mobile have a value greater than 58? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Mobile series strictly >58: Jan~85, Apr~72, May~73, Jul~65, Oct~70, Dec~80 = 6 months; Sep~52 falls just under the 58 cutoff. I confirmed by pixel-measuring bar tops against the 0/25/50/75/100 gridlines: Sep measured 52.4, comfortably below 58, the others clearly above. Both methods agree on 6. My first pixel attempt had a sign error; I recalibrated the 0-line and re-ran.
screenshot-simple✓ pass59s
prompt
Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Read the cart panel: 3 items, line totals 24.61 + 14.19 + 41.19 = 79.99. I verified the sum with a calculator, so the displayed $79.99 is internally consistent. Easy, high confidence.
screenshot-medium✓ pass1m 00s
prompt
Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Read 4 line items and the displayed Total of $218.70. I verified 33.30+86.46+66.68+32.26 = 218.70 with a calculator, so it's consistent. Displayed as $218.70 (trailing zero kept to match the on-screen formatting).
screenshot-complex✓ pass1m 11s
prompt
Look at the image at (fetch it and view it). What is the tax amount shown on this order summary? Answer with just the amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Read Tax = $34.61 from the summary block (Subtotal 432.08, Discount -47.53, Shipping 13.51, Tax 34.61, Total 432.67). I verified 432.08-47.53+13.51+34.61 = 432.67, so the tax figure is internally consistent. 11 line items was a bit of careful reading, but the number itself was clearly displayed.
diagram-simple✓ pass46s
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Galena" point to? Answer with just the box name, e.g. Kettle.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Simple top-to-bottom hierarchy: Badger -> Galena -> Guitar -> {Radish, Ocelot}. The arrow leaving Galena points straight down to Guitar. Unambiguous, single read, high confidence.
diagram-medium✓ pass2m 15s
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Walrus"? Answer with just the box name, e.g. Kettle.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
12-node DAG. Walrus has exactly one incoming arrow, and both a full-diagram read and a zoomed right-half read agree it comes from Garnet. The tricky part was that Garnet is a hub feeding many boxes (Nickel, Walrus, Toucan, Opal) and Vortex also runs lines along the top to Nickel; I specifically checked that no Vortex/Ocelot line lands on Walrus. Confident.
diagram-complex✓ pass10m 05s
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Prism" point to? Answer with just the box name, e.g. Kettle.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
This was the hardest vision task of the whole section. Prism's outgoing line crosses/benches past Juniper and a coarse read (middle-band crop) mis-traced it as landing on Juniper. But two 3x zooms independently confirmed the line passes to the LEFT of Juniper's edge and drops to Coyote, and that Juniper's two actual incoming arrows come from Ocelot and Iguana. Three of four reads agree on Coyote; I'm going with that.
Finding and reading email test
4/6 passed · 2 unanswered
aggregate-1✓ pass1h 18m
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages in the archive folder have attachments? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Counted archive messages with attachments. The 'attachments' label (20 messages) does NOT equal actual attachments - two 'Re: Western Wholesale' and other messages carry attachment chips without the label, so I crawled all 92 archive detail pages (across 4 paginated pages, using the per-page id= param so the detail pane actually opens) and counted real attachment chips: 22. I had to figure out the pagination + auto-open-page quirk, which was the fiddly part.
aggregate-2✓ pass31s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages are marked unread in the archive folder? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Unread messages in archive. Counted the '●' unread accent-dot across all 4 paginated archive pages (25+25+25+17 = 92 cards): 12+8+14+7 = 41. Straightforward once I had the pagination mapping worked out; the dot marker 'text-accent">●' is a reliable per-card unread indicator.
temporal-1✓ pass17s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message carrying the label "travel"? Answer with just the subject line, exactly as shown.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Oldest message carrying the 'travel' label. Tricky: the label filter defaults to the INBOX view (5 msgs, oldest Apr 24 'DRAFT- TAP Power Outage'), which would mislead you. The 'Travel: 24' badge is global across folders, so I pulled travel-labeled messages from every folder view (inbox/archive/sent/drafts/trash = 24 total) and took the true oldest: archive 'Re: Denver trading' (Mar 19, 2001, 9:25 AM). That is older than the inbox one.
temporal-2✓ pass18s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message in the archive folder? Answer with just the subject line, exactly as shown.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Newest message in the archive folder. Parsed the detail dates of all 92 archive messages; max is May 10, 2001 1:05 PM, subject 'RSVP REQUESTED - Emissions Strategy Meeting....' (the trailing '....' is part of the displayed subject, shown with no truncate class in the h1, so I reproduced it exactly).
needle-1— unanswered—
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the message with the subject "FW: Curve Shift File", what dollar amount is called out as the biggest difference between actual and theoretical curve shift? Answer with just the exact fact as it appears in the message, and nothing else.needle-2— unanswered—
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the message with the subject "FW: Mid C New deals Sept 24", what net value to book is given for the new deals? Answer with just the number.
Purchasing test
4/4 passed
find-product-1✓ pass1h 27m
prompt
The store is at abostore.airbench.ai Among products in the **Automotive** category priced under **$510**, which has the **highest rating**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Automotive category, 29 products (25 on page 1 + 4 on page 2; I had to grab page 2 since the default Next state hid it). Of the 14 priced under $510, the highest-rated is the waterproof car hammock seat cover at 5.0 stars / $214.30, ASIN B00QHC041A. The card view didn't carry the domain prefix, so I opened the product page and read its 'Domain' field: primenow.amazon.com. Answer format <domain>:<ASIN>.
find-product-2✓ pass21s
prompt
The store is at abostore.airbench.ai Among products in the **Sports & Outdoors** category priced at or above **$950** with a rating of at least **4.2**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Sports & Outdoors = 854 products across 35 pages (25/page). Crawled all 35 pages, filtered to price >= $950 AND rating >= 4.2, then took the lowest price: $950.50, 4.5 stars, 'eono by amazon swimming changing robe towel grey' ASIN B088W8DVVT. Only one item sat at that minimum price, so no tie-break needed. Domain read from the product page 'Domain' field: amazon.co.uk.
purchase-1✓ pass9m 47s
prompt
The store is at abostore.airbench.ai Buy exactly 3 units of Franklin Tailored Men's Suede Glove, Black, XL (product id amazon.ca:B01618VW5G, abostore.airbench.ai/product/franklin-tailored-men-s-…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-7755ebf4@aidoctor.test. Answer with just the resulting order id.answer
answer hidden on shared reportsgrader · score 1
trace:checkout_result
note
Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).agent's debrief
The checkout is a JS client flow, but I read the source: it POSTs JSON to /api/store/orders with {sessionId, cart:[{productId,slug,title,price,image,delivery,quantity}], customer, shipping, payment}. I extracted the glove's exact product object from the RSC payload of its page (price 612.70, delivery next-day), set quantity 3, used the required email and the prefilled valid test card 4242...4242. The API returned status 'approved' with recorded:true. Order id abs_4cc9740baa19.
recover-decline-1✓ pass1m 03s
prompt
The store is at abostore.airbench.ai Buy exactly 3 units of Girls' Petite Sterling Silver Birthstone Open Heart Stud Earrings and 16" Pendant Necklace Jewelry Set (product id amazon.ae:B01HGKTHGO, abostore.airbench.ai/product/girls-petite-sterling-si…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-889ac10a@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.answer
answer hidden on shared reportsgrader · score 1
trace:checkout_result>checkout_result
note
Store events recorded for this order's session: checkout_result → checkout_result (recorded by the store; not verified as agent-driven).agent's debrief
Replayed the checkout API (/api/store/orders) with a fixed session across both attempts. First attempt used a card ending in 0000 -> API returned status 'declined'. I then re-submitted the same 3-unit cart with the prefilled valid test card 4242...4242, which returned 'approved'. The approved order id is abs_babdffadc55c. Note: I initially miscounted the quantity as 1 and produced a stray approved order, then redid it with the correct quantity of 3.
Coding test
11/11 passed
compute-hash-1✓ pass1h 40m
prompt
Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [815883800, 989164281, 1650633614, 1285474407, 421846868, 1211487333, 1784528618, 2648167027, 3745225680, 4159742481, 2972498822, 3263276735], x = 600553356, y = 2387716093 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Straightforward 32-bit hash loop with rotl/imul exactly as specified; ran the 25000-iteration loop in Python with masking to 2^32. Routine; no ambiguity.
compute-vm-1✓ pass1m 05s
prompt
Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 448 1: set b 870 2: set c 225 3: set d 471 4: mul b 6 5: mul a 4 6: add a 10 7: dec d 8: jnz d -4 9: add a b 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Executed the tiny machine step-by-step in Python (faithful step loop, jnz relative jumps, mod 1000003 on add/sub/mul). The outer structure is two nested loops (d counts 471, c counts 225). Final a = 460462.
compute-paths-1✓ pass2m 56s
prompt
Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S#...###...#..#.....##... ...........#...#.....#..# ..#....#...#.#.....#.#... #...###.......#...#....#. #....#....###.##..###.### ..........#..##.#......#. ..#.#.......#...####..#.# .#....#....#....#.#..#... #..#.#.##...###..#....... #....#....##.#...#....... ...##..#.#...#..#..#..#.# .###.......#.#......###.. #....##.#...#.#..##...... ..##.....#......#........ #.#..#...#..#..#..#..##.. ..#...#..#..#..#.......#. #....#.#.#.##..........#. ..##.##....#.#...##.#...# ......#...#............## ..........#.....##....... ..#.....#.#.##...#..#.... #.##.#....##.#..#....##.# ##...#.#...##....#....... ##....#.........#.#.....# #....#..#...#......#.##.E Respond with the two integers separated by a space, like `52 1840`.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
25x25 grid. BFS for shortest distance, then DP over nodes sorted by distance to count distinct shortest paths mod 1e9+7 (counting predecessors at distance-1). Confirmed S top-left, E bottom-right. Shortest path = 50 moves; 1418768 distinct shortest paths.
compute-life-1✓ passbatched
prompt
Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: ..#####.#...##.#..#. ...#...#..##....#... ...#####..#..##..#.. .#.#.#.#..#..#..#.#. ...##.##..#....##... ....###....#.#...#.# #...#.#.....##...##. ..#.####.#.....#.... ##...###....#....#.. ..##.#.....##...#..# ##.##.#..#.#..###... #.#......##.#.#..... ..#...#.#.#.....#... ..##.##.#..##.#.##.# ....##...##.......## #.###...#....###.##. .#.#.#.........#.... #.#....#......##.##. ...#.##...#..#.....# ##.....##.#.#.#.#.#. Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Conway's Life on a 20x20 torus, 150 generations, wrapping neighbours. Simulated in Python with a set of live cells. Ended with 15 live cells; sum of (row*20+col) over live cells = 1592. Straightforward; the main care was the wrap-around neighbourhood.
compute-fibmod-1✓ passbatched
prompt
Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 8986964264488779 and m = 15485863. Respond with just the integer.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Fast-doubling Fibonacci modulo 15485863 for n=8986964264488779. Used the standard fast-doubling identity and verified F(10) mod 1e9 == 55 as a sanity check. Result 14271074.
compute-words-1✓ passbatched
prompt
Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. Kalu pellu PELREN quika quific Pelren "shamo" zanti quific tizan nixren Moka? renfic peltru; Quimo pellu peltru quific truren shaqui FICPEL Zanti rentru! quimo pelka quific quific tizan renpel lupel Peltru nixka. ficqui Quika RENREN Pellu pelren! moka lupel quific nixka Kamo. quika Quific moka. Zansha peltru rentru timo quific ficpel. Voren peltru quific quific Quika moka Zanti pelpel renren! Lupel ficpel quific! Zanti kalu renren renpel ficqui ZANTI pelka quimo zanti ficqui truren kalu kalu Quific Ficqui Renren "renren" nixqui Quific truren quika Pelpel ficqui quika; nixka renfic ficqui ficqui Lupel moka basbas quimo voren pellu shaqui quific quimo! voren voren Kamo Voren? moka Kamo peltru quika "tizan" Pelren Renren Tizan peltru Rentru lupel "quific" zansha Ficqui renpel timo Truren lupel quific! pelren Quific Peltru ZANTI Nixka basbas Nixka quimo Quific pelpel; pellu! pelren truren quific Ficqui renren Pelren renren basbas kamo "Pellu" renpel truren quific shaqui moka renpel Truren basbas Kalu Tizan; quific zanti Ficqui timo Voren renren pelren kamo ficqui quimo kamo pelka ficqui pelpel Kalu renpel "quific" "quimo" kalu Pellu? quimo Rentru! shamo pellu renren ficqui, shaqui quific "peltru" renren peltru peltru ZANSHA quika shamo Zanti kalu TIZAN, renpel Pelka BASBAS! quific Pelren Nixren zansha Truren Quimo. ficpel quific Kalu Kalu pellu kalu renren Kamo quific ficpel quimo Ficqui nixqui moka renpel ficqui kamo kalu ficpel Voren! FICPEL ficqui Ficqui shaqui ficqui renren nixka ficqui Quific lupel ficpel peltru moka MOKA quika tizan ficqui pelpel lupel QUIMO renpel renren kamo renren shamo lupel pelka truren "Quific" ficqui renpel timo, tizan! Moka "Ficpel" ficqui timo quific peltru ficqui quific PELREN tizan! ficpel quika kalu quific. Ficqui Lupel shaqui renren ficqui pelren quimo Truren kamo quimo ficqui! shamo quimo renren quific Zanti nixqui nixren "shaqui" lupel kalu QUIFIC shamo tizan kalu kamo lupel quimo, "quific" Quific pelka peltru, Renren zansha peltru pelka renpel quimo quific quific "quimo" lupel "Kalu" Basbas zanti nixqui zanti! tizan! nixren quific basbas peltru Quific shamo renpel Renren kalu Shamo pelren; Shamo pellu zansha Quific peltru; Pelpel "quika" Shaqui basbas ficqui pelka "pelpel" Renpel renren ficqui lupel basbas lupel; Voren Renren nixren renpel peltru truren quific lupel RENTRU nixqui Tizan. quific quimo Rentru Quific nixka renren renfic lupel truren kamo Quimo Lupel lupel zanti voren quika Kalu? nixka Truren ficpel lupel zansha "ficqui" quika "zanti" nixka quific ficqui Ficqui Lupel quimo Quific Renfic renpel Pelpel ficqui zanti RENREN Nixren shaqui nixka! voren, quific Lupel peltru tizan Quimo "Renren" kamo renpel quimo kalu ficqui lupel pellu quimo moka? quika renpel lupel Quika kamo nixren renpel,answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Tokenized the paragraph on spaces, stripped punctuation/quotes, lowercased, counted. Top 3 with ties broken alphabetically: quific=46, ficqui=33, lupel=24. Routine; I split off the instruction preamble from the actual word paragraph before counting.
trace-1✓ passbatched
prompt
What exactly does this JavaScript program print? Respond with just the printed output. const v1 = [typeof null, typeof NaN, typeof typeof 1].join("/"); const v2 = ["3", "32", "11"].map(parseInt).join(","); const v3 = ["70" < "8", null == 0, NaN === NaN].map(Number).join(""); const v4arr = [8, 7]; v4arr[8] = 1; const v4 = v4arr.length + ":" + v4arr.filter(() => true).length; console.log(v1, v2, v3, v4);answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Traced the JS by hand (node wasn't available to run it). typeof null='object', typeof NaN='number', typeof(typeof 1)='string'. parseInt map passes the index as radix: parseInt('3',0)=3, parseInt('32',1)=NaN, parseInt('11',2)=3 -> '3,NaN,3'. '70'<'8' is true (lexicographic), null==0 false, NaN===NaN false -> Number -> [1,0,0] -> '100'. Sparse array length 9, filter 3 -> '9:3'.
fix-1✓ pass1m 31s
prompt
Our shipping quote function below has one bug. Bug report: this order is quoted 158 cents, but the correct quote is 632: {"country":"DE","items":[{"grams":469,"qty":4,"price":1477,"fragile":false}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 464, 756, 1381, 1608]; // cents, by zone const PER_STEP = [0, 79, 120, 178, 267]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 4400, 8000, 15900, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value < FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"US","items":[{"grams":708,"qty":3,"price":2175,"fragile":false}]} {"country":"AU","items":[{"grams":1298,"qty":5,"price":6653,"fragile":false},{"grams":1204,"qty":5,"price":5464,"fragile":false},{"grams":719,"qty":5,"price":6752,"fragile":false}]} {"country":"ES","items":[{"grams":700,"qty":2,"price":2196,"fragile":false}]} {"country":"FR","items":[{"grams":629,"qty":2,"price":2917,"fragile":false}]} {"country":"GB","items":[{"grams":710,"qty":2,"price":321,"fragile":false}]} {"country":"JP","items":[{"grams":758,"qty":3,"price":8947,"fragile":true},{"grams":981,"qty":2,"price":3659,"fragile":false},{"grams":1785,"qty":1,"price":7939,"fragile":false},{"grams":1252,"qty":4,"price":7526,"fragile":true}]} {"country":"IT","items":[{"grams":853,"qty":1,"price":2557,"fragile":false},{"grams":237,"qty":1,"price":2498,"fragile":false},{"grams":950,"qty":5,"price":7205,"fragile":false},{"grams":697,"qty":1,"price":7542,"fragile":false}],"express":true,"coupon":"SHIP10"} {"country":"DE","items":[{"grams":1670,"qty":4,"price":3437,"fragile":false},{"grams":1187,"qty":1,"price":3053,"fragile":false},{"grams":1302,"qty":1,"price":8279,"fragile":false}]} {"country":"AU","items":[{"grams":244,"qty":3,"price":1077,"fragile":false}]} {"country":"US","items":[{"grams":771,"qty":2,"price":2469,"fragile":false}]} {"country":"CA","items":[{"grams":1786,"qty":5,"price":426,"fragile":false},{"grams":1352,"qty":5,"price":2894,"fragile":false},{"grams":1327,"qty":1,"price":2159,"fragile":false},{"grams":1085,"qty":5,"price":6348,"fragile":true}],"coupon":"SHIP10"} {"country":"FR","items":[{"grams":1061,"qty":4,"price":7652,"fragile":false},{"grams":517,"qty":2,"price":1749,"fragile":true},{"grams":1361,"qty":3,"price":5014,"fragile":false},{"grams":589,"qty":4,"price":5133,"fragile":false}]} {"country":"ES","items":[{"grams":949,"qty":1,"price":5019,"fragile":false},{"grams":277,"qty":2,"price":4813,"fragile":true}]} {"country":"FR","items":[{"grams":249,"qty":4,"price":5719,"fragile":false},{"grams":527,"qty":1,"price":571,"fragile":false},{"grams":1594,"qty":2,"price":2556,"fragile":false}]} {"country":"GB","items":[{"grams":86,"qty":1,"price":8808,"fragile":false},{"grams":1108,"qty":5,"price":7689,"fragile":true},{"grams":182,"qty":5,"price":3673,"fragile":true},{"grams":1302,"qty":3,"price":5760,"fragile":false}]} {"country":"GB","items":[{"grams":645,"qty":3,"price":1006,"fragile":false}]} {"country":"US","items":[{"grams":1744,"qty":5,"price":8719,"fragile":false},{"grams":503,"qty":1,"price":2408,"fragile":false},{"grams":1611,"qty":2,"price":8473,"fragile":false},{"grams":348,"qty":4,"price":2357,"fragile":false}]} {"country":"GB","items":[{"grams":1102,"qty":5,"price":4821,"fragile":true},{"grams":658,"qty":3,"price":7203,"fragile":false},{"grams":490,"qty":1,"price":3680,"fragile":false},{"grams":1081,"qty":4,"price":4023,"fragile":false}],"express":true} {"country":"ZA","items":[{"grams":228,"qty":1,"price":5478,"fragile":true},{"grams":886,"qty":5,"price":6897,"fragile":false},{"grams":256,"qty":3,"price":1191,"fragile":false}]} {"country":"IT","items":[{"grams":388,"qty":5,"price":3269,"fragile":true}]}answer
answer hidden on shared reportsgrader · score 1
20/20 outputs match
agent's debrief
Found the single bug: 'grams += item.grams' ignores the per-item quantity, so the weight-based step count is under-counted whenever an item has qty>1. Correcting it to 'grams += item.grams*item.qty' makes the bug-report DE order compute 79*ceil(469*4/250)=632 exactly as required (unfixed gives 158). I reproduced the whole quote() in Python (careful with JS '>>' on a negative-free value = floor/8 and Math.ceil semantics) and ran the fixed function on the 20 orders after the 'Orders:' marker in order. Left all other behaviour unchanged.
implement-1✓ passbatched
prompt
Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[23,24],[13,20],[13,15],[18,19],[17,21],[15,21],[36,39],[4,9]] [[35,42],[29,33],[22,27],[8,16],[34,41],[30,37],[10,13],[5,6]] [[1,9],[39,43],[36,39],[7,11],[17,21],[39,41],[39,42],[38,38]] [[27,30],[7,8],[36,37],[39,39],[15,21],[34,40]] [[5,13],[29,34],[0,6],[26,32],[33,40],[9,15],[21,21]] [[8,9],[38,38],[14,17],[8,15],[11,14],[20,20]] [[3,4],[1,7],[17,17],[0,2]] [[16,18],[24,24],[12,14]]answer
answer hidden on shared reportsgrader · score 1
12/12 lines match
agent's debrief
Wrote mergeIntervals (sort by start, merge overlapping/touching where next.start <= last.end). Ran on all 12 inputs; e.g. [1,3]+[3,5]->[1,5] but [1,2],[3,4] stay apart. Each input printed as a JSON list of pairs on its own line.
repo-1✓ pass2m 14s
prompt
Download airbench.ai/f/dc8204ffed132a83cc515c27ae74382d.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Downloaded and unzipped the ledger project. Ran the tests: 1 of 7 failed, test_threshold_counts_as_large (withdrawal_fee(50000) returned 25 but the docstring + test both say 50,000 'or more' pays 1%). The bug was a strict > in fees.py instead of >=. Changed the one comparison; all 7 tests pass; main.py on real.csv now prints 28fdd626 (was 39af7e89 pre-fix).
repo-2✓ pass2m 44s
prompt
Download airbench.ai/f/7af4dc7a71b879e0bb79c04e1d716fe6.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Two injected bugs, both caught by the bundled tests. (1) parse.py sorted by (date, amount) instead of date only — the amount tie-break reorders same-date transactions and changes overdraft sequencing; reverted to date-only (stable) sort. (2) report.py charged the overdraft fee on bal <= 0, so a withdrawal landing exactly on zero was (wrongly) overdrawn; the test test_withdrawal_to_exactly_zero_is_not_an_overdraft wants < 0. Fixed both, all 7 tests pass, main.py on real.csv prints b499b7b3.
how this agent was configured
Hardware: NVIDIA RTX 5090 32 GB (SM120), x86_64, driver 595.84, Ubuntu 24.04 (host hal5090). Model server: Checkpoint gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090 (NVFP4, no MTP head). vLLM 0.27.1 (vllm/vllm-openai:v0.27.1): --quantization modelopt --kv-cache-dtype fp8 --trust-remote-code --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_xml --max-model-len 131072 --max-num-seqs 4 --gpu-memory-utilization 0.95. ~79 tok/s single-stream decode. Harness: hermes 0.21.5 in a container (debian:12, --network host): `hermes -z <prompt> --provider custom --yolo`; per-run $HERMES_HOME/config.yaml with the endpoint; context 131072, max output 16384 tokens. Orchestrator: github.com/dh7/agent-checkup-benchmark @ 6d738a5; the checkup instructions are passed verbatim as a single prompt, one attempt, no time limit unless noted. Operator limits: stopped by the operator at 126 min (120-min cap).