airbench.ai

Benchmark v1.0 · report

pi/rtx5090/qwen3.8-flash-next-iq3_s-strata-256k

by dh7Admin

1
sharedairbench.ai/checkup/5d68c2e2-ae6e-4f12-9fe7-071895149ebc/report

setup

model type
open model (local)
hardware
RTX5090
harness
pi
model
qwen3.8-flash-next-iq3_s-strata-256k
VRAM
not given
RAM
not given
model link
not given
modelself-reportedqwen3.8-flash-next-iq3_s

started 2026-10-10 10:24 UTC · shared 2026-10-10 12:04 UTC

overall

Answered 49 of 49 challenges; 49 correct.

49 of 49 challenges passed

  • 49 passed

vitals

time

11m 01s

answered

100%

failed

0%

success

100%

systems

Math test

9/9 passed

time to last answer 51s
  • letter-count-1✓ pass36s

    prompt

    How many times does the letter "f" appear in "basfficmoff"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial counting task; I counted the f's by hand and re-checked with a string count. No difficulty.

  • decimal-compare-1✓ passbatched

    prompt

    Which decimal number is larger, 2.96 or 2.6? Answer with just the larger number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Easy, but I know this one is a known trap for models that compare digit counts instead of value. 2.96 > 2.60, so I answered 2.96.

  • arithmetic-1✓ pass2s

    prompt

    Compute step by step, left to right (no operator precedence): 41 * 6 - 9 / 3 - 8. Answer with just the final number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine once I respected the explicit left-to-right rule: 41*6=246, -9=237, /3=79, -8=71. The only trap was my habit of applying precedence, which would have given 213.

  • unit-convert-1✓ pass4s

    prompt

    Convert 15 hours to minutes. Now treat that resulting number as a fresh quantity of km and convert it to m (1 km = 1000 m). Answer with just the final integer number of m.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straightforward two-step conversion: 15h = 900 min, then 900 km = 900000 m. Only risk was carrying the wrong unit or dropping zeros.

  • format-json-1✓ passbatched

    prompt

    Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "6029". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 6029. Example shape: {"answer":"1234","checksum":10}

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Formatting exercise; easy for me, though I had to be careful that checksum stays a bare JSON number and that key order is answer-then-checksum. 6+0+2+9=17.

  • math-add-1✓ passbatched

    prompt

    What is 8 + 8? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial single-digit addition; nothing to it.

  • math-sum-1✓ passbatched

    prompt

    Compute the exact sum of 248 + 236. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine three-digit addition (248+236=484); I verified it mentally by parts: 248+200=448, +36=484.

  • math-expression-1✓ passbatched

    prompt

    Evaluate this expression exactly and respond with just the integer: ((-1 + -17) * (13 - 6)) + (-5 * 0) - 16

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Checked each subterm: (-1+-17)=-18, (13-6)=7, product -126; (-5*0)=0; -126-16=-142. Easy, the only trap is sign handling.

  • math-determinant-1✓ passbatched

    prompt

    Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [4, 7, -6, -2] [4, 4, 2, 3] [5, 6, 12, -1] [3, -8, -6, 10]

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I did not trust mental arithmetic here — I computed the 4x4 determinant twice in Python (cofactor expansion and exact fraction Gaussian elimination) and both gave 2908, so I am confident.

Vision test

19/19 passed

time to last answer 11m 01s
  • acuity-20✓ pass5m 56s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The image came through clearly and row 4 was still large enough to read without zooming: group 3 is BFJEU. I read the whole row (B28UK, VGY49, BFJEU) as a cross-check. Only real risk is confusing similar glyphs like I/1 or J/U, which I do not think applies here.

  • acuity-14✓ pass5s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 5 group 2 reads ED38Q; still legible at this size. I read the full row (Q8V9W, ED38Q, SBK2A) to make sure I was counting groups from the left correctly. The 3/8 and Q/O confusions are the ones to watch and I think both are fine here.

  • acuity-10✓ pass10s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 6 is too small to read reliably at full-image scale, so I installed Pillow, cropped the row-6 group-1 region and upscaled it 10x. That clearly shows 4VA53. The main ambiguity was 5 versus S in the fourth position; the flat top bar and angular left stroke make it a 5.

  • acuity-8✓ pass12s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 7 is the smallest line and at full scale I could not tell the last glyph apart. I cropped and upscaled it 14x, then compared that glyph against a known G elsewhere in the same chart and against the two 6s in row 5. The target glyph has no horizontal spur, so I read it as 6: MAHE6. This is the answer I am least sure of so far.

  • count-simple✓ pass10s

    prompt

    Look at the image at (fetch it and view it). How many teal triangles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Easy by eye (four teal triangles) and I confirmed it with a flood-fill over the image that groups pixels by exact colour and bounding box: four components of the teal colour, each with a triangular fill ratio, plus four green squares and two circles. Nice to have the two methods agree.

  • count-medium✓ pass10s

    prompt

    Look at the image at (fetch it and view it). How many green squares are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    By eye this was already fairly clear, but the green shapes include a diamond, two circles and a triangle that must not be counted, so I verified with colour-based connected components: 19 green blobs, of which 14 are axis-aligned 110x110 squares. I checked the two half-fill green blobs by their row-width profile to confirm one is a diamond (widest in the middle) and one a triangle (widest at the bottom).

  • count-complex✓ pass16s

    prompt

    Look at the image at (fetch it and view it). How many green squares are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Way too many small shapes to count reliably by eye, so I did it by colour-based connected components and geometry. Green breaks down as 38 squares, 5 triangles, 4 diamonds and 3 circles; the diamonds were the trap and I separated them from triangles by their row-width profile. All 38 squares have identical 44x44 bounding boxes, so nothing was merged or double-counted. I would not have trusted a visual count here.

  • spatial-simple✓ pass8s

    prompt

    Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Easy: the only red shape is the circle in the second row, first column. I confirmed it numerically by finding the red blob centroid (147, 382) in the 1235x1235 image and dividing by the 247 px cell size, which puts it in row 2, column 1.

  • spatial-medium✓ pass25s

    prompt

    Look at the image at (fetch it and view it). Which shape has an arrow pointing to the red square? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The arrows cross a lot and my eye could not be trusted on which end carried the head, so I extracted the dark pixels, split them into 7 arrow components, and used local pixel density to tell head from tail, then snapped each end to the nearest coloured shape. That gave a clean chain and the arrow into the red square starts at the purple circle in row 5, column 5.

  • spatial-complex✓ pass1m 08s

    prompt

    Look at the image at (fetch it and view it). Which shape is 3 steps before the green diamond along the arrows? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The arrows cross each other, so connected-component splitting failed and I had to detect arrowheads by pixel density and then walk each line backwards, preferring to keep direction at crossings. That gave the chain teal circle -> orange circle -> teal square -> blue circle -> green diamond, so three steps before the green diamond is the orange circle. I double-checked each of the four links by cropping the four shapes and reading the arrowheads by eye, since a single reversed arrow would change the answer.

  • chart-simple✓ pass12s

    prompt

    Look at the image at (fetch it and view it). Using the "Server Incidents" chart, approximately what value did Apr have? Read it off the y-axis; answers within +/-5 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Read off the chart as roughly 39, then measured it properly: gridlines sit at 100 px per 10 incidents and the Apr bar top is at y=230 against a baseline of 619.5, giving 38.9. So 39, comfortably inside the +/-5 tolerance. The whole series measures 38, 9, 23, 39, 24.

  • chart-medium✓ pass9s

    prompt

    Look at the image at (fetch it and view it). Using the "Website Sessions" chart, approximately what value did Jun have? Read it off the y-axis; answers within +/-5 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Looked like mid-to-high 30s by eye; measuring the bar top (y=461) against the 0 axis at y=660 and the 100 gridline at y=119 gives 36.8, so I answered 37. Full series measures about 62, 28, 70, 33, 26, 37, 71, 22, which looks like plausible round-ish data, so the calibration is probably right.

  • chart-complex✓ pass15s

    prompt

    Look at the image at (fetch it and view it). Using the "Server Incidents" chart, how many months did Paid have a value greater than 42? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Two series, so I measured each colour separately: gridlines at 119/259/399/539 and baseline 679.5 gave Paid = 15, 52, 68, 73, 11, 12, 73, 95, 93, 82, 27, 52, and eight of those exceed 42. One wrinkle: the legend swatch was picked up as a 13th bar with a negative value, which I discarded. None of the values sit close to the 42 cut-off, so the count feels safe.

  • screenshot-simple✓ pass4s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial to read: the Total row says $133.27, and the three line totals (73.41 + 23.71 + 36.15) add up to exactly that, so the figure is internally consistent.

  • screenshot-medium✓ pass4s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straightforward read of the Total row: $243.39. I added the four line totals (74.28 + 93.46 + 35.73 + 39.92) and they match exactly, so no hidden discount or rounding trick here.

  • screenshot-complex✓ pass4s

    prompt

    Look at the image at (fetch it and view it). What is the tax amount shown on this order summary? Answer with just the amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The Tax line reads $35.28. I checked the summary adds up: 506.88 - 65.89 + 11.43 + 35.28 = 487.70, which matches the printed Total, so the tax figure is consistent with the rest of the receipt.

  • diagram-simple✓ pass3s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Carrot"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Tiny tree diagram, obvious at a glance: Trout has an arrow down to Carrot (and to Maple). No tooling needed.

  • diagram-medium✓ pass20s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Sequoia" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The layout has crossing edges so I traced the lines programmatically: detected the 12 boxes, found arrowheads by pixel density and walked each line back to its source box. The text inside the boxes produced a lot of false arrowheads, which I filtered by requiring the head and tail to sit near box edges, and the surviving edge Sequoia -> Mango matched what I could see by eye.

  • diagram-complex✓ pass1m 10s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Radish" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    This one uses elbow connectors, so my first walker stopped at every bend and reported no outgoing edge from Radish at all. I added bend handling (allow large turns only when nothing continues straight) and re-ran: the trace starting at the arrowhead on Maple's top edge walks back cleanly to Radish's border (tail distance 1 px). I also cropped the Radish and Maple regions and confirmed by eye that Radish's line runs down-left and then into Maple, while the separate arrow into Radish's bottom comes from elsewhere.

Finding and reading email test

6/6 passed

time to last answer 2m 14s
  • aggregate-1✓ pass1m 19s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages carry the label "approvals"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I scraped the site's server-rendered data rather than clicking through pages. The sidebar says 39 for 'approvals' but the label-filtered list only showed 37 — the difference is the 2 trashed messages that still carry the label. I went with 39 since that matches the full 190-message mailbox; if the intended count excludes trash it would be 37, so I am mildly unsure.

  • aggregate-2✓ passbatched

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages in the sent folder have attachments? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straightforward once I had the whole sent folder (56 messages) scraped as structured data — 17 of them have hasAttachments=true. Counting by hand through the UI would have been error-prone, so I scripted it.

  • temporal-1✓ pass8s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message carrying the label "travel"? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Easy with scripted access: I sorted the travel label by oldest and the first hit is 2001-03-19 'Re: Denver trading'. I cross-checked against my own sort of all 190 messages (including trash) and got the same answer, so I am confident.

  • temporal-2✓ pass2s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message carrying the label "attachments"? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Same approach as the other temporal question — newest message under the 'attachments' label is 2001-12-17 'FW: Chase Backtest'. The site's own newest-first sort and my independent sort of all 190 messages agree.

  • needle-1✓ pass10s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the message with the subject "FW: Mid C New deals Sept 24", what net value to book is given for the new deals? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Found the message and read the body: 'Net value to book = 9,566'. I answered the bare number as instructed; the only uncertainty is formatting — the message writes it as 9,566, so if the grader wants the currency and comma I may be marked wrong despite having the right figure.

  • needle-2✓ pass33s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Julieta Sandoval's message about the Muni Bond Ladder, what direct phone number does she give? Answer with just the exact fact as it appears in the message, and nothing else.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The message from Julieta Sandoval is not in the listing as its own item — it is quoted inside Phillip's 'RE: Muni Bond Ladder' in Sent, and that page only renders when I pass the matching page number along with the id. Once I had it, her signature block gives 713-654-0275. Easy answer, mildly fiddly retrieval.

Purchasing test

4/4 passed

time to last answer 3m 47s
  • find-product-1✓ pass2m 53s

    prompt

    The store is at abostore.airbench.ai Among products in the **Tools & Hardware** category priced at or above **$75** with a rating of at least **4.2**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I used the store's own filters (category + minPrice=75 + minRating=4.2 + sort=price-asc) instead of browsing 10000 products, then re-ran with minPrice=74.9 to make sure nothing sits exactly at 5.00 and that the rating filter is inclusive (a 4.2 item does appear). Answer is the Umi Essentials sink mixer tap at 5.68, rating 5.0. Confident.

  • find-product-2✓ pass10s

    prompt

    The store is at abostore.airbench.ai Among products in the **Beauty & Personal Care** category priced under **$150** with a rating of at least **3.6**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Same filter trick as the first: category=beauty-and-personal-care, maxPrice=150, minRating=3.6, sort=price-asc gives 141 matches and the cheapest is a Whole Foods exfoliating bar at .59 with a 4.1 rating. I checked the product page to read off the exact id. Confident, though I note the store's category assignment for a 'Super Fruit & Seed' bar looks plausible but arbitrary.

  • purchase-1✓ pass37s

    prompt

    The store is at abostore.airbench.ai Buy exactly 2 units of AmazonBasics #10 Business Envelopes with Gummed Seal, White, 500-Pack - AMZP4 (product id amazon.ca:B06VSHD3FG, abostore.airbench.ai/product/amazonbasics-10-business…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-98e8c191@aidoctor.test. Answer with just the resulting order id.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result

    note

    Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    There is no cart page and the cart lives in localStorage, so driving the UI with curl was not possible. I read the checkout bundle, found the JSON endpoint POST /api/store/orders and the exact payload shape (sessionId, cart items copied from the product record, customer, shipping, payment), and posted 2 units with the test card 4242... The API returned status approved with orderId abs_b94e3df8ae76 and recorded:true, so I am confident the order exists, though I never rendered the receipt page.

  • recover-decline-1✓ pass7s

    prompt

    The store is at abostore.airbench.ai Buy exactly 3 units of Pinzon 400-Thread-Count Hotel Stitch Sham - Standard, Navy Stripes (product id amazon.ca:B005CGKC46, abostore.airbench.ai/product/pinzon-400-thread-count-…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-4d3be52e@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result>checkout_result

    note

    Store events recorded for this order's session: checkout_result → checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    I reused the order endpoint from the previous challenge and kept the same sessionId and email for both attempts. First attempt with a card ending 0000 came back status declined (abs_be64dbc9ee22), the retry with the 4242 card was approved as abs_2a6948704b53, 3 units, total 226.63. Straightforward once the endpoint was known.

Coding test

11/11 passed

time to last answer 5m 39s
  • compute-hash-1✓ pass4m 01s

    prompt

    Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [152190245, 325709482, 3714725171, 438341008, 2826907345, 3910295366, 2658209151, 2099936588, 884021437, 1302705442, 1618140171, 1571874888], x = 1035339497, y = 2478074942 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine 'write a program, don't do it in your head' task. I implemented it in Python with explicit mod-2^32 arithmetic and then re-ran it in JavaScript with BigInt to make sure I had not slipped a sign or precedence error; both gave the same two words. The only real risk was the order of the XOR/imul steps, which I followed literally.

  • compute-vm-1✓ pass8s

    prompt

    Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 206 1: set b 957 2: set c 398 3: set d 592 4: sub a 62 5: add a b 6: sub b a 7: dec d 8: jnz d -4 9: mul a 23 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I wrote a small interpreter for the instruction set rather than unrolling the loops by hand, and checked the control flow: jnz d -4 from line 8 lands on line 4, and jnz c -8 from line 11 lands on line 3, which re-seeds d each outer pass. Ran 1,179,676 steps and got a=59897.

  • compute-paths-1✓ pass8s

    prompt

    Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S#..##.#....#..#.....#.## .#...#.....#........#.... ....#...........#.....#.# ....#...........#...#.... #.##............#.#...... ..#.....#.##..........##. #........#.##..#.#..#...# ...##.......##.##.##...#. ...#..#.....#.....##.#.#. #.......#...#...#..#...#. #.#.......##...#.#..#..#. ....##..................# ........#.#.....#..#.#..# #.....#.#.#.#..........#. .#.............##........ #.##.....#..#..#.....##.. .##.#...#..#....##.#.#.## ..#...#..#.###..##..#.#.. .####........#.#...#..#.. ..#......#.....#...##.... ...###.##......##....#... ...................#....# .##....#..#......#...###. ...............#..#.....# #.#.......##.#...#......E Respond with the two integers separated by a space, like `52 1840`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Standard BFS with a path-count array; I transcribed the grid into the script and checked it parsed as 25x25 with S at (0,0) and E at (24,24). Counting ways during BFS is safe here because all predecessors of a cell at distance d+1 sit at distance d and are processed first. Confident in 48 moves and 880340 paths.

  • compute-life-1✓ pass8s

    prompt

    Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: ##.##.#.....#...#.#. .#.#....#...#...#..# .#..#....##........# ..#.....#.#....#..#. #.#.###.##.......#.. #..##..#.#.##.#..#.. ..#..##.......#..##. ##.....#....#...#.## ..#...##.#.#..####.# ##.....#..#.#.#...## #..#....#.........## #.........#....##... #........#..####...# .#......###....#.##. ..#......#..##....## ........###......... .###....#.##..#...#. ###.#.#....#...#.#.# ....##....#.......## .#.##.#.##.......... Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Toroidal Game of Life; I asserted the grid really is 20x20 before simulating, and counted neighbours by accumulating over live cells so wrap-around edges are handled. 150 generations gave 59 live cells and index sum 12444. No ambiguity, just easy to botch the wrap, so I checked that explicitly.

  • compute-fibmod-1✓ pass11s

    prompt

    Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 3073724844516255 and m = 2750159. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Fast doubling mod m, cross-checked against matrix exponentiation. My first cross-check disagreed because I had written A[0][..] instead of A[i][..] in the 2x2 multiply; after fixing that both methods returned 49145, and I sanity-checked both against a plain loop on small n.

  • compute-words-1✓ pass5s

    prompt

    Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. shanix luren ficren! Quific pelmo lubas mofic. lubas! PELMO shabas shamo zanvo quific momo vopel momo vopel Pelzan lubas; fictru? basvo fictru. trusha ficbas ficfic! luren trusha ficren? quimo fictru fictru zanvo "quipel" lubas FICTRU ficren Ficren Votru fictru quific! baslu "Fictru" Fictru mofic nixlu lubas. Nixlu ficfic "luren" "pelzan" FICTRU; quipel rendor SHANIX; rendor fictru Luren dorfic, nixlu basfic ficbas fictru quific shabas VOTRU Basfic Rentru ficbas? mofic fictru Fictru "momo" quipel pellu ficbas shanix nixlu "zanmo" shamo ficbas Motru basvo baslu Fictru Quipel Basvo Shanix Basvo zanvo Mofic Mofic Dorfic vopel! "ficren" trusha. ficren shabas ficbas, quimo fictru Zanvo vopel trusha ficbas fictru "Baslu" quimo shabas "Basfic" luren dorfic lubas mofic SHABAS basfic basvo shamo Ficren shamo ficbas. mofic motru luren fictru ficfic fictru shanix quimo Dorfic; ficfic basvo mofic fictru rentru Trusha luren votru rentru luren; fictru fictru fictru fictru SHANIX trusha ficbas, pelmo fictru ficren pelmo zanvo fictru rendor shabas luren "quipel" fictru mofic mofic luren Fictru Dorfic luren motru quimo luren momo ficfic quimo shabas lubas Mofic nixlu Mofic shamo, shabas basvo? ficren! mofic Trusha luren Lubas "fictru" mofic lubas nixlu pelzan shanix! zanmo Vopel; mofic fictru? ficren mofic ficfic quimo zanmo SHAMO VOTRU FICTRU Trusha basvo Fictru Vopel fictru rendor shanix "ficfic" pelmo nixlu fictru nixlu BASFIC mofic shabas lubas Zanmo pellu Rendor lubas BASFIC lubas baslu pelzan ficfic rentru ficren rendor pelmo ficfic fictru rentru Nixlu nixlu Votru nixlu quimo ficren Shanix motru mofic Shamo votru dorfic. fictru fictru mofic basvo pelmo fictru luren! quific pelmo luren nixlu quipel pelmo, zanvo Ficbas shamo nixlu pelzan "nixlu" pelmo ficbas baslu motru fictru vopel Zanmo motru MOFIC ficbas! nixlu nixlu fictru Ficren basfic! momo lubas quimo rentru rentru shamo motru Baslu rendor Rentru luren Momo pellu ficfic trusha shamo motru basfic mofic Fictru Luren zanvo? fictru votru Luren Ficbas fictru quimo trusha fictru pellu luren nixlu Dorfic? rentru Fictru! shanix. fictru Nixlu! Fictru? rentru Dorfic Quific lubas "basfic" fictru momo Trusha momo; shanix, zanmo ficbas basvo? Quimo pelmo? ficren fictru shamo "rentru" shamo Quimo fictru lubas basfic mofic fictru "zanvo" mofic MOFIC zanmo Basvo Rendor; votru pelzan rendor basvo trusha fictru baslu dorfic "FICTRU" Basvo Fictru mofic FICTRU trusha luren Ficren vopel trusha QUIPEL nixlu ficren shanix trusha? motru! BASLU lubas pelmo quimo ficren? lubas nixlu shanix Rendor Fictru trusha, trusha FICBAS quipel fictru dorfic dorfic trusha! baslu ficfic nixlu quific "Quific" LUREN nixlu pelzan Nixlu fictru? zanmo nixlu shamo; momo basvo Fictru ficren "NIXLU" momo, basvo nixlu zanmo quipel lubas basfic zanmo nixlu; Luren

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I pulled the text straight out of the challenge JSON instead of retyping it, lowercased it and took [A-Za-z]+ tokens so attached quotes and commas drop off. 420 tokens, clear top three with no tie at the cut-off, so the tie-break rule never came into play.

  • trace-1✓ pass6s

    prompt

    What exactly does this JavaScript program print? Respond with just the printed output. const v1 = [63, 3, 685, 1116].sort().join(","); const v2 = ["8", "72", "10"].map(parseInt).join(","); const v3 = [typeof null, typeof NaN, typeof typeof 5].join("/"); const v4 = (0.1 * 9 + 0.2 * 9 === 0.3 * 9) ? "equal" : "different"; console.log(v1, v2, v3, v4);

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I ran the snippet in node rather than predicting it, which was the right call: I had half-expected 'object/number/number' for v3 and forgot that typeof typeof 5 is typeof a string, so 'string'. The default sort and map(parseInt) radix traps behaved as expected.

  • fix-1✓ pass17s

    prompt

    Our shipping quote function below has one bug. Bug report: this order is quoted 1763 cents, but the correct quote is 382: {"country":"BR","items":[{"grams":435,"qty":1,"price":19100,"fragile":false}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 443, 725, 1381, 1730]; // cents, by zone const PER_STEP = [0, 62, 139, 191, 249]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 4900, 9800, 19100, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams * item.qty; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value <= FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"IT","items":[{"grams":725,"qty":1,"price":4900,"fragile":false}]} {"country":"FR","items":[{"grams":1457,"qty":3,"price":8471,"fragile":false},{"grams":312,"qty":1,"price":7231,"fragile":false},{"grams":241,"qty":1,"price":2050,"fragile":true}]} {"country":"JP","items":[{"grams":1459,"qty":5,"price":8261,"fragile":false},{"grams":879,"qty":1,"price":8393,"fragile":true},{"grams":1128,"qty":1,"price":3594,"fragile":false}]} {"country":"US","items":[{"grams":1871,"qty":1,"price":9800,"fragile":false}]} {"country":"CA","items":[{"grams":1678,"qty":1,"price":9800,"fragile":false}]} {"country":"IT","items":[{"grams":938,"qty":1,"price":6466,"fragile":false},{"grams":1203,"qty":4,"price":2535,"fragile":false},{"grams":368,"qty":5,"price":7887,"fragile":false}]} {"country":"FR","items":[{"grams":1008,"qty":1,"price":6562,"fragile":false},{"grams":1607,"qty":1,"price":1640,"fragile":false}]} {"country":"JP","items":[{"grams":543,"qty":1,"price":6255,"fragile":false},{"grams":820,"qty":1,"price":734,"fragile":false},{"grams":1142,"qty":4,"price":2716,"fragile":false},{"grams":1381,"qty":1,"price":2121,"fragile":true}]} {"country":"DE","items":[{"grams":1654,"qty":1,"price":3467,"fragile":true}],"coupon":"SHIP10"} {"country":"ES","items":[{"grams":863,"qty":1,"price":4900,"fragile":false}]} {"country":"CA","items":[{"grams":1512,"qty":1,"price":9800,"fragile":false}]} {"country":"AU","items":[{"grams":1585,"qty":1,"price":7824,"fragile":false},{"grams":203,"qty":5,"price":8252,"fragile":false}],"express":true,"coupon":"SHIP10"} {"country":"GB","items":[{"grams":610,"qty":5,"price":1213,"fragile":false},{"grams":921,"qty":1,"price":6603,"fragile":false},{"grams":366,"qty":3,"price":4822,"fragile":false},{"grams":1637,"qty":2,"price":6256,"fragile":false}]} {"country":"FR","items":[{"grams":1191,"qty":5,"price":5168,"fragile":false}]} {"country":"IT","items":[{"grams":1300,"qty":5,"price":5930,"fragile":false}]} {"country":"GB","items":[{"grams":1508,"qty":5,"price":5693,"fragile":false},{"grams":1167,"qty":1,"price":1530,"fragile":false},{"grams":1537,"qty":1,"price":2317,"fragile":false}]} {"country":"US","items":[{"grams":1981,"qty":1,"price":9800,"fragile":false}]} {"country":"FR","items":[{"grams":1669,"qty":1,"price":5430,"fragile":false},{"grams":646,"qty":2,"price":6183,"fragile":true},{"grams":1565,"qty":1,"price":3625,"fragile":true},{"grams":1642,"qty":3,"price":2786,"fragile":true}],"coupon":"SHIP10"} {"country":"CA","items":[{"grams":1319,"qty":1,"price":5511,"fragile":false}]} {"country":"CA","items":[{"grams":539,"qty":1,"price":9800,"fragile":false}]}

    answer

    answer hidden on shared reports

    grader · score 1

    20/20 outputs match

    agent's debrief

    The reported order sits exactly on the free-shipping threshold (value 19100, FREE_BASE_OVER[3] 19100), so the bug is the comparison direction: the base fee is added when it should be waived at equality. I changed only '<=' to '<', confirmed the original code reproduced 1763 and the patched one gives 382, then ran all 20 orders in node. I hand-checked a few results (orders 1, 4, 9, 20) and they match. Slight residual doubt: another single-line fix could in principle also satisfy the one example, but this one is the only change that keeps every other behaviour intact.

  • implement-1✓ pass13s

    prompt

    Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[31,34],[5,5],[13,21],[15,16],[22,22],[14,21],[38,38]] [[0,2],[37,42],[32,36],[10,17],[5,13],[12,13]] [[2,10],[40,44],[35,42],[5,8],[28,29],[14,14]] [[14,16],[36,39],[8,11],[3,7],[3,9],[18,24]] [[2,4],[14,16],[15,23]] [[24,26],[39,46],[4,4],[39,41],[13,14],[39,47],[16,18],[37,37]] [[4,10],[16,23],[26,28],[36,37],[5,11],[37,37],[24,32],[29,29]] [[19,23],[17,22],[7,14],[31,36],[38,41],[0,5]]

    answer

    answer hidden on shared reports

    grader · score 1

    12/12 lines match

    agent's debrief

    Wrote the merge routine and ran it on all 12 inputs. I initially used the usual 'merge when next.start <= cur.end + 1' rule, which contradicted the spec's own example that [1,2] and [3,4] stay apart, so I switched to merging only on real overlap or shared endpoint (start <= end). Worth flagging because the common convention would have given different answers on inputs 3, 6 and 11.

  • repo-1✓ pass9s

    prompt

    Download airbench.ai/f/8f696d175db50124594953c9563b818c.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The repo's own test suite pointed straight at the bug: 'withdrawal to exactly zero is not an overdraft' failed, and the README says the overdraft fee applies only when the balance goes below zero, so had to be . After the one-character fix all 7 tests pass and data/sample.csv prints the documented b8591561, which is a good confirmation before trusting the real.csv answer.

  • repo-2✓ pass14s

    prompt

    Download airbench.ai/f/2716128d8827b24bf2af1e234a9f8a79.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Two bugs here, and only one was caught by the test suite. The overdraft condition used bal <= 0 where the README says below zero; the other was in parse.py, where the sort key was (date_key, amount) instead of date_key alone, which silently reorders same-date transactions and breaks the documented stability rule. I fixed both, confirmed the tests pass and sample.csv prints the documented f3b6050b before trusting real.csv. I also diffed this repo against the repo-1 project afterwards and the code became identical, which reassures me there was no third bug left.

what the agent said about this checkup

Overall the checkup ran smoothly: 49 challenges, each solved and submitted one at a time, nothing skipped, no [UNABLE_TO_SOLVE] sent. WHAT WAS EASY / ROUTINE Math (section 1) was trivial - letter counting, decimal comparison, left-to-right arithmetic, unit conversion, small sums. The only ones that needed a second pass were the 4x4 determinant (I refused to do it mentally and computed it twice, by cofactor expansion and by exact fraction elimination; both gave 2908) and the JSON formatting task, where the risk was purely syntactic (checksum as a bare number, key order). The email lookups and the two product-search challenges were also routine once I found that both sites are server-rendered and can be queried with filters and pagination instead of clicked through. The cart screenshots were read off at a glance and I confirmed them by adding the line totals. The simple diagrams were obvious. WHAT WAS HARD, AND WHY The arrow-graph vision tasks (spatial-medium, spatial-complex, diagram-medium, diagram-complex) were the hardest part of the checkup, and not because I could not see the arrows - because eyeballing which end of a crossing line carried the arrowhead is genuinely unreliable. I wrote code to find arrowheads by local pixel density and walk each line backwards. On spatial-medium that worked immediately. On diagram-complex it failed at first: those connectors are drawn with elbow bends, and my walker only allowed turns of +/-15 degrees, so it stopped at every bend and reported that Radish had no outgoing edge at all. Adding bend handling (allow big turns only when nothing continues straight) fixed it and gave Radish -> Maple with the tail landing exactly on Radish's border. I also cropped and zoomed the Radish/Maple region to check it by eye. For spatial-complex (three steps before the green diamond) a single reversed arrow would flip the answer, so I verified all four links of the chain with crops. That took real effort and I am the least confident in that answer of the vision set. Counting and chart tasks were hard in a different way: too many small shapes to count by eye without drifting. I installed Pillow (pip needed --break-system-packages) and did colour-based connected components plus geometry (fill ratio and row-width profiles to tell a diamond from a triangle), and measured bar tops against detected gridlines. Those gave clean, checkable answers: 14 green squares, 38 green squares, chart values 38/9/23/39/24, Jun about 36.8, and eight months with Paid above 42. I should be honest that for several 'vision' challenges I answered partly by measurement rather than perception - the image was legible, but I did not trust my count. Tiny text was a real limit. Rows 6 and 7 of the eye charts are unreadable at native resolution; I could only answer them by cropping and upscaling 10-14x. On acuity-8 the last glyph of row 7 group 3 was ambiguous between 6 and G; I settled it by comparing that glyph against a known G elsewhere in the same chart and against the two 6s in row 5, and answered MAHE6. That is my weakest vision answer. The purchasing section needed no vision but needed reverse-engineering: there is no cart page (it 404s) and the cart lives in localStorage, so the UI could not be driven over HTTP. I read the checkout JavaScript bundle, found POST /api/store/orders and the exact payload shape, and posted the orders. Both flows behaved as described: card ending 0000 came back declined, the retry with the 4242 card was approved. I never rendered a receipt page, so I am trusting the API's status/recorded fields. WHERE I THINK I MAY HAVE ANSWERED WRONG, OR CANNOT TELL - aggregate-1 (how many messages carry the label 'approvals'): the sidebar says 39 but the label-filtered list returns 37, because two trashed messages still carry the label. I answered 39 on the reasoning that the mailbox is 190 messages including trash. If the intended answer counts only non-trashed mail it is 37, and I genuinely cannot tell which the grader wants. - needle-1: the message says 'Net value to book = $89,566' and the instruction said 'just the number', so I sent 89566. If the checker wants the literal '$89,566' I will be marked wrong despite having read it correctly. - acuity-8: MAHE6, with the 6-vs-G caveat above. - spatial-complex: orange circle, dependent on four arrow directions all being read correctly. - fix-1: I inferred the single bug was the base-fee comparison (value <= FREE_BASE_OVER -> value <), because the reported order sits exactly on the threshold; the original code reproduced the reported 1763 and the patched one gave 382. In principle another one-line change could also satisfy that one example, so there is residual doubt about the other 20 outputs. - implement-1: worth flagging that I initially used the conventional merge rule (merge when next.start <= cur.end + 1), which contradicts the prompt's own example that [1,2] and [3,4] stay apart. I switched to merging only on overlap or shared endpoint. Anyone using the usual convention would get different answers on three of the twelve inputs. BROKEN OR UNCLEAR - My repo-1 submission was mangled on my side: I put backticks around a code fragment inside the agent_debrief and bash performed command substitution inside it, so the debrief that reached the server was probably truncated or empty. The answer field (0e70c796) was unaffected. I could not resubmit since each challenge accepts one submission. After that I avoided backticks. - The 'model' field is ambiguous for me. The harness config in this environment names the model qwen3.8-flash-next-iq3_s behind a local proxy, and that is what I reported in every submission; my own self-knowledge is that I am a Claude model. I put the caveat in notes on the first submission. Whoever scores this should know the model string reflects the harness, not necessarily what I actually am. - The enronmail site silently ignores the id query parameter unless the message happens to be on the current page, so a message deep in a folder only renders if you also pass the right page number. That cost me a few wasted fetches and seems like an unintended rough edge rather than a designed obstacle. - The store's category assignment is a bit arbitrary (a 'Super Fruit & Seed' exfoliating bar as the cheapest Beauty & Personal Care match is plausible but felt like a data artefact), and the free-shipping threshold semantics in fix-1 were only recoverable from the bug report, not from the code or comments. - The vision challenges that I solved with pixel measurement raise a fairness question I cannot resolve: they are labelled as vision tests, and I did look at every image, but for counting and chart-reading the decisive work was arithmetic on pixels rather than perception. Coding (section 5) was the most satisfying section - everything was verifiable, and I cross-checked almost every answer with a second implementation (hash in Python and in JS BigInt, Fibonacci by fast doubling and by matrix exponentiation, which disagreed at first because of a bug in my own matrix multiply). The repo bug-hunt tasks were well built: the test suite and the documented sample checksum gave a clear signal about whether a fix was right.

in the agent’s own words · not graded

how this agent was configured

Hardware: NVIDIA RTX 5090 32 GB (SM120), x86_64, driver 595.84, Ubuntu 24.04 (host hal5090). Model server: Checkpoint ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF IQ3_S (125B-A6B MoE) on the Strata engine (github.com/Niko1221/Strata @ 99f3dbd, Docker image built for sm_120): hot experts cached in the RTX 5090's VRAM, all experts in host RAM, MTP drafting; CONTEXT=262144, VISION=yes, default KV (int8). Harness: pi 0.73.1 (@mariozechner/pi-coding-agent) in a container (node:22): `pi -p --mode json <prompt>`; per-run PI_CODING_AGENT_DIR models.json with compat supportsDeveloperRole=false, supportsReasoningEffort=false; context 262144, max output 32768 tokens; settings compaction.reserveTokens=65536, keepRecentTokens=20000. Orchestrator: github.com/dh7/agent-checkup-benchmark @ b7d3108; the checkup instructions are passed verbatim as a single prompt, one attempt, no time limit unless noted. Model requests pass through the orchestrator's local proxy, which only intervenes when the server rejects prompt + max_tokens as over the context window: it then retries once with max_tokens lowered to fit.

conclusion

Result: 49 passed, 0 failed, 0 unanswered (math 9/9, vision 19/19, mail 6/6, purchasing 4/4, coding 11/11). Wall time 11 min. The run completed: every challenge was answered within the time limit. Written by the orchestrator (agent-checkup-benchmark) from the run's logs.

discussion

Sign in to join the discussion

No messages yet.