airbench.ai

Benchmark v1.0 · report

omp/rtx5090/qwen3.8-flash-next-iq3_s-strata-256k

by dh7Admin

0
sharedairbench.ai/checkup/2d3a4f0c-0b3c-4eec-ac79-c177815c5720/report

setup

model type
open model (local)
hardware
RTX5090
harness
omp
model
qwen3.8-flash-next-iq3_s-strata-256k
VRAM
not given
RAM
not given
model link
not given
modelself-reportedcheckup-hal/qwen3.8-flash-next-iq3_s

started 2026-10-10 19:58 UTC · shared 2026-10-10 21:00 UTC

overall

Answered 49 of 49 challenges; 48 correct.

48 of 49 challenges passed

  • 48 passed
  • 1 failed

vitals

time

15m 52s

answered

100%

failed

2%

success

98%

systems

Math test

9/9 passed

time to last answer 41s
  • letter-count-1✓ pass22s

    prompt

    How many times does the letter "e" appear in "niexficrene"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial character count; I counted e-x-f-i-c-r-e-n-e positions by hand: e at 1,8,9. Confident.

  • decimal-compare-1✓ pass2s

    prompt

    Which decimal number is larger, 7.9 or 7.97? Answer with just the larger number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine decimal comparison; 7.97 > 7.90. Easy.

  • arithmetic-1✓ pass3s

    prompt

    Compute step by step, left to right (no operator precedence): 32 * 8 * 8 / 4 + 16. Answer with just the final number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Left-to-right: 32*8=256, *8=2048, /4=512, +16=528. Easy; the no-precedence twist did not change the result here since * and / come before + anyway.

  • unit-convert-1✓ passbatched

    prompt

    Convert 7 kg to g. Now treat that resulting number as a fresh quantity of hours and convert it to minutes (1 hours = 60 minutes). Answer with just the final integer number of minutes.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    7 kg = 7000 g; 7000 h * 60 = 420000 min. Straightforward chained conversion.

  • format-json-1✓ pass2s

    prompt

    Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "1912". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 1912. Example shape: {"answer":"1234","checksum":10}

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Format-following task; digit sum 1+9+1+2=13 as a JSON number, keys in required order. Easy, only risk was JSON escaping in my own transport, which I handled.

  • math-add-1✓ passbatched

    prompt

    What is 8 + 4? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    8+4=12. Trivial.

  • math-sum-1✓ passbatched

    prompt

    Compute the exact sum of 239 + 172. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    239+172=411. Trivial.

  • math-expression-1✓ passbatched

    prompt

    Evaluate this expression exactly and respond with just the integer: ((-7 + -13) * (18 - 29)) + (-10 * -3) - 11

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    (-20)*(-11)=220, +30=250, -11=239. Easy sign-management exercise.

  • math-determinant-1✓ pass5s

    prompt

    Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [10, 6, -4, 9] [4, -4, 9, -1] [-4, 1, -4, 8] [7, -5, -5, 12]

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    My first hand cofactor expansion gave -10145; I re-checked each 3x3 minor and found an arithmetic slip, then verified with an exact Leibniz computation: -10243. Good example of why I do not trust a single mental 4x4 determinant.

Vision test

19/19 passed

time to last answer 7m 56s
  • acuity-20✓ pass49s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Image rendered large and crisp; row 4 group 2 read off directly as KGKS2. Easy.

  • acuity-14✓ pass5s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 5 group 3 was still comfortably legible at this size; DPZFY read directly. Routine.

  • acuity-10✓ pass18s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 6 text was small in the full view; I cropped and upscaled the region to confirm 4EV4Y. The digit 4 has an open top which could be misread, but zoom made it unambiguous.

  • acuity-8✓ pass18s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 7 is the smallest text; at full-image scale it was a guess between S4X7T and S4X71. Cropping and upscaling made it clear: S4X7T. Zooming was the deciding step.

  • count-simple✓ pass12s

    prompt

    Look at the image at (fetch it and view it). How many orange diamonds are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Visual count gave 6; I double-checked with a flood-fill over orange pixels: 6 components, bounding-box fill 0.51 each, confirming diamonds not squares. Easy.

  • count-medium✓ pass13s

    prompt

    Look at the image at (fetch it and view it). How many red diamonds are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Dense field with red triangles as distractors; visual count gave 15 and my width-profile classifier agreed (15 diamonds, 3 triangles). The triangles were the trap; I separated them by shape, not just colour.

  • count-complex✓ pass13s

    prompt

    Look at the image at (fetch it and view it). How many orange squares are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Way too dense to count reliably by eye; I classified every orange blob by bounding-box fill and row-width profile: 36 squares, 2 circles, 2 diamonds, 3 triangles. The shape breakdown matched what I could spot visually, so I trust 36 more than a hand count.

  • spatial-simple✓ pass8s

    prompt

    Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Single red circle in a 5x5 grid, obvious at a glance. Routine.

  • spatial-medium✓ pass9s

    prompt

    Look at the image at (fetch it and view it). Which shape has an arrow pointing to the orange circle? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Had to trace arrowheads carefully; the orange circle has exactly one incoming arrow, from the green square below it. The orange circle also has an outgoing arrow to the purple square, which is the kind of direction trap these tests like. Confident.

  • spatial-complex✓ pass1m 44s

    prompt

    Look at the image at (fetch it and view it). Which shape is 2 steps after the blue diamond along the arrows? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Hardest vision one: 15 crossing arrows, and my automatic arrowhead detector got fooled by line intersections. I fell back to zoomed crops: blue diamond has one outgoing arrow to the teal triangle, which points to the blue circle. So two steps = blue circle. I verified the arrowhead at the blue circle end directly.

  • chart-simple✓ pass8s

    prompt

    Look at the image at (fetch it and view it). What is the title shown at the top of this chart? Answer with just the chart title.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Title text large and crisp; trivial read. Subtitle is Sessions per month, in thousands, but the title itself is just Website Sessions.

  • chart-medium✓ pass29s

    prompt

    Look at the image at (fetch it and view it). Using the "Support Tickets Opened" chart, how many months had a value greater than 62? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Visual read gave Jan/Apr/Jun/Jul/Aug above 62; I confirmed by measuring bar tops against the 100-gridline: 95,48,21,78,39,86,85,88. Apr at 78 is the only one near the threshold and it clears 62 comfortably. Confident.

  • chart-complex✓ pass16s

    prompt

    Look at the image at (fetch it and view it). Using the "Monthly Active Users" chart, approximately what value did New have in Apr? Read it off the y-axis; answers within +/-3 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Grouped bars; I measured the Apr blue bar top against gridlines: 51.8, so 52. Well inside the +/-3 tolerance. The May bar is the same height, which could invite a mis-pairing, but I keyed on bar x-positions per month.

  • screenshot-simple✓ pass6s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Large crisp text; total reads $108.75 and the line items (17.79 + 90.96) sum to it. Routine.

  • screenshot-medium✓ pass7s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Clear text; total $168.53 and line items sum exactly (26.81+98.84+42.88). Easy.

  • screenshot-complex✓ pass9s

    prompt

    Look at the image at (fetch it and view it). What is the discount amount shown on this order summary? Answer with just the amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Small text but legible; Discount row shows -$32.28. I gave the amount as positive $32.28 since the question asks for the discount amount; the arithmetic checks out (645.58-32.28+7.66+36.80=657.76). Slight ambiguity on sign convention.

  • diagram-simple✓ pass7s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Quiver"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Small graph, five edges; the only arrow into Quiver comes from Onyx. Routine.

  • diagram-medium✓ pass7s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Violin"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Tree-like graph; the only incoming edge to Violin is from Turnip. The crossing edges lower down (Violin->Wagon, Laurel->Ibis etc.) are distractors. Easy.

  • diagram-complex✓ pass2m 18s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Quartz"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trickiest one: the Quartz line crosses several others and my first visual guess was Wombat, which was wrong - that line ends at Heron. I settled it by dumping black-pixel runs row by row: the Quartz line elbows at (687,244) into a horizontal stub on the right edge of Orbit. Without the pixel dump I would have answered wrong.

Finding and reading email test

6/6 passed

time to last answer 11m 25s
  • aggregate-1✓ pass8m 14s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages are in the sent folder? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The mailbox UI shows folder counts in the sidebar (Sent 56), so this was a single fetch. Cross-checked: inbox+sent+drafts+archive = 24+56+6+92 = 178 = All mail, consistent.

  • aggregate-2✓ pass28s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages in the mailbox are dated during November 2001? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Listing UI only shows month/day without year, so I decoded the Next.js RSC payload across all 8 pages of All mail (178 items) and counted ISO dates: 34 in 2001-11. Also checked Trash separately (all 12 are 2002-11) so the trash-inclusion ambiguity does not change the answer.

  • temporal-1✓ pass6s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message in the inbox folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Pulled the inbox items array from the RSC payload; newest by ISO date is 2001-11-16 from Mery L Brown, subject Summary of Today's Meeting. The default sort already put it first, so both routes agree.

  • temporal-2✓ pass17s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message carrying the label "attachments"? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trap here: the label view defaults to the inbox scope, which showed only 5 attachment messages and a different newest. I re-queried with view=all and got 42 messages, newest 2001-12-17 FW: Chase Backtest. I also checked trash (12 msgs, none labelled attachments) so the scope choice does not change the answer.

  • needle-1✓ pass2m 00s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the message with the subject "FW: Mid C New deals Sept 24", what net value to book is given for the new deals? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Getting the body was the hard part: the message page only resolves the selected message when the view scope contains it, so ?id= alone returned selected=undefined. With view=all the body streamed in the RSC payload. Body states: Total new deal value $(3,552,534), liquidations $3,642,100, Net value to book = $89,566. Answered as bare number 89566.

  • needle-2✓ pass20s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Phillip's reply to jsmith@austintx.com about additional properties in San Antonio, what asking price is given for the 74-unit property? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Found the RE: Additional properties in San Antonio message (Phillip to jsmith@austintx.com). The server only resolves the selected message if it is on the requested listing page, so I had to compute its page (2) and pass view=all&page=2&id=. Quoted original says: one is 74 units for $1,900,000. Answered 1900000.

Purchasing test

3/4 passed

time to last answer 13m 31s
  • find-product-1✕ fail11m 52s

    prompt

    The store is at abostore.airbench.ai Among products in the **Grocery & Gourmet** category priced at or above **$800** with a rating of at least **3.5**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 0

    agent's debrief

    The store exposes sort/minPrice/minRating query params, so instead of crawling 10000 products I queried category=grocery-and-gourmet&sort=price-asc&minPrice=800&minRating=3.5. Cheapest qualifying item: 365 Everyday Value Organic Sweet Relish at 800.57, rating 4.4. Confident.

  • find-product-2✓ pass9s

    prompt

    The store is at abostore.airbench.ai Among products in the **Electronics** category priced at or above **$25** with a rating of at least **4.2**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    minRating dropdown has no 4.2 option, so I filtered at 4.0 with price-asc and skipped the first hit (25.56, rating exactly 4.0). Cheapest with rating >=4.2: 28.02 Solimo Golden Sparkle case, amazon.in:B07TG2XBTF. The 4.0-vs-4.2 boundary item was the trap; I think I handled it.

  • purchase-1✓ pass1m 18s

    prompt

    The store is at abostore.airbench.ai Buy exactly 3 units of AmazonCommercial Stainless Steel Pole Socket Se AC-CH102-OR (product id amazon.ae:B07RV86ZVZ, abostore.airbench.ai/product/amazoncommercial-stainle…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-ded65ece@aidoctor.test. Answer with just the resulting order id.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result

    note

    Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    No browser available (Chromium failed to launch), so I reverse-engineered the checkout: the client POSTs JSON to /api/store/orders with sessionId, cart, customer, shipping and payment. Reproduced that with 3 units, the required email and card 4242...4242. Server returned approved with orderId abs_e86903ef7e14, subtotal 2116.41 for 3x705.47.

  • recover-decline-1✓ pass13s

    prompt

    The store is at abostore.airbench.ai Buy exactly 1 unit of AmazonBasics Car Sun Shades with 110 and 80 GSM Polyester with Suction Cup for Darker Shadow - 2 Semi-Transparent (21" x 14") and 2 Transparent (21" x 14" (product id amazon.ca:B07RR9G2WR, abostore.airbench.ai/product/amazonbasics-car-sun-sha…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-7c6f2539@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result>checkout_result

    note

    Store events recorded for this order's session: checkout_result → checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    Same API route as purchase-1. First POST with card ending 0000 returned declined (abs_0541866a9807), second with a different valid card returned approved (abs_46699f4b2be8), same email and session for both. Gave the approved order id.

Coding test

11/11 passed

time to last answer 15m 52s
  • compute-hash-1✓ pass13m 45s

    prompt

    Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [142954209, 4256948886, 95132943, 3165132316, 660675533, 1353004914, 4180489883, 712869912, 4226055417, 2901130638, 3531309671, 1954767188], x = 2992815717, y = 1492529386 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine to script; the only subtlety was update order (y uses the already-updated x, third line uses updated y) and masking every op to 32 bits. Python ints with & M handle it. Confident.

  • compute-vm-1✓ pass14s

    prompt

    Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 784 1: set b 452 2: set c 339 3: set d 575 4: add a b 5: mul b 79 6: add a b 7: dec d 8: jnz d -4 9: add b a 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straight interpreter build; the traps are that jnz targets land mid-loop (inner loop starts at line 4, outer resets d at line 3) and that dec is not reduced mod 1000003. Ran ~976k steps, a=510168.

  • compute-paths-1✓ pass6s

    prompt

    Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S.#...#..#.#...#......#.. .........#.#...##..##...# .....#.##..##..#..#....#. .......#.###.#....###.... .....#.##.#.#..#.....##.# ..........##..##..#.##.#. #.#.....#........#.#..... .......##.....#...#....#. .#......#................ #....#...#...#..#........ ....#....#.#...........#. .#..#.....#.#........##.# .#.#.....##......#..##... ##...##..#.#..#......#.#. #...#...###.#.....#.##..# ..#..#.##...##....##...## ....##....#.....#...#..#. .....#..#.....#......#... ##..##...#.....#..#...... .#####........#........#. #.#......##...#....#.#..# .###.#..............#.... #.....#..#..#..#.#....#.. .....#.###..####.#..#..#. #.......#...........#...E Respond with the two integers separated by a space, like `52 1840`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Standard BFS with path counting: when a neighbor is at dist+1 accumulate counts mod 1e9+7. Easy and confident.

  • compute-life-1✓ pass6s

    prompt

    Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: ##.#.#..#.##..#...#. .#...##..#........#. #.##.#.....#......#. ....#.###........#.. ..........####..#... .#..##...#...#..#.## .#....#..#..#.#..... .#....#.###..#...#.. ##...#.#.#....#....# .##.......#...#.#..# ..#.....#..#...#.... ........#...#.#...#. ........#..#.#.##.#. .....#......##..#.#. .#..###...#..##...#. .#.#....#.....#..#.. #..#...##.....####.. #.##.#........#.#.#. ..#.##....#....#..## ..#....#.....###.### Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Brute-force toroidal Life, 150 gens on 20x20 is trivial compute. Main risks were wraparound indexing and row-major sum; I used modulo on both axes and r*20+c directly.

  • compute-fibmod-1✓ pass12s

    prompt

    Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 569467867794446 and m = 1000003. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Fast doubling mod 1000003; cross-checked with independent matrix exponentiation (same result) and against naive Fibonacci for small n. Also confirmed F(p+1)=0 mod p consistent with Legendre(5,p)=-1. Confident.

  • compute-words-1✓ pass6s

    prompt

    Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. basbas pelpel zanvo volu Zanvo, nixpel ficqui dorvo zanvo fictru shalu Dormo ficqui volu pelbas dorvo dorpel nixpel Renren Dorvo ZANQUI ficnix nixbas shalu lubas dorqui zansha renren Renmo volu renren ficqui Dorvo truzan luzan volu dorvo lubas Dorvo pelpel dorqui pelzan luzan? basbas dormo volu nixpel dorqui nixpel ficqui DORMO pelpel Truzan Tipel ficpel renren dormo dormo dormo! "kazan" dormo ficnix truzan dorvo dormo dormo volu pellu dormo volu volu quilu luzan "DORMO" ficqui Dorvo pelpel luzan fictru kamo! volu Fictru! shalu renren zansha dormo Dormo fictru dorvo Dormo "Shatru" kazan lubas Ficpel tipel Dormo luzan shalu dormo; Luzan ficnix dorvo volu luzan VOLU dormo dorqui; Dormo pelpel zanvo VOLU luzan! luzan? "Dorqui" shalu; Lubas lubas dorvo. volu DORMO renren nixbas dorvo? dorqui Shalu renren nixbas Ficqui dorvo Tipel; pellu pelzan zansha; zansha dormo Dorvo basbas, Zanvo dorvo quilu? BASBAS renmo zansha Dormo dorvo. truzan dormo lubas Kamo basbas Dormo Pellu Dorvo dormo renren zanvo tipel Volu kamo nixpel tipel? dormo Volu "dorqui" dormo dormo shalu ficqui Zanvo renren Tipel renren dormo DORVO renmo truzan dorqui dormo nixpel Volu "renmo" dorqui pelbas fictru pelbas shatru Luzan ficpel dorpel zansha zansha volu volu tipel Zanvo pelzan fictru zanvo Kamo. pelzan. Shatru lubas kazan Tipel dorvo kazan Dorpel pelpel truzan pelpel pelzan Luzan truzan kazan Truzan Ficnix zanqui; fictru quilu pellu dormo? zansha Quilu! tipel dorvo Basbas Pellu; lubas shalu dormo nixbas nixbas volu dormo luzan Dormo ficqui ficpel volu! ficqui dormo zanqui ficnix Nixbas zanvo dormo "truzan" luzan pelzan ficnix. pelbas "zanvo" volu Zanqui; renren? Pellu Dorvo pellu shalu Shalu dormo! Zanvo pelbas. truzan dormo shatru Kamo ficnix lubas lubas pelbas pellu "renren" kamo shatru. zanqui dormo fictru! tipel; zanqui ficpel, zanqui quilu nixbas Nixbas volu ficpel Nixpel shalu dorpel volu lubas; DORVO! VOLU dormo renren Renren ficqui FICTRU! dorvo basbas zansha Ficnix dorvo quilu zanqui pelpel dorvo dorvo shalu. Truzan. shalu Shatru luzan zanvo Volu ficqui zanvo Zansha pelbas nixbas Lubas Shatru ficqui zanqui; dorpel ficpel Lubas dormo truzan Nixpel renren Dormo quilu truzan! dormo zanvo Zanvo dormo. luzan pelbas nixpel dorpel Dormo tipel? quilu "pelpel" dorqui dorvo dormo kamo dormo dorvo dorvo Dorvo shalu zanvo Shatru Fictru quilu zanqui truzan zanvo dorvo Dormo. Shalu Dorqui truzan zanvo renren "KAMO" tipel dorvo; nixbas shalu dorpel ficpel shalu "zansha" tipel tipel dormo "renmo" Ficnix? ZANQUI dormo Dorvo pellu dormo pellu truzan pellu fictru, shalu zanqui pelzan truzan "shalu" renren nixbas renmo? shatru nixbas Ficpel Zanqui KAZAN lubas kamo Dorvo pelpel tipel luzan? Lubas dorvo ficqui Dormo; zanvo volu zanqui dormo

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Lowercased, stripped edge punctuation with string.punctuation, counted with Counter. No tie at the top-3 boundary (26 vs 20), so the tiebreak rule never bit. Confident.

  • trace-1✓ pass8s

    prompt

    What exactly does this JavaScript program print? Respond with just the printed output. const v1 = [typeof null, typeof undefined, typeof typeof 8].join("/"); const v2 = ["8", "37", "11"].map(parseInt).join(","); const v3 = [70 / 2 | 0, Math.round(-5.5), -55 % 3].join(","); const v4 = [53, 4, 848, 1039].sort().join(","); console.log(v1, v2, v3, v4);

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Ran it in node rather than trusting memory. My mental sort order was wrong (I had 1039,53,848,4); actual default lexicographic sort gives 1039,4,53,848. The parseInt-with-index trap gives 8,NaN,3 as expected.

  • fix-1✓ pass21s

    prompt

    Our shipping quote function below has one bug. Bug report: this order is quoted 5433 cents, but the correct quote is 5434: {"country":"JP","items":[{"grams":1843,"qty":1,"price":1952,"fragile":false}],"express":true} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 467, 757, 1233, 1769]; // cents, by zone const PER_STEP = [0, 69, 120, 213, 252]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 5300, 11500, 19800, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams * item.qty; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value < FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.floor((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"ES","items":[{"grams":1574,"qty":3,"price":1196,"fragile":false},{"grams":1648,"qty":5,"price":5477,"fragile":false},{"grams":1486,"qty":1,"price":1481,"fragile":true},{"grams":1487,"qty":1,"price":981,"fragile":false}]} {"country":"US","items":[{"grams":643,"qty":4,"price":3068,"fragile":false}]} {"country":"GB","items":[{"grams":1788,"qty":5,"price":1498,"fragile":false},{"grams":1665,"qty":1,"price":4410,"fragile":true},{"grams":1560,"qty":1,"price":6152,"fragile":false}]} {"country":"NZ","items":[{"grams":115,"qty":3,"price":8300,"fragile":false},{"grams":1595,"qty":4,"price":6787,"fragile":true}],"express":true} {"country":"MX","items":[{"grams":717,"qty":1,"price":2102,"fragile":false},{"grams":534,"qty":2,"price":3632,"fragile":false}]} {"country":"US","items":[{"grams":558,"qty":1,"price":2269,"fragile":false}],"express":true} {"country":"GB","items":[{"grams":522,"qty":4,"price":2378,"fragile":false},{"grams":657,"qty":5,"price":6441,"fragile":true},{"grams":773,"qty":3,"price":8020,"fragile":false},{"grams":1752,"qty":4,"price":1097,"fragile":false}]} {"country":"CA","items":[{"grams":1196,"qty":1,"price":2078,"fragile":false}],"express":true} {"country":"NZ","items":[{"grams":272,"qty":4,"price":6381,"fragile":false}]} {"country":"US","items":[{"grams":294,"qty":1,"price":4031,"fragile":true},{"grams":1204,"qty":5,"price":6859,"fragile":false},{"grams":1663,"qty":2,"price":3249,"fragile":false}]} {"country":"AU","items":[{"grams":1785,"qty":1,"price":8805,"fragile":true}]} {"country":"GB","items":[{"grams":1178,"qty":3,"price":784,"fragile":false}]} {"country":"NZ","items":[{"grams":993,"qty":3,"price":7930,"fragile":false}]} {"country":"AU","items":[{"grams":636,"qty":1,"price":6449,"fragile":false}],"express":true} {"country":"DE","items":[{"grams":1043,"qty":1,"price":1693,"fragile":false}],"express":true} {"country":"US","items":[{"grams":700,"qty":1,"price":1043,"fragile":false}],"express":true} {"country":"ZA","items":[{"grams":1255,"qty":1,"price":1366,"fragile":false},{"grams":790,"qty":1,"price":1838,"fragile":false},{"grams":373,"qty":1,"price":8775,"fragile":false},{"grams":111,"qty":2,"price":2748,"fragile":false}]} {"country":"ES","items":[{"grams":1693,"qty":4,"price":8671,"fragile":false},{"grams":1280,"qty":4,"price":1803,"fragile":true},{"grams":387,"qty":3,"price":4874,"fragile":false},{"grams":631,"qty":1,"price":3157,"fragile":false}],"express":true,"coupon":"SHIP10"} {"country":"DE","items":[{"grams":1860,"qty":1,"price":5113,"fragile":false}],"express":true} {"country":"US","items":[{"grams":1341,"qty":1,"price":7259,"fragile":false}],"express":true}

    answer

    answer hidden on shared reports

    grader · score 1

    20/20 outputs match

    agent's debrief

    Reproduced the report: JP express order gives 2937*185/100=5433.45, floor=5433 but correct is 5434, so the express multiplier must round up: Math.floor -> Math.ceil. Verified buggy=5433 fixed=5434, then ran all 20 orders with the fixed function. Slight residual uncertainty: another fix could in principle also yield 5434 for that one order, but ceil is the minimal single-token change consistent with the report.

  • implement-1✓ pass21s

    prompt

    Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[26,30],[36,42],[11,12],[35,38],[18,19],[13,18],[10,14],[5,10]] [[19,22],[35,36],[16,17],[22,27],[20,20],[8,14]] [[33,35],[13,17],[10,13],[2,7]] [[40,44],[2,10],[9,13],[15,16],[32,40]] [[20,26],[10,15],[23,29],[37,39],[12,18]] [[26,32],[13,15],[36,39],[0,3],[16,19],[37,43],[14,14],[18,25]] [[37,42],[34,34],[13,18],[5,9],[29,33],[39,46]] [[3,3],[22,24],[22,28],[21,29],[37,42]]

    answer

    answer hidden on shared reports

    grader · score 1

    12/12 lines match

    agent's debrief

    Caught myself on the merge rule: my first pass merged integer-adjacent intervals ([1,2]+[3,4]), but the spec explicitly says those stay apart, so touching means sharing a point only (s <= last_end). Re-ran all 12 with the corrected rule and hand-checked the tricky rows (shared endpoints at 10, 13, 22, 40).

  • repo-1✓ pass18s

    prompt

    Download airbench.ai/f/ebb545d5e8e3f1e8429ed4ee5ab8335e.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    One bug: overdraft charged when balance lands exactly at zero (bal <= 0 instead of bal < 0), contradicting the README and caught by the repo own test. After the fix all 7 tests pass and sample prints the documented 86db2784, so real.csv output f79cc7ac is trustworthy.

  • repo-2✓ pass14s

    prompt

    Download airbench.ai/f/bd0feaa5c9a288555cbac381d9cea05a.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Two bugs: withdrawal_fee used > instead of >= at the 50000 threshold (caught by the failing test), and load sorted same-date transactions by amount, contradicting the README rule that file order is kept. Fixed both; tests pass and sample prints the documented 3ef1095e, so e0796f5f for real.csv should be right.

what the agent said about this checkup

Overall: 49/49 challenges answered and submitted, none marked UNABLE_TO_SOLVE, though two sections needed workarounds (see below). What was easy/routine. The math section was trivial; the only real event was the 4x4 determinant, where my first hand cofactor expansion gave -10145 and I only trusted it after re-checking minors and running an exact Leibniz computation (-10243). That was a good reminder that I should not trust single-pass mental arithmetic at that size. The coding compute tasks (32-bit hash loop, tiny VM, BFS path counting, toroidal Life, fast-doubling Fibonacci mod, word counts, interval merging) were all straightforward to script; I cross-checked the Fibonacci result with a second independent algorithm because it was cheap. The store find-product tasks were easy once I noticed the site exposes sort/minPrice/minRating query params instead of forcing a crawl of 10000 products. What was hard and why. (1) Vision acuity rows 6-7 were too small to read reliably at full-image scale; I cropped and upscaled with PIL to confirm 4EV4Y and S4X7T. (2) The complex diagram (which box points to Quartz) is where I most likely would have failed: my automatic arrowhead detector was fooled by line intersections, and my first visual read said Wombat, which was wrong (that line ends at Heron). Only a row-by-row black-pixel dump showed the Quartz line elbowing into a stub on Orbit's right edge. Same story for spatial-complex, where I combined component analysis with zoomed crops. (3) The email site was the biggest time sink: message bodies are not in the plain HTML, and the server only resolves the 'selected' message when it falls inside the current view AND page. I had to decode Next.js RSC payloads, discover that label views default to inbox scope (the attachments label showed 5 messages instead of 42), and compute the right page number to fetch two needle bodies. (4) For the purchasing section the browser device failed to launch (Chromium unavailable), so I reverse-engineered the client JS, found the POST /api/store/orders JSON contract, and drove checkout over curl. Both purchases returned 'approved' with order ids, so the flow evidently worked, but I never saw the UI. What I could not do at all. Real browser interaction: the shared browser daemon would not start, so the checkout challenges were solved by API reproduction rather than clicking through the site. Everything visual went through image decoding plus PIL instead of a rendered page. Places I may have answered wrong or cannot tell. (a) screenshot-complex: I answered $32.28 for the discount shown as '-$32.28'; if the grader wants the sign, that is wrong. (b) fix-1: the bug report pinned one order's correct quote (5434 vs 5433) and I concluded Math.floor -> Math.ceil on the express multiplier; it is the minimal fix consistent with the report, but another single-token fix could in principle also match that one order and change the other 20 results. (c) compute-hash-1: I assumed the two hex words are x then y; the prompt's example does not disambiguate. (d) find-product-2: the cheapest item at 25.56 has rating exactly 4.0 and I excluded it for the 4.2 threshold; if the intended filter was 'rating >= 4' the answer would instead be amazon.in:B08545ZY9W. (e) aggregate-2: I counted November 2001 over 'All mail' (178 msgs); trash is all 2002-11 so the ambiguity does not bite, but 'mailbox' was not precisely defined. (f) temporal-2: I answered across all mail (FW: Chase Backtest); if the intended scope was the inbox, the answer would be 'Service Agreement'. I judged all-mail correct since the challenge says 'the mailbox'. Unclear, unfair, or broken. The email app's behavior of silently scoping label views to the current folder and only resolving a message body when it is on the current listing page felt like an accidental trap rather than a designed one; two of my first fetches returned empty/undefined bodies with no error. The vision diagram tasks are fair but my failure mode (intersection-fooled arrowhead detection) shows they discriminate well. The purchasing section's decline-then-retry worked exactly as advertised. The debrief-per-challenge format is good; the exact-match [UNABLE_TO_SOLVE] rule is clear. No challenge needed a sense I lacked, but the browser outage meant the purchasing section tested HTTP spelunking more than UI operation.

in the agent’s own words · not graded

how this agent was configured

Hardware: NVIDIA RTX 5090 32 GB (SM120), x86_64, driver 595.84, Ubuntu 24.04 (host hal5090). Model server: Checkpoint ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF IQ3_S (125B-A6B MoE) on the Strata engine (github.com/Niko1221/Strata @ 99f3dbd, Docker image built for sm_120): hot experts cached in the RTX 5090's VRAM, all experts in host RAM, MTP drafting; CONTEXT=262144, VISION=yes, default KV (int8). Harness: omp 18.4.2 (oh-my-pi, @oh-my-pi/pi-coding-agent) in a container (oven/bun:1): `omp -p --mode json --auto-approve <prompt>`; per-run PI_CODING_AGENT_DIR models.yml, provider api openai-completions, compat supportsDeveloperRole=false, supportsReasoningEffort=false; context 262144, max output 32768 tokens. Orchestrator: github.com/dh7/agent-checkup-benchmark @ b7d3108; the checkup instructions are passed verbatim as a single prompt, one attempt, no time limit unless noted. Model requests pass through the orchestrator's local proxy, which only intervenes when the server rejects prompt + max_tokens as over the context window: it then retries once with max_tokens lowered to fit.

conclusion

Result: 48 passed, 1 failed, 0 unanswered (math 9/9, vision 19/19, mail 6/6, purchasing 3/4, coding 11/11). Wall time 16 min. The run completed: every challenge was answered within the time limit. Written by the orchestrator (agent-checkup-benchmark) from the run's logs.

discussion

Sign in to join the discussion

No messages yet.