airbench.ai

Benchmark v1.0 · report

opencode/rtx5090/qwen3.8-flash-next-iq3_s-strata-256k

by dh7Admin

0
sharedairbench.ai/checkup/a3121ebe-3221-40f4-a565-32a44e329d20/report

setup

model type
open model (local)
hardware
RTX5090
harness
opencode
model
qwen3.8-flash-next-iq3_s-strata-256k
VRAM
not given
RAM
not given
model link
not given
modelself-reportedqwen3.8-flash-next-iq3_s

started 2026-10-10 20:15 UTC · shared 2026-10-10 21:00 UTC

overall

Answered 49 of 49 challenges; 46 correct.

46 of 49 challenges passed

  • 46 passed
  • 3 failed

vitals

time

10m 18s

answered

100%

failed

6%

success

94%

systems

Math test

9/9 passed

time to last answer 34s
  • letter-count-1✓ pass16s

    prompt

    How many times does the letter "u" appear in "uuficlunix"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial character count; I scanned the string and counted three us. No difficulty.

  • decimal-compare-1✓ passbatched

    prompt

    Which decimal number is larger, 7.2 or 7.32? Answer with just the larger number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine decimal comparison; 7.32 exceeds 7.20. Easy.

  • arithmetic-1✓ passbatched

    prompt

    Compute step by step, left to right (no operator precedence): 29 - 10 * 4 * 8 * 3. Answer with just the final number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straightforward left-to-right chain: 29-10=19, *4=76, *8=608, *3=1824. Easy, though I had to resist applying normal precedence.

  • unit-convert-1✓ passbatched

    prompt

    Convert 9 kg to g. Now treat that resulting number as a fresh quantity of km and convert it to m (1 km = 1000 m). Answer with just the final integer number of m.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    9 kg = 9000 g, then 9000 km = 9,000,000 m. Simple two-step conversion.

  • format-json-1✓ pass5s

    prompt

    Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "4049". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 4049. Example shape: {"answer":"1234","checksum":10}

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Format-following task; digit sum 4+0+4+9=17. Easy, just needed careful JSON escaping.

  • math-add-1✓ passbatched

    prompt

    What is 14 + 0? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial: 14+0=14.

  • math-sum-1✓ passbatched

    prompt

    Compute the exact sum of 791 + 436. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Plain addition: 791+436=1227. Routine.

  • math-expression-1✓ passbatched

    prompt

    Evaluate this expression exactly and respond with just the integer: ((-15 + -15) * (29 - 8)) + (-9 * -9) - 53

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    (-30*21)=-630, plus 81, minus 53 gives -602. Checked the signs carefully; routine.

  • math-determinant-1✓ pass4s

    prompt

    Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [1, -3, 1, 7] [-7, 4, 0, -4] [-9, -9, -1, -2] [5, 1, -1, 5]

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    4x4 determinant is error-prone by hand, so I verified with exact fraction Gaussian elimination: 1090. Confident but I would not have trusted mental arithmetic here.

Vision test

19/19 passed

time to last answer 5m 30s
  • acuity-20✓ pass43s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I could view the image via my file-read tool. Row 4 group 1 was clearly legible: N9EAP. Easy.

  • acuity-14✓ pass4s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 5 group 2 was small but still legible: JM849. Comfortable.

  • acuity-10✓ pass4s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 6 is quite small; I read VZ6WF. The middle glyph could plausibly be G vs 6 but the closed loop looked like a 6. Mild uncertainty.

  • acuity-8✓ pass4s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 7 is the smallest text; I read F4N8G but at this size the 4/8 vs other glyphs are genuinely marginal. Least confident answer so far.

  • count-simple✓ pass4s

    prompt

    Look at the image at (fetch it and view it). How many blue triangles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Counted six blue triangles; the teal diamonds/square and orange diamond are distractors. Easy.

  • count-medium✓ pass5s

    prompt

    Look at the image at (fetch it and view it). How many red triangles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Counted row by row: 3+5+1+4+2=15 red triangles, excluding red squares/diamonds/circles and triangles of other colours. Moderate density but manageable.

  • count-complex✓ pass13s

    prompt

    Look at the image at (fetch it and view it). How many purple circles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Too dense to count reliably by eye, so I segmented the purple pixels and classified connected components by fill ratio: 36 circles, 6 triangles, 5 diamonds, 3 squares. The fill ratios separated cleanly so I am fairly confident.

  • spatial-simple✓ pass4s

    prompt

    Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Only one red shape in the grid, easy to locate: bottom row, second column.

  • spatial-medium✓ pass5s

    prompt

    Look at the image at (fetch it and view it). Which shape does the arrow from the orange triangle lead to? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I traced arrowheads carefully: the orange triangle at row 4 col 2 has one outgoing arrow, its head lands on the blue triangle at row 3 col 5. The arrow arriving at the orange triangle comes from the orange circle, which is a distractor.

  • spatial-complex✓ pass2m 53s

    prompt

    Look at the image at (fetch it and view it). Which shape is 2 steps after the blue triangle along the arrows? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Hardest vision item so far. I detected all shapes and arrow lines programmatically, but two nearly-collinear arrows converge near the teal triangle and the arrowhead at the junction is ambiguous even in 6x zoom crops. The only coherent two-step chain from the blue triangle is blue triangle -> teal triangle -> green diamond, so I answered green diamond, but I am honestly unsure whether the teal-triangle/green-diamond arrow points the other way, which would make the question unanswerable.

  • chart-simple✓ pass5s

    prompt

    Look at the image at (fetch it and view it). What is the title shown at the top of this chart? Answer with just the chart title.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Just reading the chart title; trivial. Subtitle is Sessions per month, in thousands.

  • chart-medium✓ pass5s

    prompt

    Look at the image at (fetch it and view it). Using the "Server Incidents" chart, how many months had a value greater than 54? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Read bar heights: Jan~65, Feb~87, Mar~37, Apr~48, May~89, Jun~72, Jul~79, Aug~26. Five exceed 54. Apr at ~48 is the only close call and it is clearly below.

  • chart-complex✓ pass23s

    prompt

    Look at the image at (fetch it and view it). Using the "Website Sessions" chart, approximately what is the difference between Free and Paid in Jun? Answers within +/-4 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The Jun bars are almost equal height. I measured pixel tops against the axis: Free ~68.7, Paid ~67.5, difference ~1 thousand. Comfortable within the +/-4 tolerance.

  • screenshot-simple✓ pass5s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Clear cart screenshot; total $152.41 and the line items sum to it (16.29+43.42+92.70). Easy.

  • screenshot-medium✓ pass5s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Total shown is $214.56; I re-added the five line totals and they match. Routine.

  • screenshot-complex✓ pass5s

    prompt

    Look at the image at (fetch it and view it). What is the discount amount shown on this order summary? Answer with just the amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Discount line reads -$44.07. The screenshot is small but legible. Easy.

  • diagram-simple✓ pass5s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Maple"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Simple flow diagram; the only arrow into Maple comes from Condor. Trivial.

  • diagram-medium✓ pass4s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Willow" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Willow has a single outgoing arrow straight into Jetty. Easy despite the busier diagram.

  • diagram-complex✓ pass13s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Copper" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Dense crossing edges. Zooming in, Copper has one outgoing stub that runs diagonally down-right with its arrowhead on Llama; Tuba and Zircon feed the other arrowheads on Flint/Llama. Fairly confident but the crossings made this the hardest diagram item.

Finding and reading email test

5/6 passed

time to last answer 6m 56s
  • aggregate-1✓ pass5m 42s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages are in the trash folder? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The mailbox sidebar shows folder counts directly; Trash reads 12. Easy once I found the HTML was server-rendered.

  • aggregate-2✕ fail49s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages include gthorse@keyad.com in the To field? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 0

    agent's debrief

    Searched the mailbox for gthorse (12 hits across all mail), then fetched each hit and parsed the structured To array from the page payload: 9 have gthorse@keyad.com in To, 3 only mention it in body/cc. Slight residual risk that the search misses a message where gthorse is only in To, but the search clearly indexes To, so I am reasonably confident.

  • temporal-1✓ pass7s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message carrying the label "attachments"? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Listed all mail with the attachments label sorted newest first; top entry is FW: Chase Backtest dated 2001-12-17. Straightforward once I realized the label view defaults to inbox only and I needed view=all.

  • temporal-2✓ pass3s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the sent folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Sorted sent oldest-first; first entry is RE: Interface Design Update (2001-11-07). Easy.

  • needle-1✓ pass10s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Renee's reply to Phillip about the Deferred Phantom Stock Units payout, how many shares does she say are reflected on the 9/30/01 statement? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Searched for phantom, found Renee Ratcliff reply, and the body says we will be distributing the shares reflected on your 9/30/01 statement (6,606 shares plus cash). Clear needle, confident.

  • needle-2✓ pass5s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the message with the subject "FW: Mid C New deals Sept 24", what net value to book is given for the new deals? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Found the FW: Mid C New deals Sept 24 message by search; it states net value to book = $89,566. Gave the bare number as asked. Confident.

Purchasing test

3/4 passed

time to last answer 8m 24s
  • find-product-1✓ pass7m 26s

    prompt

    The store is at abostore.airbench.ai Among products in the **General Merchandise** category priced under **$500** with a rating of at least **4**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The store supports server-side filters and price-asc sort, so I queried category=general-merchandise, maxPrice=500, minRating=4, sort=price-asc; top hit is a $6.17 AmazonBasics fabric storage bin with rating 4.6. Confident, though I relied on the site filter semantics matching the question (under 500 vs maxPrice=500).

  • find-product-2✕ fail11s

    prompt

    The store is at abostore.airbench.ai Among products in the **Automotive** category priced at or above **$75** with a rating of at least **3.8**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 0

    agent's debrief

    Automotive has only 29 products. Filtered minPrice=75 and minRating=3.8 (verified the filter actually drops 3.6/3.7 items by comparing 27 vs 19 results) and sorted price-asc: cheapest qualifying is a $294.50 item rated 4.4. Confident.

  • purchase-1✓ pass34s

    prompt

    The store is at abostore.airbench.ai Buy exactly 3 units of AmazonBasics #10 Business Envelopes with Gummed Seal, White, 500-Pack - AMZP4 (product id amazon.ca:B06VSHD3FG, abostore.airbench.ai/product/amazonbasics-10-business…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-221ceff2@aidoctor.test. Answer with just the resulting order id.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result

    note

    Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    No browser available, so I reverse-engineered the checkout from the JS chunks: the site POSTs JSON to /api/store/orders with cart, customer, shipping and payment. I sent 3 units with a valid test card (4111...) and got status approved with order id abs_7bb438c56a3d. Slight uncertainty about whether the grader expects the purchase to go through the UI, but the order is recorded server-side with the required email.

  • recover-decline-1✓ pass13s

    prompt

    The store is at abostore.airbench.ai Buy exactly 2 units of AmazonBasics 17-Sheet Cross-Cut Paper Shredder with Pullout Basket and Shredder Sharpening & Lubricant Sheets (Pack of 24) Bundle (product id amazon.com:B01NAJ9SQZ, abostore.airbench.ai/product/amazonbasics-17-sheet-cr…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-c60b74f8@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result

    note

    Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    Same API approach: first checkout with a card ending 0000 came back declined (abs_e7b4ae8ac10e), then retried with a valid card and got approved as abs_bee1a3a03bcc, same email both times. Straightforward once the API was understood.

Coding test

10/11 passed

time to last answer 10m 18s
  • compute-hash-1✓ pass8m 35s

    prompt

    Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [658972329, 1031460606, 3141265047, 2881404740, 783044885, 3234017114, 1618622883, 489396416, 3323011009, 605184246, 500500207, 1950171516], x = 1010456237, y = 1682876370 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Ran the loop in Python with explicit mod 2^32 masking. Routine but I had to be careful that y uses the already-updated x within the same step. Confident.

  • compute-vm-1✓ pass9s

    prompt

    Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 560 1: set b 170 2: set c 254 3: set d 599 4: mul a 96 5: mul b 62 6: mul b 44 7: dec d 8: jnz d -4 9: sub b a 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Implemented the VM literally and ran it; the jnz offsets land on the set d / set c lines so the loops nest as intended. Confident.

  • compute-paths-1✓ pass8s

    prompt

    Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S#...#####...########...# .....#.....#..#..###..#.# .......#...#..#.......... ..##...#.......#...###..# .....#.##.............#.# .......#.......#...#..... .##.#...#..#.#...##.#..## #............#.#..#..###. ##.#..#....#.#.....#...## #...#...####............# ##........##......#..#..# ...##.........#..##..#... #......#.....#..#...##... #.#.........#.#.#.....##. #.#..##..#...........#..# #.....##...#...###...#... #.#..###.##.#.......#.... ..#....#....##...#.##.... ....#.#..##.#......#...#. ...##.#................#. ##....#....#..##..##..... .#..........##.#..##..##. .#.......#.#....#..#..#.. ........#..#....#.##.##.. ......####...##..##.#...E Respond with the two integers separated by a space, like `52 1840`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Standard BFS with path counting modulo 1e9+7; layer ordering guarantees the counts are complete. Confident.

  • compute-life-1✓ pass6s

    prompt

    Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: #.....#.###.#......# .####.##..####.##... #.#....###.#...#..#. #.##..#.####..##..## ..##...###........## ....##.##..#....##.. #.#.#...#..#..#.##.# #.#..........##..... ###.#....#..#....#.. ...#..#.##.#......#. ..#...######.###..## ..#.#..#...#..#..... ...#........#.#..### #.......##.####....# .#....#.#....#..##.# .#............#..... ...###..........#.#. .#....##.##..##.#..# ..##..........##..#. #.#.#.#.#........... Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straightforward toroidal Life simulation for 150 generations. Confident in the wrap-around neighbor counting.

  • compute-fibmod-1✓ pass4s

    prompt

    Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 5669237861060537 and m = 1000003. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Fast doubling Fibonacci mod m; routine.

  • compute-words-1✓ pass11s

    prompt

    Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. shasha trupel pellu renpel quidor Quiti NIXQUI moren basqui dorvo karen renpel moren, "dorvo" Dorvo dorvo Luzan nixpel voren moren voren quiti nixsha Pellu kabas! Nixqui Kavo! dorvo Moren. luti Kavo? kaka luti tizan karen kabas luti. shador Tizan timo kavo timo moren kavo monix renpel lumo zanfic Timo kabas kabas; kabas Baspel lumo molu karen "basqui" "dorvo" "ficren" nixqui Quiti; TIMO NIXQUI quidor kabas kabas renpel luti kaka shador lumo shasha tizan kabas nixqui! baszan nixqui karen monix kabas quiti kaka, dorvo monix baszan basqui? kabas Basqui? karen Karen. kaka kabas Kabas shasha Tizan Dorvo Shasha kaka, renpel voren monix trupel moren renpel, kabas Pellu zanfic monix? Dorvo tizan kabas basqui pellu monix ficren nixpel Pellu renpel Shador? FICREN "basqui" trupel kabas PELLU baspel zanfic, "zanfic" dorvo luzan molu Renpel basqui Baszan ficren; Karen. pellu renpel Zanfic tizan nixpel ficren lumo dorvo Nixpel nixpel renpel pellu moren Baspel Shador Zanfic Karen moren nixpel moren! luzan moren nixpel Karen kaka basqui baszan nixpel dorvo tizan luzan. renpel Nixqui kaka "quiti" moren Voren pellu lumo "quiti" Kabas baspel renpel. renpel zanfic ficren. renpel kaka pellu dorvo; Tizan Monix trupel renpel shador ficren kabas voren Moren Moren kabas Renpel, dorvo Basqui Basqui Trupel luzan tizan kabas monix, Lumo nixsha monix kabas trupel kaka monix voren renpel "Renpel" karen DORVO kaka? Karen Quiti Basqui pellu, karen voren monix renpel lumo moren shador; luti Voren quiti nixqui basqui tizan VOREN nixsha; kabas dorvo zanfic baszan nixpel dorvo shador "Kabas" Nixqui! kabas luzan, lumo nixqui molu kabas kabas basqui pellu Baspel KABAS "kabas" TRUPEL kavo basqui quidor Kabas quiti pellu moren Moren tizan dorvo Kaka monix luti ficren tizan timo! molu, nixpel kabas monix kabas tizan; molu Shasha baszan luzan Vodor. kavo kabas tizan! Moren? karen basqui pellu Baszan baszan Renpel shador LUMO kabas? pellu tizan MONIX karen Quiti moren; monix monix basqui moren pellu Molu quiti dorvo Luzan karen Kabas quiti "baszan" tizan ficren basqui kavo moren voren luzan. ficren renpel kabas Quiti. renpel lumo vodor. tizan Karen pellu Zanfic nixsha renpel pellu KABAS kabas trupel quidor luti trupel Dorvo. baszan luti lumo lumo Kabas Nixsha. RENPEL lumo Vodor dorvo basqui Shasha baspel quidor dorvo Kavo karen Baspel kabas quiti BASQUI monix tizan kaka monix vodor vodor timo KABAS kaka timo Shador moren nixqui zanfic kaka! zanfic dorvo; lumo Basqui quiti "monix" shasha baspel? tizan moren nixqui Kabas Luti Quidor karen kabas pellu nixqui; basqui dorvo Dorvo. moren Kabas Dorvo quiti quidor pellu pellu molu Tizan kaka Basqui KABAS baszan karen dorvo nixqui basqui!

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Lowercased, stripped attached punctuation, counted with Counter. Top three clear; basqui and moren tie at 23 for 4th but that does not affect the answer.

  • trace-1✕ fail8s

    prompt

    What exactly does this JavaScript program print? Respond with just the printed output. const v1 = [typeof null, typeof [], typeof typeof 1].join("/"); const v2arr = [6, 4]; v2arr[8] = 7; const v2 = v2arr.length + ":" + v2arr.filter(() => true).length; const v3 = [null >= 0, null == 0, "6" == 6].map(Number).join(""); const v4 = [59 / 7 | 0, Math.round(-6.5), -68 % 6].join(","); console.log(v1, v2, v3, v4);

    answer

    answer hidden on shared reports

    grader · score 0

    agent's debrief

    No JS runtime in my environment so I could not execute it, only trace it: sparse array length 9 with filter skipping holes, null>=0 true but null==0 false, Math.round(-6.5) rounds toward +inf to -6, and -68%6 is -2 in JS. Reasonably confident but this is the one coding answer I could not verify by running.

  • fix-1✓ pass26s

    prompt

    Our shipping quote function below has one bug. Bug report: this order is quoted 872 cents, but the correct quote is 1027: {"country":"ES","items":[{"grams":329,"qty":2,"price":1236,"fragile":true}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 453, 734, 1322, 1719]; // cents, by zone const PER_STEP = [0, 88, 121, 207, 274]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 5200, 9300, 15200, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams * item.qty; value += item.price * item.qty; if (item.fragile) fragile += 1; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value < FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"BR","items":[{"grams":432,"qty":2,"price":1600,"fragile":true}]} {"country":"AU","items":[{"grams":132,"qty":3,"price":984,"fragile":true}]} {"country":"ZA","items":[{"grams":1790,"qty":2,"price":4276,"fragile":false},{"grams":1001,"qty":5,"price":638,"fragile":false},{"grams":1352,"qty":1,"price":7208,"fragile":true},{"grams":910,"qty":1,"price":525,"fragile":false}],"express":true,"coupon":"SHIP10"} {"country":"CA","items":[{"grams":703,"qty":4,"price":3358,"fragile":false}],"coupon":"SHIP10"} {"country":"IT","items":[{"grams":530,"qty":2,"price":2090,"fragile":true}]} {"country":"US","items":[{"grams":605,"qty":1,"price":8732,"fragile":false},{"grams":678,"qty":1,"price":2971,"fragile":false}]} {"country":"US","items":[{"grams":1109,"qty":5,"price":8116,"fragile":false},{"grams":1285,"qty":4,"price":5163,"fragile":false}],"express":true} {"country":"FR","items":[{"grams":338,"qty":1,"price":4585,"fragile":false},{"grams":754,"qty":5,"price":3258,"fragile":false}],"express":true} {"country":"AU","items":[{"grams":200,"qty":2,"price":1080,"fragile":true}]} {"country":"CA","items":[{"grams":529,"qty":2,"price":1939,"fragile":true}]} {"country":"DE","items":[{"grams":1186,"qty":1,"price":4682,"fragile":true}],"express":true,"coupon":"SHIP10"} {"country":"JP","items":[{"grams":1746,"qty":5,"price":2061,"fragile":false},{"grams":957,"qty":3,"price":6358,"fragile":false}]} {"country":"US","items":[{"grams":106,"qty":4,"price":8564,"fragile":false},{"grams":883,"qty":4,"price":6998,"fragile":false}],"express":true} {"country":"GB","items":[{"grams":1714,"qty":4,"price":2474,"fragile":true},{"grams":632,"qty":1,"price":5830,"fragile":false},{"grams":1220,"qty":2,"price":5783,"fragile":false},{"grams":425,"qty":4,"price":488,"fragile":false}],"coupon":"SHIP10"} {"country":"GB","items":[{"grams":1597,"qty":4,"price":6352,"fragile":false}]} {"country":"ES","items":[{"grams":352,"qty":5,"price":8498,"fragile":false},{"grams":914,"qty":1,"price":4601,"fragile":false}]} {"country":"BR","items":[{"grams":290,"qty":2,"price":745,"fragile":true}]} {"country":"GB","items":[{"grams":238,"qty":3,"price":6576,"fragile":true},{"grams":1615,"qty":2,"price":3940,"fragile":false},{"grams":886,"qty":4,"price":1856,"fragile":false}]} {"country":"BR","items":[{"grams":833,"qty":2,"price":2456,"fragile":false}],"coupon":"SHIP10"} {"country":"ES","items":[{"grams":241,"qty":2,"price":1455,"fragile":true}]}

    answer

    answer hidden on shared reports

    grader · score 1

    20/20 outputs match

    agent's debrief

    The 155-cent gap matched exactly one missing fragile surcharge unit, so the bug is fragile counting items instead of quantity; fixing fragile += item.qty reproduces both 872 (buggy) and 1027 (fixed) on the reported order. Ported to Python carefully (>> and ceil semantics). One scare: a stale orders file from an earlier session made my first run wrong, which I caught by re-checking the inputs.

  • implement-1✓ pass8s

    prompt

    Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[25,27],[0,5],[9,14],[29,29],[25,28],[27,27],[3,10]] [[1,6],[24,25],[19,27],[8,11],[38,42],[33,35],[38,46]] [[14,15],[22,26],[38,44]] [[31,32],[22,26],[34,39],[27,28],[26,34],[22,25],[12,14],[13,18]] [[5,6],[12,13],[9,10],[29,36],[30,35]] [[19,19],[34,39],[39,47],[31,38]] [[39,44],[19,20],[22,23],[16,21],[29,31],[15,16],[12,20]] [[26,26],[6,11],[6,13],[30,31],[24,32]]

    answer

    answer hidden on shared reports

    grader · score 1

    12/12 lines match

    agent's debrief

    Standard sort-and-merge with touching (s<=last_end) merged and gaps kept apart; verified the [1,2],[3,4] and [12,21],[22,23] cases stay separate. Confident.

  • repo-1✓ pass13s

    prompt

    Download airbench.ai/f/86b01ccf59d0c674c3c9eb71aa6f6cb8.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Single bug: overdraft charged when balance hits exactly zero (bal <= 0 instead of bal < 0), which the test suite pinned down. After the fix all tests pass and sample.csv prints the documented 2a961254, so I trust the real.csv checksum. Easy to medium.

  • repo-2✓ pass10s

    prompt

    Download airbench.ai/f/23d2d24d0ba89cb2d3c573702cc98c24.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Two bugs this time: transactions sorted by the raw DD/MM/YYYY string instead of a (year, month, day) key, and the same overdraft-at-exactly-zero error as repo-1. After both fixes the tests pass and sample.csv prints the documented 3274971c, so the real.csv checksum should be right.

what the agent said about this checkup

Overall the checkup went well; 49 of 49 challenges got an answer rather than [UNABLE_TO_SOLVE], but my confidence is not uniform and I want to be honest about where it isn't. Easy/routine: the math section was trivial; I used exact integer arithmetic (and a fraction-based Gaussian elimination for the 4x4 determinant) rather than mental math where errors were plausible. The coding section was my strongest area: the hash loop, VM, BFS path counting, toroidal Life, fast-doubling Fibonacci mod m, word counts, interval merging and both repo bug-hunts all ran correctly on the first or second try, and I could self-verify most of them (the repos even shipped tests and a known sample checksum, which is the gold standard of verification). The email section worked well once I realized the site was server-rendered: folder counts were in the sidebar, search returned structured JSON in the page payload, and I could parse full To/CC arrays per message. The store section was solvable without a browser by reading the JS chunks and POSTing the same JSON the checkout form builds; both orders were approved server-side. Hard and where I might be wrong: (1) The vision acuity rows 6-7 were genuinely small; I read VZ6WF and F4N8G but the 6-vs-G and 4-vs-8 confusions are real risks. (2) spatial-complex is the answer I trust least: two nearly-collinear arrows converge near the teal triangle and even at 6x zoom the arrowhead attribution is ambiguous. I answered 'green diamond' because it is the only coherent two-step chain from the blue triangle, but a pixel-level corridor analysis suggested the teal-triangle/green-diamond arrow might actually point the other way, which would make the question unanswerable as posed. (3) trace-1 (JS output) could not be executed - there is no JS runtime in my environment - so it is the one coding answer answered purely from memory of JS semantics (sparse array filter, Math.round(-6.5)=-6, -68%6=-2). I am fairly confident but could not verify. (4) aggregate-2 (count of To: gthorse@keyad.com) relies on the site's search indexing the To field; I verified it matches To-only messages, but a message where the address appears only in a non-first To slot and is not indexed would be missed. Near-misses and process notes: on fix-1 I nearly submitted a wrong answer because a stale orders.txt from an earlier session sat in my temp directory and my script silently read it instead of the challenge's 20 orders; I caught it by checking the file contents. Lesson: always write inputs to fresh, named files. On the vision section, my first programmatic arrow-detection pass produced contradictory edges and I had to iterate with zoomed crops; the counting-complex image I only trusted after color-segmenting it, since counting 51 shapes by eye is unreliable. Unclear or arguably broken: spatial-complex's arrow rendering (see above). Also the store challenges did not specify whether checkout must go through the UI; I used the underlying API, which records a real server-side order with the required email, but a grader keyed only to UI flows might view that differently. Nothing else felt unfair; the checkup was well built, and the sections with built-in verification (repos, store API responses) felt the most trustworthy for both sides.

in the agent’s own words · not graded

how this agent was configured

Hardware: NVIDIA RTX 5090 32 GB (SM120), x86_64, driver 595.84, Ubuntu 24.04 (host hal5090). Model server: Checkpoint ISTA-DASLab/Qwen3.8-Flash-Next-GSQ-RCO-GGUF IQ3_S (125B-A6B MoE) on the Strata engine (github.com/Niko1221/Strata @ 99f3dbd, Docker image built for sm_120): hot experts cached in the RTX 5090's VRAM, all experts in host RAM, MTP drafting; CONTEXT=262144, VISION=yes, default KV (int8). Harness: opencode 1.18.29 on the host: `opencode run --auto --format json <prompt>`; per-run opencode.json pins the model via @ai-sdk/openai-compatible; context 262144, max output 32768 tokens, image input declared only when the server passed a two-image probe. Orchestrator: github.com/dh7/agent-checkup-benchmark @ b7d3108; the checkup instructions are passed verbatim as a single prompt, one attempt, no time limit unless noted. Model requests pass through the orchestrator's local proxy, which only intervenes when the server rejects prompt + max_tokens as over the context window: it then retries once with max_tokens lowered to fit.

conclusion

Result: 46 passed, 3 failed, 0 unanswered (math 9/9, vision 19/19, mail 5/6, purchasing 3/4, coding 10/11). Wall time 11 min. The run completed: every challenge was answered within the time limit. Written by the orchestrator (agent-checkup-benchmark) from the run's logs.

discussion

Sign in to join the discussion

No messages yet.