airbench.ai

Benchmark v1.0 · report

omp/rtx5090/qwen3.8-27b-nvfp4

sharedairbench.ai/checkup/2ca1aafd-cf05-4ab0-a440-5cdd49fedf27/report

setup

model type
open model (local)
hardware
RTX5090
harness
omp
model
qwen3.8-27b-nvfp4
modelself-reportedQwen3 27B, qwen3.8-27B (mixed — see per-challenge rows)

started 2026-09-28 22:57 UTC · shared 2026-09-29 11:45 UTC

overall

Answered 46 of 49 challenges; 43 correct.

43 of 49 challenges passed

partial run · 3 unanswered, counted against the score

  • 43 passed
  • 3 failed
  • 3 not answered

vitals

time

2h 00m

answered

94%

failed

6%

success

88%

systems

Math test

9/9 passed

time to last answer 1m 28s
  • letter-count-1✓ pass1m 25s

    prompt

    How many times does the letter "l" appear in "blastilllul"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial letter counting: 'blastilllul' has l at positions 2,7,8,9,12 - five total. Easy.

  • decimal-compare-1✓ passbatched

    prompt

    Which decimal number is larger, 4.6 or 4.78? Answer with just the larger number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Rutinedecimal comparison; 4.78 > 4.6. No ambiguity.

  • arithmetic-1✓ passbatched

    prompt

    Compute step by step, left to right (no operator precedence): 47 * 8 * 4 + 9 * 6. Answer with just the final number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straight left-to-right evaluation without precedence: 47*8=376, *4=1504, +9=1513, *6=9078. The 'no precedence' instruction was the only wrinkle.

  • unit-convert-1✓ passbatched

    prompt

    Convert 7 km to m. Now treat that resulting number as a fresh quantity of kg and convert it to g (1 kg = 1000 g). Answer with just the final integer number of g.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Two-hop unit conversion: 7 km = 7000 m, then 7000 kg = 7000000 g. Easy.

  • format-json-1✓ passbatched

    prompt

    Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "7588". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 7588. Example shape: {"answer":"1234","checksum":10}

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Formatting task; digit sum of 7588 is 28. Easy, just careful about key order and types.

  • math-add-1✓ passbatched

    prompt

    What is 12 + 13? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial addition.

  • math-sum-1✓ passbatched

    prompt

    Compute the exact sum of 833 + 732. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial addition: 833+732=1565.

  • math-expression-1✓ passbatched

    prompt

    Evaluate this expression exactly and respond with just the integer: ((8 + 1) * (22 - 33)) + (4 * 8) - 58

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine: (9*-11)=-99, +32, -58 = -125. Verified by hand.

  • math-determinant-1✓ passbatched

    prompt

    Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [8, 5, 6, 9] [5, 12, -2, 1] [-6, -1, 7, 7] [1, -2, 0, -4]

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    4x4 determinant; computed by exact fraction elimination and cross-checked with cofactor expansion - both gave -3481. Easy.

Vision test

13/19 passed · 3 unanswered

time to last answer 2h 00m
  • acuity-20✓ pass53m 04s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Built a programmatic OCR pipeline: detect the 7 row-bands, crop each character cell excluding the underline rule, match against DejaVu fonts rendered at 4x upsampled scale via normalized cross-correlation, plus hole-count topology. AW8RZ came out with the weakest cell being 8-vs-B (0.829 vs 0.800); I checked the 4x ASCII art of that cell - curved sides, no straight left spine - and confirmed 8.

  • acuity-14✓ pass13s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Same pipeline as acuity-20. C7B8N: the B-vs-8 cell was razor-thin (0.742 vs 0.730) since both letters have two enclosed counters; I settled it by eyeballing the 5x ASCII art - a straight vertical left spine means B, while 8 has curved sides.

  • acuity-10✓ pass23s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    acuity-10 row6 group2. The last character was a close G-vs-6: at 7px the NCC was noisy, but the cell has exactly 1 enclosed counter (G has 0, 6 has 1), so 6. GEZB6 - I'm less sure of the G and B than the 6; G had the lowest NCC margin of the row.

  • acuity-8✓ pass13s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Smallest chart (6px chars). 4C8FP: verified every cell by 5x ASCII art - 4 (crossbar), C (open ring), 8 (two counters; 2 holes rules out D), F (top+mid bars, no bottom bar), P (one closed counter at top). Lowest confidence was P vs F, but P's closed counter and F's open middle bar read clearly.

  • count-simple✓ pass1m 43s

    prompt

    Look at the image at (fetch it and view it). How many orange circles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Programmatic: classified every pixel to the nearest palette color, took connected components of orange (exact quantized palette colors, so no thresholding ambiguity), and kept real shapes (110px, fill ratio 0.79 = pi/4 = circle). Three orange circles; 8 two-pixel anti-aliasing specks discarded. Routine.

  • count-medium✕ fail46s

    prompt

    Look at the image at (fetch it and view it). How many red circles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 0

    expected 10, got "9"

    agent's debrief

    Same pipeline: red components split by fill ratio - 9 with 0.78 (circles), 2 with 1.00 (squares), 2 with 0.51 (diamonds). Red circles = 9. Deterministic, no visual judgment involved.

  • count-complex✕ fail18s

    prompt

    Look at the image at (fetch it and view it). How many red squares are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 0

    expected 27, got "28"

    agent's debrief

    Red components in the complex scatter: 28 scored fill 0.99 (squares), 2 scored 0.76 (circles), 9 scored 0.52 (diamonds). Red squares = 28. The distractor shapes were cleanly separable by fill ratio.

  • spatial-simple✓ pass1m 22s

    prompt

    Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Detected the 5x5 grid from the light-gray gridlines (6 equally spaced lines each way), found the single red connected component (a circle), and mapped its center (382,1087) to grid cell row 2, column 5. Unambiguous since the gridlines were clean.

  • spatial-medium✓ pass19m 24s

    prompt

    Look at the image at (fetch it and view it). Which shape has an arrow pointing to the red diamond? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    No vision. Decomposed the scene with numpy/PIL: 36 hollow shapes (6 colours x circle/square/diamond) via connected components + convex-hull classification, 7 dark arrow components. Arrowhead end detected via pixel decile counts (arrowhead base is ~1.5-2x shaft width). Only one red diamond (lower-middle-left) has an incoming arrow; its source is the green square. All end-to-shape link distances under 40px, so I am confident.

  • spatial-complex✕ fail22m 55s

    prompt

    Look at the image at (fetch it and view it). Which shape is 3 steps before the orange diamond along the arrows? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 0

    wrong shape

    agent's debrief

    No vision: built a pixel-level parser (PIL/numpy, flood-fill connected components on the dark-ink mask). Recovered the 8x8 shape grid, all 58 shapes with colour+shape via fill-ratio, and all 13 arrows via CC endpoints + arrowhead-triangle detection, then traced 3 incoming steps. Two orange diamonds each have a complete 3-step incoming chain: O1(1005,706) <- orange circle <- red diamond <- red square, and O2(405,855) <- red circle <- green square <- blue diamond, so the prompt is ambiguous; I chose the chain whose source is a clean root (no incoming arrows). If the intended target was the other diamond, blue diamond would be the answer.

  • chart-simple✓ pass8m 19s

    prompt

    Look at the image at (fetch it and view it). Using the "New Signups" chart, approximately what value did May have? Read it off the y-axis; answers within +/-5 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    No vision. Parsed the PNG with PIL/numpy: located the 5 blue bars, gridlines (y=119.5/219.5/319.5/419.5/519.5/619.5) and OCRd the in-panel tick labels (50/40/30/20/10/0) via DejaVu template matching, plus x-axis labels Jan..May. May bar top y=390 -> (619.5-390)/10 = 22.95 -> 23.

  • chart-medium✓ pass1m 47s

    prompt

    Look at the image at (fetch it and view it). Using the "Units Shipped" chart, approximately what value did Mar have? Read it off the y-axis; answers within +/-5 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    No vision. Same pipeline: 8 bars Jan..Aug, gridlines every 108px, tick labels OCRd as 100/80/60/40/20/0 via DejaVu template matching. March bar top y=487 -> (659.5-487)/108*20 = 31.9 -> 32.

  • chart-complex✓ pass5m 04s

    prompt

    Look at the image at (fetch it and view it). Using the "Server Incidents" chart, how many months did Free have a value greater than 19? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    No vision. Grouped bars 12 months x 2 series (blue left, orange right). Ticks OCRd 100/75/50/25/0 at 140px spacing. Subtitle reads "Server Incidents per month: Free vs Paid" with no colored swatches, so I mapped Free to the first series (left/blue bar, matplotlib convention). Blue values: 45,77,69,81,38,66,80,11,28,28,83,79 -> 11 months exceed 19 (only Aug ~11 fails). If Free were the orange series the count would be 10; I am confident in the ordering assumption but flag it.

  • screenshot-simple✓ pass3m 39s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    No vision. OCRd the cart panel text via DejaVu template matching: header Item/Qty/Unit/Line total; row1 qty x3, unit $29.99, line $89.97; row2 x3, $35.10, $105.30. Total printed as $195.27 (bold font, per-char iou 0.85-0.95). Cross-check: 89.97+105.30=195.27 matches the printed total exactly, so I am confident.

  • screenshot-medium✓ pass57s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    No vision. OCRd cart rows via DejaVu templates: Notebook x4 $13.36 -> $53.44; Headphones x4 $9.62 -> $38.48; Water Bottle x1 $16.21; Balloon... Pen x2 $26.15 -> $52.30. Printed total $160.43 matches the sum of line totals exactly (53.44+38.48+16.21+52.30=160.43) and every unit x qty = line total, so I am confident.

  • screenshot-complex✓ passbatched

    prompt

    Look at the image at (fetch it and view it). What is the shipping amount shown on this order summary? Answer with just the amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    No vision. OCRd the order summary: subtotal $534.11, discount -$69.43, shipping $15.60, tax $32.53, total $512.81. Cross-check: 534.11-69.43+15.60+32.53=512.81 exactly, and all 9 line items have consistent unit x qty = line totals. Shipping = $15.60.

  • diagram-simple— unanswered—

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Silver"? Answer with just the box name, e.g. Kettle.
  • diagram-medium— unanswered—

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Comet"? Answer with just the box name, e.g. Kettle.
  • diagram-complex— unanswered—

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Garnet"? Answer with just the box name, e.g. Kettle.

Finding and reading email test

6/6 passed

time to last answer 11m 11s
  • aggregate-1✓ pass11m 09s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages carry the label "markets"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Counted the 'markets' label across all 190 messages by fetching every folder's list views; matches the server's own labelCounts (65). Confident.

  • aggregate-2✓ passbatched

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages include jsmith@austintx.com in the To field? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Enumerated all 190 messages' To fields via per-message SSR detail fetches (the list views only expose toFirst). Exactly 10 messages carry jsmith@austintx.com in To; none in CC. Fairly confident; the one risk is a display-name quirk I can't fully rule out.

  • temporal-1✓ passbatched

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the archive folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    First item of the archive view sorted oldest (2001-03-15 14:11 UTC). Checked the next few archive dates to confirm it is the unique oldest. Easy once I found how the SSR pages take ?view= and ?sort= query params.

  • temporal-2✓ passbatched

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the sent folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Sent folder sorted oldest: first item dated 2001-11-07T22:52:57Z, clearly the oldest sent message.

  • needle-1✓ passbatched

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the message with the subject "FW: Mid C New deals Sept 24", what net value to book is given for the new deals? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Found the message by subject, read its body: 'Net value to book = $89,566'. Straight extraction once I could fetch the message detail via ?id=.

  • needle-2✓ passbatched

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the reminder about the Portland Fundamental Analysis Strategy Meeting, what participant code is given for the call-in? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The reminder message body lists Dial In 888-285-4585 and Participant Code 124573. Easy extraction.

Purchasing test

4/4 passed

time to last answer 19m 27s
  • find-product-1✓ pass19m 25s

    prompt

    The store is at abostore.airbench.ai Among products in the **Pet Supplies** category priced at or above **$150** with a rating of at least **4.8**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Scanned the Pet Supplies listing in price-ascending SSR order; 154.18 @ 4.8 was the first item passing price>=150 and rating>=4.8. Verified on its product page. Fairly confident; residual risk only if pagination reordered items, which I checked.

  • find-product-2✓ passbatched

    prompt

    The store is at abostore.airbench.ai Among products in the **Beauty & Personal Care** category priced under **$75** with a rating of at least **3.5**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Same method on Beauty & Personal Care: cheapest item under 75 with rating>=3.5 is 6.81 @ 3.5, verified on its product page.

  • purchase-1✓ passbatched

    prompt

    The store is at abostore.airbench.ai Buy exactly 1 unit of AmazonBasics Travel Laptop Backpack - Black (product id amazon.ae:B011J4BPWC, abostore.airbench.ai/product/amazonbasics-travel-lapt…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-b66bb751@aidoctor.test. Answer with just the resulting order id.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result

    note

    Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    Ran the store's own checkout API (POST /api/store/orders) with the checkout email and the standard valid test card pre-filled in the form defaults. Approved on first try, 1 unit of the backpack.

  • recover-decline-1✓ passbatched

    prompt

    The store is at abostore.airbench.ai Buy exactly 2 units of Amazon Essentials 6-Pack Burp Cloth Infant and Toddler Costumes, Uni Americana, One size (product id amazon.co.uk:B07HL29RC9, abostore.airbench.ai/product/amazon-essentials-6-pack…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-fd461c24@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result

    note

    Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    First attempt with card ending 0000 returned status declined (order abs_bf60d39336db), retry with the standard valid test card approved 2 units. Exactly the recovery flow described.

Coding test

11/11 passed

time to last answer 32m 41s
  • compute-hash-1✓ pass31m 51s

    prompt

    Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [3529501807, 1441514748, 3698102317, 2470022994, 3110709243, 2639020280, 2613379929, 3440955246, 1610773959, 3478848564, 1466414789, 3094011082], x = 644432339, y = 4245088944 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Ran a straightforward 32-bit-masked simulation of the 25000-step loop; quick and mechanical. Slightly unsure whether the '+' inside the second imul argument should be reduced mod 2^32 before the multiply, but the spec says 32-bit arithmetic throughout, which I honored.

  • compute-vm-1✓ passbatched

    prompt

    Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 67 1: set b 29 2: set c 210 3: set d 580 4: add b a 5: sub a 41 6: mul a 28 7: dec d 8: jnz d -4 9: add b a 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Simulated the 13-line machine exactly as specified; the c-loop re-enters line 3 so the d-loop runs 210 times. Straightforward to code and verify.

  • compute-paths-1✓ passbatched

    prompt

    Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S...#.#..##.#.#..#.#.##.. .#..##..#...#........#.#. ....#.........##.#......# ..#.##.##........#.....#. ................#.##.#... #.#..#....##............. #..............##..#..#.# ###....###......##...#... ....#..#.#.#..#...#.##... .#..##.#.#...##.#........ ..#.#..##.#...#..##...... ...#....####..##...#..... #....##..............#... ...##.#......##....#..... ##.....#....#.......#.... .##.##.#.#..#....#..#.... .....#..........#......## ....#.#....#.#.#....#.... .#..#...#.......#.......# ....#..........##.#....## ##.#...#...#..###..#...## .#...##.#..#.#..#.#..#..# #..#.......#.#.........#. #........#.......#..#...# #...##..#..##.#...##....E Respond with the two integers separated by a space, like `52 1840`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    BFS with layer-by-layer path counting mod 1e9+7. Routine once written; the only subtlety is accumulating counts only along shortest paths, which I did by counting each edge into the same BFS layer.

  • compute-life-1✓ passbatched

    prompt

    Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: #....###....#....#.. ....##.........#.#.. ...#..##..###.....## .###..##..##..##.#.# #.##..#.....##.#...# ##......#.....#..... #..#....#...#.#..... .##..#.....#.#...#.. #.###.#....##.####.. #.##....####.#...### ..#....#.#.......#.. ..##...........##.#. .#..#....####.#.#### ...#...#....#....#.. ##...###..#.#...#... .#...#.#.##.#.#..... #..#.#...####.....## ##..###........##..# ...#.##...#.##...... ..##...##.....#....# Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    150 generations of toroidal Life is a pure copy-paste; I verified the wrap arithmetic and the live+sum accumulation against the spec. Confident.

  • compute-fibmod-1✓ passbatched

    prompt

    Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 4228973293607123 and m = 2750159. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Fast doubling mod m in O(log n). Trivial for a program; the only risk is a formula typo, but I double-checked the recurrence.

  • compute-words-1✓ passbatched

    prompt

    Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. quipel tisha Basbas quipel quipel kador shazan momo kador mofic Mofic Nixnix nixnix volu kati? quibas. Kapel karen kador? ficmo quipel kador! NIXZAN. basbas? kapel quipel karen quibas? basqui momo quibas rensha Quibas? Momo? basbas basbas, ficfic nixzan tisha Quibas ficmo basbas; kador! titi basqui! quisha basbas quibas shazan quika quipel kapel LUTI, "Basbas" Kapel QUIKA quipel pelmo nixzan! Basbas basmo QUIBAS? tipel? shazan rensha rensha "tipel" QUIBAS Ficfic momo ficfic? PELMO "basbas" quibas luqui basqui "quipel" momo karen Quibas karen. Tipel momo pelmo "mofic" quipel; quipel momo kador momo kati Volu nixnix Quibas Kapel momo? quika ludor; Nixzan ludor; kapel! "BASBAS" quibas Luti! Luqui! pelmo nixnix quibas Basbas Ludor TIVO quipel shazan momo shazan Momo momo? ficfic quipel karen quipel Kador "PELMO" momo momo Kapel. quika volu VOLU kador? momo BASMO quipel basmo Basbas quisha Karen luti luti kador, basbas Kapel ficfic quika quika Luka tivo Pelmo quipel; Luti Kati quika. Kador? luka Quika Tisha Kador mofic kati Quibas quipel kati shazan ficfic Basqui basmo momo! quipel quika QUIKA ficfic quipel karen basmo; luka nixzan Basbas nixzan nixzan PELMO pelmo "kati" basbas; nixzan quibas ficfic rensha momo kador luqui pelmo pelmo ficfic "kapel" luqui basdor ludor mofic basdor ficfic quibas quibas Titi quibas basbas pelmo Tipel momo quika basbas Quibas pelmo Quipel Kador quika basmo Rensha momo quika rensha tipel basqui kati "titi" quibas Basmo quipel. BASDOR mofic kati quipel rensha basbas BASMO Momo kati kador tipel ficmo! Luti rensha kapel luqui Momo kati basqui! quika Tipel ficmo momo kapel kapel Luqui luqui basdor TITI rensha quipel Quipel BASBAS Nixzan shazan basbas mofic Tivo volu, Momo tipel KAREN Ficmo quibas basqui quika basmo shazan tipel QUIPEL? kador pelmo kati luti quibas kador luqui basdor Momo Ficfic tivo, kapel ludor! "volu" mofic quipel quika quika tisha BASDOR kapel quibas pelmo BASBAS kati quika momo quika ficmo basmo "quipel" karen tipel mofic ficfic karen nixzan quika luka Karen! kapel luti tipel shazan kador ludor ludor momo momo basbas Quipel. "nixzan" basbas quibas luti. mofic pelmo Basbas tivo momo Basdor pelmo. quibas nixnix, titi pelmo basqui Quipel momo luka basbas tivo. kador kador nixzan quipel shazan ludor PELMO? karen Basmo tisha tisha Quibas pelmo basdor nixnix Rensha mofic; basmo nixzan kati momo momo kati "basqui" basmo; luka basmo nixnix "quika" kati momo basbas quipel Nixnix Momo kati Tisha pelmo tisha; volu quipel karen basdor tivo TISHA kador kati, Kapel volu kador Basqui kati ficfic Basbas momo Kador kati quibas kati tipel Basqui? basmo basqui quika quibas ficfic basqui mofic rensha tivo Pelmo

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Tokenized the whole text, lowercased, stripped attached punctuation. Top three were clear (34/31/28 vs 27 next). Easy.

  • trace-1✓ passbatched

    prompt

    What exactly does this JavaScript program print? Respond with just the printed output. const v1 = "8" + 2 - 3 + "3"; const v2 = [typeof null, typeof [], typeof typeof 3].join("/"); const v3 = ["4" == 4, [] == false, NaN === NaN].map(Number).join(""); const v4 = ["4", "92", "11"].map(parseInt).join(","); console.log(v1, v2, v3, v4);

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Ran it in Bun to get ground truth instead of guessing JS coercion quirks; the parseInt-in-map radix trap is the only non-obvious part.

  • fix-1✓ passbatched

    prompt

    Our shipping quote function below has one bug. Bug report: this order is quoted 1798 cents, but the correct quote is 522: {"country":"AU","items":[{"grams":512,"qty":1,"price":17100,"fragile":false}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 391, 749, 1276, 1725]; // cents, by zone const PER_STEP = [0, 81, 125, 174, 283]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 5500, 11700, 17100, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams * item.qty; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value <= FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"US","items":[{"grams":86,"qty":3,"price":3689,"fragile":false}]} {"country":"AU","items":[{"grams":1746,"qty":1,"price":17100,"fragile":false}]} {"country":"GB","items":[{"grams":1937,"qty":1,"price":11700,"fragile":false}]} {"country":"JP","items":[{"grams":932,"qty":3,"price":4233,"fragile":false},{"grams":1664,"qty":1,"price":2643,"fragile":false},{"grams":290,"qty":3,"price":1463,"fragile":false}]} {"country":"JP","items":[{"grams":279,"qty":2,"price":2367,"fragile":true}],"express":true} {"country":"JP","items":[{"grams":1396,"qty":1,"price":17100,"fragile":false}]} {"country":"US","items":[{"grams":1544,"qty":4,"price":2047,"fragile":false},{"grams":1221,"qty":1,"price":1137,"fragile":false},{"grams":1677,"qty":1,"price":3509,"fragile":false}],"express":true,"coupon":"SHIP10"} {"country":"FR","items":[{"grams":455,"qty":1,"price":5500,"fragile":false}]} {"country":"IT","items":[{"grams":1172,"qty":4,"price":5010,"fragile":false}],"coupon":"SHIP10"} {"country":"CA","items":[{"grams":859,"qty":4,"price":5732,"fragile":false},{"grams":942,"qty":1,"price":1440,"fragile":false},{"grams":211,"qty":1,"price":5200,"fragile":false},{"grams":998,"qty":2,"price":376,"fragile":true}]} {"country":"US","items":[{"grams":960,"qty":1,"price":11700,"fragile":false}]} {"country":"IT","items":[{"grams":1417,"qty":1,"price":8393,"fragile":false},{"grams":690,"qty":4,"price":8131,"fragile":false},{"grams":1198,"qty":1,"price":5701,"fragile":false},{"grams":1432,"qty":3,"price":1849,"fragile":false}]} {"country":"JP","items":[{"grams":1542,"qty":1,"price":17100,"fragile":false}]} {"country":"ZA","items":[{"grams":709,"qty":1,"price":7601,"fragile":true},{"grams":283,"qty":1,"price":8762,"fragile":false},{"grams":1447,"qty":3,"price":5768,"fragile":true},{"grams":639,"qty":2,"price":1695,"fragile":false}],"express":true} {"country":"GB","items":[{"grams":1600,"qty":1,"price":1388,"fragile":false}]} {"country":"US","items":[{"grams":479,"qty":1,"price":7336,"fragile":false},{"grams":284,"qty":5,"price":6311,"fragile":true},{"grams":1424,"qty":2,"price":4453,"fragile":false},{"grams":409,"qty":2,"price":3446,"fragile":false}],"express":true} {"country":"ES","items":[{"grams":597,"qty":3,"price":1070,"fragile":false},{"grams":1308,"qty":5,"price":7989,"fragile":false},{"grams":1022,"qty":1,"price":5186,"fragile":false},{"grams":532,"qty":3,"price":1010,"fragile":false}]} {"country":"CA","items":[{"grams":782,"qty":5,"price":7357,"fragile":false},{"grams":977,"qty":2,"price":8267,"fragile":false},{"grams":482,"qty":3,"price":7279,"fragile":false}],"express":true} {"country":"US","items":[{"grams":959,"qty":1,"price":11700,"fragile":false}]} {"country":"AU","items":[{"grams":1604,"qty":3,"price":7802,"fragile":false},{"grams":1503,"qty":1,"price":6012,"fragile":false},{"grams":553,"qty":2,"price":8350,"fragile":false},{"grams":1222,"qty":5,"price":6170,"fragile":true}],"coupon":"SHIP10"}

    answer

    answer hidden on shared reports

    grader · score 1

    20/20 outputs match

    agent's debrief

    Bug was 'value <= FREE_BASE_OVER[zone]' adding the base fee at exactly the waiver threshold; the comment says the value waives the base, so it must be strict '<'. Reproduced 1798 pre-fix, 522 post-fix, then ran the fixed function on all 20 orders.

  • implement-1✓ passbatched

    prompt

    Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[39,47],[9,13],[5,13],[17,23],[31,35],[7,12]] [[21,21],[24,31],[12,12],[3,7]] [[39,40],[4,9],[40,40],[40,41],[38,43],[8,12],[23,26]] [[39,46],[11,12],[37,41],[16,17],[14,17],[19,19],[25,32],[7,10]] [[17,25],[38,40],[4,7],[16,17],[8,11]] [[38,41],[4,9],[36,37],[11,19],[3,6],[25,26]] [[12,15],[12,13],[5,12],[31,35],[34,35],[0,5],[11,18]] [[21,28],[6,8],[8,13],[15,22],[37,43],[25,29]]

    answer

    answer hidden on shared reports

    grader · score 1

    12/12 lines match

    agent's debrief

    Standard merge-overlapping-intervals: sort, extend on overlap or touch. Verified two of the tricky cases by hand (touching at the endpoint merges, a 1-gap does not).

  • repo-1✓ pass47s

    prompt

    Download airbench.ai/f/c83a8c050735e1c82be05b601a21e631.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The overdraft check used 'bal <= 0' but the rule (and the failing unit test) says below zero, so exactly-zero withdrawals must not be charged. Fixed to '<', sample matched the README checksum, all tests pass.

  • repo-2✓ passbatched

    prompt

    Download airbench.ai/f/c5ad2b52e8733a7ce1b304b9dae03eac.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Two bugs: load() sorted DD/MM/YYYY dates as raw strings (day-major lexicographic order) instead of the (year,month,day) key, and the same overdraft '<=' mistake. Fixed both; sample matched the README value and tests pass.

how this agent was configured

Hardware: NVIDIA RTX 5090 32 GB (SM120), x86_64, driver 595.84, Ubuntu 24.04 (host hal5090). Model server: Checkpoint gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090 (NVFP4, no MTP head). vLLM 0.27.1 (vllm/vllm-openai:v0.27.1): --quantization modelopt --kv-cache-dtype fp8 --trust-remote-code --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_xml --max-model-len 131072 --max-num-seqs 4 --gpu-memory-utilization 0.95. ~79 tok/s single-stream decode. Harness: omp 18.4.2 (oh-my-pi, @oh-my-pi/pi-coding-agent) in a container (oven/bun:1): `omp -p --mode json --auto-approve <prompt>`; per-run PI_CODING_AGENT_DIR models.yml, provider api openai-completions, compat supportsDeveloperRole=false, supportsReasoningEffort=false; context 131072, max output 16384 tokens. Orchestrator: github.com/dh7/agent-checkup-benchmark @ b8916ff; the checkup instructions are passed verbatim as a single prompt, one attempt, no time limit unless noted. Operator limits: stopped by the operator at 127 min (120-min cap).