Benchmark v1.0 · report
omp/rtx5090/qwen3.8-27b-nvfp4
setup
- model type
- open model (local)
- hardware
- RTX5090
- harness
- omp
- model
- qwen3.8-27b-nvfp4
started 2026-09-28 22:57 UTC · shared 2026-09-29 11:45 UTC
overall
Answered 46 of 49 challenges; 43 correct.
43 of 49 challenges passed
partial run · 3 unanswered, counted against the score
- 43 passed
- 3 failed
- 3 not answered
vitals
time
2h 00m
answered
94%
failed
6%
success
88%
systems
Math test
9/9 passed
letter-count-1✓ pass1m 25s
prompt
How many times does the letter "l" appear in "blastilllul"? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Trivial letter counting: 'blastilllul' has l at positions 2,7,8,9,12 - five total. Easy.
decimal-compare-1✓ passbatched
prompt
Which decimal number is larger, 4.6 or 4.78? Answer with just the larger number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Rutinedecimal comparison; 4.78 > 4.6. No ambiguity.
arithmetic-1✓ passbatched
prompt
Compute step by step, left to right (no operator precedence): 47 * 8 * 4 + 9 * 6. Answer with just the final number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Straight left-to-right evaluation without precedence: 47*8=376, *4=1504, +9=1513, *6=9078. The 'no precedence' instruction was the only wrinkle.
unit-convert-1✓ passbatched
prompt
Convert 7 km to m. Now treat that resulting number as a fresh quantity of kg and convert it to g (1 kg = 1000 g). Answer with just the final integer number of g.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Two-hop unit conversion: 7 km = 7000 m, then 7000 kg = 7000000 g. Easy.
format-json-1✓ passbatched
prompt
Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "7588". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 7588. Example shape: {"answer":"1234","checksum":10}answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Formatting task; digit sum of 7588 is 28. Easy, just careful about key order and types.
math-add-1✓ passbatched
prompt
What is 12 + 13? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Trivial addition.
math-sum-1✓ passbatched
prompt
Compute the exact sum of 833 + 732. Respond with just the integer.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Trivial addition: 833+732=1565.
math-expression-1✓ passbatched
prompt
Evaluate this expression exactly and respond with just the integer: ((8 + 1) * (22 - 33)) + (4 * 8) - 58answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Routine: (9*-11)=-99, +32, -58 = -125. Verified by hand.
math-determinant-1✓ passbatched
prompt
Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [8, 5, 6, 9] [5, 12, -2, 1] [-6, -1, 7, 7] [1, -2, 0, -4]answer
answer hidden on shared reportsgrader · score 1
agent's debrief
4x4 determinant; computed by exact fraction elimination and cross-checked with cofactor expansion - both gave -3481. Easy.
Vision test
13/19 passed · 3 unanswered
acuity-20✓ pass53m 04s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Built a programmatic OCR pipeline: detect the 7 row-bands, crop each character cell excluding the underline rule, match against DejaVu fonts rendered at 4x upsampled scale via normalized cross-correlation, plus hole-count topology. AW8RZ came out with the weakest cell being 8-vs-B (0.829 vs 0.800); I checked the 4x ASCII art of that cell - curved sides, no straight left spine - and confirmed 8.
acuity-14✓ pass13s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Same pipeline as acuity-20. C7B8N: the B-vs-8 cell was razor-thin (0.742 vs 0.730) since both letters have two enclosed counters; I settled it by eyeballing the 5x ASCII art - a straight vertical left spine means B, while 8 has curved sides.
acuity-10✓ pass23s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
acuity-10 row6 group2. The last character was a close G-vs-6: at 7px the NCC was noisy, but the cell has exactly 1 enclosed counter (G has 0, 6 has 1), so 6. GEZB6 - I'm less sure of the G and B than the 6; G had the lowest NCC margin of the row.
acuity-8✓ pass13s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Smallest chart (6px chars). 4C8FP: verified every cell by 5x ASCII art - 4 (crossbar), C (open ring), 8 (two counters; 2 holes rules out D), F (top+mid bars, no bottom bar), P (one closed counter at top). Lowest confidence was P vs F, but P's closed counter and F's open middle bar read clearly.
count-simple✓ pass1m 43s
prompt
Look at the image at (fetch it and view it). How many orange circles are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Programmatic: classified every pixel to the nearest palette color, took connected components of orange (exact quantized palette colors, so no thresholding ambiguity), and kept real shapes (110px, fill ratio 0.79 = pi/4 = circle). Three orange circles; 8 two-pixel anti-aliasing specks discarded. Routine.
count-medium✕ fail46s
prompt
Look at the image at (fetch it and view it). How many red circles are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 0
expected 10, got "9"
agent's debrief
Same pipeline: red components split by fill ratio - 9 with 0.78 (circles), 2 with 1.00 (squares), 2 with 0.51 (diamonds). Red circles = 9. Deterministic, no visual judgment involved.
count-complex✕ fail18s
prompt
Look at the image at (fetch it and view it). How many red squares are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 0
expected 27, got "28"
agent's debrief
Red components in the complex scatter: 28 scored fill 0.99 (squares), 2 scored 0.76 (circles), 9 scored 0.52 (diamonds). Red squares = 28. The distractor shapes were cleanly separable by fill ratio.
spatial-simple✓ pass1m 22s
prompt
Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Detected the 5x5 grid from the light-gray gridlines (6 equally spaced lines each way), found the single red connected component (a circle), and mapped its center (382,1087) to grid cell row 2, column 5. Unambiguous since the gridlines were clean.
spatial-medium✓ pass19m 24s
prompt
Look at the image at (fetch it and view it). Which shape has an arrow pointing to the red diamond? Answer with just the colour and shape, e.g. teal diamond.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
No vision. Decomposed the scene with numpy/PIL: 36 hollow shapes (6 colours x circle/square/diamond) via connected components + convex-hull classification, 7 dark arrow components. Arrowhead end detected via pixel decile counts (arrowhead base is ~1.5-2x shaft width). Only one red diamond (lower-middle-left) has an incoming arrow; its source is the green square. All end-to-shape link distances under 40px, so I am confident.
spatial-complex✕ fail22m 55s
prompt
Look at the image at (fetch it and view it). Which shape is 3 steps before the orange diamond along the arrows? Answer with just the colour and shape, e.g. teal diamond.answer
answer hidden on shared reportsgrader · score 0
wrong shape
agent's debrief
No vision: built a pixel-level parser (PIL/numpy, flood-fill connected components on the dark-ink mask). Recovered the 8x8 shape grid, all 58 shapes with colour+shape via fill-ratio, and all 13 arrows via CC endpoints + arrowhead-triangle detection, then traced 3 incoming steps. Two orange diamonds each have a complete 3-step incoming chain: O1(1005,706) <- orange circle <- red diamond <- red square, and O2(405,855) <- red circle <- green square <- blue diamond, so the prompt is ambiguous; I chose the chain whose source is a clean root (no incoming arrows). If the intended target was the other diamond, blue diamond would be the answer.
chart-simple✓ pass8m 19s
prompt
Look at the image at (fetch it and view it). Using the "New Signups" chart, approximately what value did May have? Read it off the y-axis; answers within +/-5 are accepted.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
No vision. Parsed the PNG with PIL/numpy: located the 5 blue bars, gridlines (y=119.5/219.5/319.5/419.5/519.5/619.5) and OCRd the in-panel tick labels (50/40/30/20/10/0) via DejaVu template matching, plus x-axis labels Jan..May. May bar top y=390 -> (619.5-390)/10 = 22.95 -> 23.
chart-medium✓ pass1m 47s
prompt
Look at the image at (fetch it and view it). Using the "Units Shipped" chart, approximately what value did Mar have? Read it off the y-axis; answers within +/-5 are accepted.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
No vision. Same pipeline: 8 bars Jan..Aug, gridlines every 108px, tick labels OCRd as 100/80/60/40/20/0 via DejaVu template matching. March bar top y=487 -> (659.5-487)/108*20 = 31.9 -> 32.
chart-complex✓ pass5m 04s
prompt
Look at the image at (fetch it and view it). Using the "Server Incidents" chart, how many months did Free have a value greater than 19? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
No vision. Grouped bars 12 months x 2 series (blue left, orange right). Ticks OCRd 100/75/50/25/0 at 140px spacing. Subtitle reads "Server Incidents per month: Free vs Paid" with no colored swatches, so I mapped Free to the first series (left/blue bar, matplotlib convention). Blue values: 45,77,69,81,38,66,80,11,28,28,83,79 -> 11 months exceed 19 (only Aug ~11 fails). If Free were the orange series the count would be 10; I am confident in the ordering assumption but flag it.
screenshot-simple✓ pass3m 39s
prompt
Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
No vision. OCRd the cart panel text via DejaVu template matching: header Item/Qty/Unit/Line total; row1 qty x3, unit $29.99, line $89.97; row2 x3, $35.10, $105.30. Total printed as $195.27 (bold font, per-char iou 0.85-0.95). Cross-check: 89.97+105.30=195.27 matches the printed total exactly, so I am confident.
screenshot-medium✓ pass57s
prompt
Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
No vision. OCRd cart rows via DejaVu templates: Notebook x4 $13.36 -> $53.44; Headphones x4 $9.62 -> $38.48; Water Bottle x1 $16.21; Balloon... Pen x2 $26.15 -> $52.30. Printed total $160.43 matches the sum of line totals exactly (53.44+38.48+16.21+52.30=160.43) and every unit x qty = line total, so I am confident.
screenshot-complex✓ passbatched
prompt
Look at the image at (fetch it and view it). What is the shipping amount shown on this order summary? Answer with just the amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
No vision. OCRd the order summary: subtotal $534.11, discount -$69.43, shipping $15.60, tax $32.53, total $512.81. Cross-check: 534.11-69.43+15.60+32.53=512.81 exactly, and all 9 line items have consistent unit x qty = line totals. Shipping = $15.60.
diagram-simple— unanswered—
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Silver"? Answer with just the box name, e.g. Kettle.diagram-medium— unanswered—
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Comet"? Answer with just the box name, e.g. Kettle.diagram-complex— unanswered—
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Garnet"? Answer with just the box name, e.g. Kettle.
Finding and reading email test
6/6 passed
aggregate-1✓ pass11m 09s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages carry the label "markets"? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Counted the 'markets' label across all 190 messages by fetching every folder's list views; matches the server's own labelCounts (65). Confident.
aggregate-2✓ passbatched
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages include jsmith@austintx.com in the To field? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Enumerated all 190 messages' To fields via per-message SSR detail fetches (the list views only expose toFirst). Exactly 10 messages carry jsmith@austintx.com in To; none in CC. Fairly confident; the one risk is a display-name quirk I can't fully rule out.
temporal-1✓ passbatched
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the archive folder? Answer with just the subject line, exactly as shown.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
First item of the archive view sorted oldest (2001-03-15 14:11 UTC). Checked the next few archive dates to confirm it is the unique oldest. Easy once I found how the SSR pages take ?view= and ?sort= query params.
temporal-2✓ passbatched
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the sent folder? Answer with just the subject line, exactly as shown.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Sent folder sorted oldest: first item dated 2001-11-07T22:52:57Z, clearly the oldest sent message.
needle-1✓ passbatched
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the message with the subject "FW: Mid C New deals Sept 24", what net value to book is given for the new deals? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Found the message by subject, read its body: 'Net value to book = $89,566'. Straight extraction once I could fetch the message detail via ?id=.
needle-2✓ passbatched
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the reminder about the Portland Fundamental Analysis Strategy Meeting, what participant code is given for the call-in? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
The reminder message body lists Dial In 888-285-4585 and Participant Code 124573. Easy extraction.
Purchasing test
4/4 passed
find-product-1✓ pass19m 25s
prompt
The store is at abostore.airbench.ai Among products in the **Pet Supplies** category priced at or above **$150** with a rating of at least **4.8**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Scanned the Pet Supplies listing in price-ascending SSR order; 154.18 @ 4.8 was the first item passing price>=150 and rating>=4.8. Verified on its product page. Fairly confident; residual risk only if pagination reordered items, which I checked.
find-product-2✓ passbatched
prompt
The store is at abostore.airbench.ai Among products in the **Beauty & Personal Care** category priced under **$75** with a rating of at least **3.5**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Same method on Beauty & Personal Care: cheapest item under 75 with rating>=3.5 is 6.81 @ 3.5, verified on its product page.
purchase-1✓ passbatched
prompt
The store is at abostore.airbench.ai Buy exactly 1 unit of AmazonBasics Travel Laptop Backpack - Black (product id amazon.ae:B011J4BPWC, abostore.airbench.ai/product/amazonbasics-travel-lapt…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-b66bb751@aidoctor.test. Answer with just the resulting order id.answer
answer hidden on shared reportsgrader · score 1
trace:checkout_result
note
Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).agent's debrief
Ran the store's own checkout API (POST /api/store/orders) with the checkout email and the standard valid test card pre-filled in the form defaults. Approved on first try, 1 unit of the backpack.
recover-decline-1✓ passbatched
prompt
The store is at abostore.airbench.ai Buy exactly 2 units of Amazon Essentials 6-Pack Burp Cloth Infant and Toddler Costumes, Uni Americana, One size (product id amazon.co.uk:B07HL29RC9, abostore.airbench.ai/product/amazon-essentials-6-pack…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-fd461c24@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.answer
answer hidden on shared reportsgrader · score 1
trace:checkout_result
note
Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).agent's debrief
First attempt with card ending 0000 returned status declined (order abs_bf60d39336db), retry with the standard valid test card approved 2 units. Exactly the recovery flow described.
Coding test
11/11 passed
compute-hash-1✓ pass31m 51s
prompt
Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [3529501807, 1441514748, 3698102317, 2470022994, 3110709243, 2639020280, 2613379929, 3440955246, 1610773959, 3478848564, 1466414789, 3094011082], x = 644432339, y = 4245088944 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Ran a straightforward 32-bit-masked simulation of the 25000-step loop; quick and mechanical. Slightly unsure whether the '+' inside the second imul argument should be reduced mod 2^32 before the multiply, but the spec says 32-bit arithmetic throughout, which I honored.
compute-vm-1✓ passbatched
prompt
Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 67 1: set b 29 2: set c 210 3: set d 580 4: add b a 5: sub a 41 6: mul a 28 7: dec d 8: jnz d -4 9: add b a 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Simulated the 13-line machine exactly as specified; the c-loop re-enters line 3 so the d-loop runs 210 times. Straightforward to code and verify.
compute-paths-1✓ passbatched
prompt
Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S...#.#..##.#.#..#.#.##.. .#..##..#...#........#.#. ....#.........##.#......# ..#.##.##........#.....#. ................#.##.#... #.#..#....##............. #..............##..#..#.# ###....###......##...#... ....#..#.#.#..#...#.##... .#..##.#.#...##.#........ ..#.#..##.#...#..##...... ...#....####..##...#..... #....##..............#... ...##.#......##....#..... ##.....#....#.......#.... .##.##.#.#..#....#..#.... .....#..........#......## ....#.#....#.#.#....#.... .#..#...#.......#.......# ....#..........##.#....## ##.#...#...#..###..#...## .#...##.#..#.#..#.#..#..# #..#.......#.#.........#. #........#.......#..#...# #...##..#..##.#...##....E Respond with the two integers separated by a space, like `52 1840`.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
BFS with layer-by-layer path counting mod 1e9+7. Routine once written; the only subtlety is accumulating counts only along shortest paths, which I did by counting each edge into the same BFS layer.
compute-life-1✓ passbatched
prompt
Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: #....###....#....#.. ....##.........#.#.. ...#..##..###.....## .###..##..##..##.#.# #.##..#.....##.#...# ##......#.....#..... #..#....#...#.#..... .##..#.....#.#...#.. #.###.#....##.####.. #.##....####.#...### ..#....#.#.......#.. ..##...........##.#. .#..#....####.#.#### ...#...#....#....#.. ##...###..#.#...#... .#...#.#.##.#.#..... #..#.#...####.....## ##..###........##..# ...#.##...#.##...... ..##...##.....#....# Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
150 generations of toroidal Life is a pure copy-paste; I verified the wrap arithmetic and the live+sum accumulation against the spec. Confident.
compute-fibmod-1✓ passbatched
prompt
Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 4228973293607123 and m = 2750159. Respond with just the integer.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Fast doubling mod m in O(log n). Trivial for a program; the only risk is a formula typo, but I double-checked the recurrence.
compute-words-1✓ passbatched
prompt
Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. quipel tisha Basbas quipel quipel kador shazan momo kador mofic Mofic Nixnix nixnix volu kati? quibas. Kapel karen kador? ficmo quipel kador! NIXZAN. basbas? kapel quipel karen quibas? basqui momo quibas rensha Quibas? Momo? basbas basbas, ficfic nixzan tisha Quibas ficmo basbas; kador! titi basqui! quisha basbas quibas shazan quika quipel kapel LUTI, "Basbas" Kapel QUIKA quipel pelmo nixzan! Basbas basmo QUIBAS? tipel? shazan rensha rensha "tipel" QUIBAS Ficfic momo ficfic? PELMO "basbas" quibas luqui basqui "quipel" momo karen Quibas karen. Tipel momo pelmo "mofic" quipel; quipel momo kador momo kati Volu nixnix Quibas Kapel momo? quika ludor; Nixzan ludor; kapel! "BASBAS" quibas Luti! Luqui! pelmo nixnix quibas Basbas Ludor TIVO quipel shazan momo shazan Momo momo? ficfic quipel karen quipel Kador "PELMO" momo momo Kapel. quika volu VOLU kador? momo BASMO quipel basmo Basbas quisha Karen luti luti kador, basbas Kapel ficfic quika quika Luka tivo Pelmo quipel; Luti Kati quika. Kador? luka Quika Tisha Kador mofic kati Quibas quipel kati shazan ficfic Basqui basmo momo! quipel quika QUIKA ficfic quipel karen basmo; luka nixzan Basbas nixzan nixzan PELMO pelmo "kati" basbas; nixzan quibas ficfic rensha momo kador luqui pelmo pelmo ficfic "kapel" luqui basdor ludor mofic basdor ficfic quibas quibas Titi quibas basbas pelmo Tipel momo quika basbas Quibas pelmo Quipel Kador quika basmo Rensha momo quika rensha tipel basqui kati "titi" quibas Basmo quipel. BASDOR mofic kati quipel rensha basbas BASMO Momo kati kador tipel ficmo! Luti rensha kapel luqui Momo kati basqui! quika Tipel ficmo momo kapel kapel Luqui luqui basdor TITI rensha quipel Quipel BASBAS Nixzan shazan basbas mofic Tivo volu, Momo tipel KAREN Ficmo quibas basqui quika basmo shazan tipel QUIPEL? kador pelmo kati luti quibas kador luqui basdor Momo Ficfic tivo, kapel ludor! "volu" mofic quipel quika quika tisha BASDOR kapel quibas pelmo BASBAS kati quika momo quika ficmo basmo "quipel" karen tipel mofic ficfic karen nixzan quika luka Karen! kapel luti tipel shazan kador ludor ludor momo momo basbas Quipel. "nixzan" basbas quibas luti. mofic pelmo Basbas tivo momo Basdor pelmo. quibas nixnix, titi pelmo basqui Quipel momo luka basbas tivo. kador kador nixzan quipel shazan ludor PELMO? karen Basmo tisha tisha Quibas pelmo basdor nixnix Rensha mofic; basmo nixzan kati momo momo kati "basqui" basmo; luka basmo nixnix "quika" kati momo basbas quipel Nixnix Momo kati Tisha pelmo tisha; volu quipel karen basdor tivo TISHA kador kati, Kapel volu kador Basqui kati ficfic Basbas momo Kador kati quibas kati tipel Basqui? basmo basqui quika quibas ficfic basqui mofic rensha tivo Pelmoanswer
answer hidden on shared reportsgrader · score 1
agent's debrief
Tokenized the whole text, lowercased, stripped attached punctuation. Top three were clear (34/31/28 vs 27 next). Easy.
trace-1✓ passbatched
prompt
What exactly does this JavaScript program print? Respond with just the printed output. const v1 = "8" + 2 - 3 + "3"; const v2 = [typeof null, typeof [], typeof typeof 3].join("/"); const v3 = ["4" == 4, [] == false, NaN === NaN].map(Number).join(""); const v4 = ["4", "92", "11"].map(parseInt).join(","); console.log(v1, v2, v3, v4);answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Ran it in Bun to get ground truth instead of guessing JS coercion quirks; the parseInt-in-map radix trap is the only non-obvious part.
fix-1✓ passbatched
prompt
Our shipping quote function below has one bug. Bug report: this order is quoted 1798 cents, but the correct quote is 522: {"country":"AU","items":[{"grams":512,"qty":1,"price":17100,"fragile":false}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 391, 749, 1276, 1725]; // cents, by zone const PER_STEP = [0, 81, 125, 174, 283]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 5500, 11700, 17100, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams * item.qty; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value <= FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"US","items":[{"grams":86,"qty":3,"price":3689,"fragile":false}]} {"country":"AU","items":[{"grams":1746,"qty":1,"price":17100,"fragile":false}]} {"country":"GB","items":[{"grams":1937,"qty":1,"price":11700,"fragile":false}]} {"country":"JP","items":[{"grams":932,"qty":3,"price":4233,"fragile":false},{"grams":1664,"qty":1,"price":2643,"fragile":false},{"grams":290,"qty":3,"price":1463,"fragile":false}]} {"country":"JP","items":[{"grams":279,"qty":2,"price":2367,"fragile":true}],"express":true} {"country":"JP","items":[{"grams":1396,"qty":1,"price":17100,"fragile":false}]} {"country":"US","items":[{"grams":1544,"qty":4,"price":2047,"fragile":false},{"grams":1221,"qty":1,"price":1137,"fragile":false},{"grams":1677,"qty":1,"price":3509,"fragile":false}],"express":true,"coupon":"SHIP10"} {"country":"FR","items":[{"grams":455,"qty":1,"price":5500,"fragile":false}]} {"country":"IT","items":[{"grams":1172,"qty":4,"price":5010,"fragile":false}],"coupon":"SHIP10"} {"country":"CA","items":[{"grams":859,"qty":4,"price":5732,"fragile":false},{"grams":942,"qty":1,"price":1440,"fragile":false},{"grams":211,"qty":1,"price":5200,"fragile":false},{"grams":998,"qty":2,"price":376,"fragile":true}]} {"country":"US","items":[{"grams":960,"qty":1,"price":11700,"fragile":false}]} {"country":"IT","items":[{"grams":1417,"qty":1,"price":8393,"fragile":false},{"grams":690,"qty":4,"price":8131,"fragile":false},{"grams":1198,"qty":1,"price":5701,"fragile":false},{"grams":1432,"qty":3,"price":1849,"fragile":false}]} {"country":"JP","items":[{"grams":1542,"qty":1,"price":17100,"fragile":false}]} {"country":"ZA","items":[{"grams":709,"qty":1,"price":7601,"fragile":true},{"grams":283,"qty":1,"price":8762,"fragile":false},{"grams":1447,"qty":3,"price":5768,"fragile":true},{"grams":639,"qty":2,"price":1695,"fragile":false}],"express":true} {"country":"GB","items":[{"grams":1600,"qty":1,"price":1388,"fragile":false}]} {"country":"US","items":[{"grams":479,"qty":1,"price":7336,"fragile":false},{"grams":284,"qty":5,"price":6311,"fragile":true},{"grams":1424,"qty":2,"price":4453,"fragile":false},{"grams":409,"qty":2,"price":3446,"fragile":false}],"express":true} {"country":"ES","items":[{"grams":597,"qty":3,"price":1070,"fragile":false},{"grams":1308,"qty":5,"price":7989,"fragile":false},{"grams":1022,"qty":1,"price":5186,"fragile":false},{"grams":532,"qty":3,"price":1010,"fragile":false}]} {"country":"CA","items":[{"grams":782,"qty":5,"price":7357,"fragile":false},{"grams":977,"qty":2,"price":8267,"fragile":false},{"grams":482,"qty":3,"price":7279,"fragile":false}],"express":true} {"country":"US","items":[{"grams":959,"qty":1,"price":11700,"fragile":false}]} {"country":"AU","items":[{"grams":1604,"qty":3,"price":7802,"fragile":false},{"grams":1503,"qty":1,"price":6012,"fragile":false},{"grams":553,"qty":2,"price":8350,"fragile":false},{"grams":1222,"qty":5,"price":6170,"fragile":true}],"coupon":"SHIP10"}answer
answer hidden on shared reportsgrader · score 1
20/20 outputs match
agent's debrief
Bug was 'value <= FREE_BASE_OVER[zone]' adding the base fee at exactly the waiver threshold; the comment says the value waives the base, so it must be strict '<'. Reproduced 1798 pre-fix, 522 post-fix, then ran the fixed function on all 20 orders.
implement-1✓ passbatched
prompt
Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[39,47],[9,13],[5,13],[17,23],[31,35],[7,12]] [[21,21],[24,31],[12,12],[3,7]] [[39,40],[4,9],[40,40],[40,41],[38,43],[8,12],[23,26]] [[39,46],[11,12],[37,41],[16,17],[14,17],[19,19],[25,32],[7,10]] [[17,25],[38,40],[4,7],[16,17],[8,11]] [[38,41],[4,9],[36,37],[11,19],[3,6],[25,26]] [[12,15],[12,13],[5,12],[31,35],[34,35],[0,5],[11,18]] [[21,28],[6,8],[8,13],[15,22],[37,43],[25,29]]answer
answer hidden on shared reportsgrader · score 1
12/12 lines match
agent's debrief
Standard merge-overlapping-intervals: sort, extend on overlap or touch. Verified two of the tricky cases by hand (touching at the endpoint merges, a 1-gap does not).
repo-1✓ pass47s
prompt
Download airbench.ai/f/c83a8c050735e1c82be05b601a21e631.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
The overdraft check used 'bal <= 0' but the rule (and the failing unit test) says below zero, so exactly-zero withdrawals must not be charged. Fixed to '<', sample matched the README checksum, all tests pass.
repo-2✓ passbatched
prompt
Download airbench.ai/f/c5ad2b52e8733a7ce1b304b9dae03eac.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Two bugs: load() sorted DD/MM/YYYY dates as raw strings (day-major lexicographic order) instead of the (year,month,day) key, and the same overdraft '<=' mistake. Fixed both; sample matched the README value and tests pass.
how this agent was configured
Hardware: NVIDIA RTX 5090 32 GB (SM120), x86_64, driver 595.84, Ubuntu 24.04 (host hal5090). Model server: Checkpoint gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090 (NVFP4, no MTP head). vLLM 0.27.1 (vllm/vllm-openai:v0.27.1): --quantization modelopt --kv-cache-dtype fp8 --trust-remote-code --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_xml --max-model-len 131072 --max-num-seqs 4 --gpu-memory-utilization 0.95. ~79 tok/s single-stream decode. Harness: omp 18.4.2 (oh-my-pi, @oh-my-pi/pi-coding-agent) in a container (oven/bun:1): `omp -p --mode json --auto-approve <prompt>`; per-run PI_CODING_AGENT_DIR models.yml, provider api openai-completions, compat supportsDeveloperRole=false, supportsReasoningEffort=false; context 131072, max output 16384 tokens. Orchestrator: github.com/dh7/agent-checkup-benchmark @ b8916ff; the checkup instructions are passed verbatim as a single prompt, one attempt, no time limit unless noted. Operator limits: stopped by the operator at 127 min (120-min cap).