Benchmark v1.0 · report
dsh/rtx5090/qwen3.8-27b-quasar-nvfp4-mtp
setup
- model type
- open model (local)
- hardware
- RTX5090
- harness
- dsh
- model
- qwen3.8-27b-quasar-nvfp4-mtp
started 2026-10-08 13:09 UTC
overall
Answered 26 of 49 challenges; 26 correct.
26 of 49 challenges passed
partial run · 23 unanswered, counted against the score
- 26 passed
- 23 not answered
vitals
time
1h 55m
answered
53%
failed
0%
success
53%
systems
Math test
9/9 passed
letter-count-1✓ pass2m 20s
prompt
How many times does the letter "c" appear in "zancmofcic"? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Counted the c in zancmofcic: 3. Trivial, verified by hand.
decimal-compare-1✓ pass10s
prompt
Which decimal number is larger, 1.5 or 1.91? Answer with just the larger number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Direct comparison: 1.91 > 1.5. No trickery here; answered straight away.
arithmetic-1✓ pass50s
prompt
Compute step by step, left to right (no operator precedence): 36 + 10 - 9 - 19 - 14. Answer with just the final number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Left-to-right: 36+10=46, -9=37, -19=18, -14=4. Done mentally, no ambiguity since precedence was explicitly disabled.
unit-convert-1✓ pass5s
prompt
Convert 15 km to m. Now treat that resulting number as a fresh quantity of hours and convert it to minutes (1 hours = 60 minutes). Answer with just the final integer number of minutes.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
15 km = 15000 m; then 15000 hours x 60 = 900000 minutes. Straight two-step conversion.
format-json-1✓ pass5s
prompt
Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "6105". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 6105. Example shape: {"answer":"1234","checksum":10}answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Exact JSON with key order answer,checksum; digit sum of 6105 is 12, emitted as a number. Formatted by hand and verified the checksum is a JSON number, not a string.
math-add-1✓ pass5s
prompt
What is 8 + 15? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
8+15=23, mental arithmetic, no ambiguity.
math-sum-1✓ pass5s
prompt
Compute the exact sum of 920 + 440. Respond with just the integer.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
920+440 = 1360, simple addition.
math-expression-1✓ pass5s
prompt
Evaluate this expression exactly and respond with just the integer: ((20 + 7) * (12 - 26)) + (8 * -2) - 19answer
answer hidden on shared reportsgrader · score 1
agent's debrief
(27 * -14) + (-16) - 19 = -378 - 16 - 19 = -413. Computed in Python as a safeguard against sign slips.
math-determinant-1✓ pass5s
prompt
Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [10, 1, -8, -8] [-3, 12, -6, -2] [-3, -3, -5, 7] [-9, 0, 6, 13]answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Expanded the 4x4 determinant two independent ways in Python (Leibniz permutation sum and exact-fraction Gaussian elimination); both agree on -6699.
Vision test
6/19 passed · 13 unanswered
acuity-20✓ pass1h 27m
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Read the eye chart directly; row 4 group 1 rendered too small to read reliably, so I located the row band and group columns by pixel analysis, then cropped and upscaled 10x to read 7TEBW unambiguously.
acuity-14✓ pass5m 52s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Located the row band and group columns via pixel analysis, cropped and upscaled 8x; row 5 group 3 reads 3UKWG.
acuity-10✓ pass4m 29s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Located the row band and group columns via pixel analysis, cropped and upscaled 12x; row 6 group 2 reads BY8GW.
acuity-8✓ pass5m 21s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Row 7 is only ~6px tall, so I cropped via pixel-located bands and upscaled both LANCZOS and NEAREST; nearest-neighbor inspection distinguished F vs P/E and 8 vs B. Row 7 group 1 reads FTVR8.
count-simple✓ pass4m 44s
prompt
Look at the image at (fetch it and view it). How many orange squares are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Segmented exact orange color (242,106,34), ran 4-neighbor connected components: 6 shapes, all 110x110 with fill ratio 0.998 (squares); total orange pixels 72480 = 6 x 12080 exactly, so no missed or overlapping shapes.
count-medium✓ pass7m 07s
prompt
Look at the image at (fetch it and view it). How many teal triangles are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Connected components over both teal shades: 19 shapes total. Fill ratio suggested 16 triangles, but two of those were diamonds with identical 0.5 fill; row-width profiles confirmed (triangle widens monotonically to bottom, diamond narrows again). True teal triangles = 14, cross-checked by hand on the full image.
count-complex— unanswered—
prompt
Look at the image at (fetch it and view it). How many orange squares are in the image? Answer with just the number.spatial-simple— unanswered—
prompt
Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).spatial-medium— unanswered—
prompt
Look at the image at (fetch it and view it). Which shape does the arrow from the green diamond lead to? Answer with just the colour and shape, e.g. teal diamond.spatial-complex— unanswered—
prompt
Look at the image at (fetch it and view it). How many shapes come after the red square along the arrows? Answer with just the number.chart-simple— unanswered—
prompt
Look at the image at (fetch it and view it). Using the "Units Shipped" chart, approximately what value did Feb have? Read it off the y-axis; answers within +/-5 are accepted.chart-medium— unanswered—
prompt
Look at the image at (fetch it and view it). Using the "New Signups" chart, approximately what value did Aug have? Read it off the y-axis; answers within +/-5 are accepted.chart-complex— unanswered—
prompt
Look at the image at (fetch it and view it). Using the "Support Tickets Opened" chart, approximately what value did Paid have in Mar? Read it off the y-axis; answers within +/-3 are accepted.screenshot-simple— unanswered—
prompt
Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.screenshot-medium— unanswered—
prompt
Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.screenshot-complex— unanswered—
prompt
Look at the image at (fetch it and view it). What is the tax amount shown on this order summary? Answer with just the amount, e.g. $12.34.diagram-simple— unanswered—
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Osprey"? Answer with just the box name, e.g. Kettle.diagram-medium— unanswered—
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Meadow"? Answer with just the box name, e.g. Kettle.diagram-complex— unanswered—
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Condor" point to? Answer with just the box name, e.g. Kettle.
Finding and reading email test
not examined · 0/6 answered
aggregate-1— unanswered—
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages in the mailbox are dated during December 2001? Answer with just the number.aggregate-2— unanswered—
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages are in the trash folder? Answer with just the number.temporal-1— unanswered—
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the sent folder? Answer with just the subject line, exactly as shown.temporal-2— unanswered—
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message in the sent folder? Answer with just the subject line, exactly as shown.needle-1— unanswered—
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Julieta Sandoval's message about the Muni Bond Ladder, what direct phone number does she give? Answer with just the exact fact as it appears in the message, and nothing else.needle-2— unanswered—
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Renee's reply to Phillip about the Deferred Phantom Stock Units payout, how many shares does she say are reflected on the 9/30/01 statement? Answer with just the number.
Purchasing test
not examined · 0/4 answered
find-product-1— unanswered—
prompt
The store is at abostore.airbench.ai Among products in the **Fashion** category priced at or above **$200** with a rating of at least **3.6**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).find-product-2— unanswered—
prompt
The store is at abostore.airbench.ai Among products in the **Office & School** category priced at or above **$250** with a rating of at least **3.5**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).purchase-1— unanswered—
prompt
The store is at abostore.airbench.ai Buy exactly 2 units of AmazonBasics Smoke and Odor Eliminator Refill, 1-Gallon, 4-Pack (product id amazon.ca:B07CVSMKPG, abostore.airbench.ai/product/amazonbasics-smoke-and-o…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-2be08635@aidoctor.test. Answer with just the resulting order id.recover-decline-1— unanswered—
prompt
The store is at abostore.airbench.ai Buy exactly 1 unit of AmazonBasics PlayStation 4 Controller Thumb Grips - 2-Pack, Black (product id amazon.ca:B07JHYQ45L, abostore.airbench.ai/product/amazonbasics-playstation…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-041100b2@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.
Coding test
11/11 passed
compute-hash-1✓ pass24m 49s
prompt
Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [3763530015, 3356996972, 3660019549, 2218527810, 1029612971, 1949329000, 902357897, 3742685534, 1385669751, 2522304164, 61947893, 911270842], x = 2911228291, y = 3717745184 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Simulated the 25000-step 32-bit loop exactly as specified (unsigned 32-bit, imul low-32, rotl) in Python, then cross-checked with an independent Node.js implementation; both agreed.
compute-vm-1✓ pass2m 13s
prompt
Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 372 1: set b 229 2: set c 226 3: set d 585 4: add a 77 5: mul b 70 6: add a b 7: dec d 8: jnz d -4 9: mul b 29 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Simulated the 13-line VM in Python (two nested loops: 226 outer x 585 inner iterations), and cross-checked with an independent Node.js simulation of the same semantics; both produced a=786802.
compute-paths-1✓ pass2m 15s
prompt
Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S...###.#.#......###..... .....#.#.#......#....##.# ..#.#.#..#......#.#...#.. ..#.##.##.....##.#.#.#... #........#......#.....#.. .#...###...#.#...#.#...## .....#.....#.#......#..#. ###....#..##...........#. ...#..#.........###.#..#. ........#....#.....##.... ........#.......#..#...#. ..###................#.#. ...#.....#.....#.#.#..##. ..##...##.............##. ...#..####.####...#.#.#.. ....##.........#......... .............#.#..#.#.... ##.#.#..##..#..#...#..... .#..#......#.........#.#. ..#....................#. ..#.........#..##........ ..#..#.....##.#....##.... ....###.#...#...##......# #.#...#........#........# .#####..#.....#.##..#.#.E Respond with the two integers separated by a space, like `52 1840`.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Ran BFS from S on the 25x25 grid for shortest distance (48) and counted shortest paths two independent ways (forward layer accumulation and dist-ordered DP); both gave 80233536 mod 1e9+7.
compute-life-1✓ pass1m 22s
prompt
Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: #.#...#.......#.#... ...#..##.#..##...#.. #..##.##.....##...#. .##..#.#...#.##....# #..##.....#...#.#... ..#..........#.#...# ...##...#..#.#.#.... ##.#.#.##..##...#.#. .....###.#.##..#..## .#.....##......#.... #...###..#..##..##.. ..#.#..#........#... ###....##.#.....#.## ...##.......####.... #.#.....#..#..#.#..# .#.#........#.##...# #.#.#.....#......#.. #.#.......#.#.#..#.. .....#...###.#.....# ###.##....#.....#.#. Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Simulated Conway Game of Life on the 20x20 torus for exactly 150 generations with wraparound neighbor counts; 14 cells survive and their row*20+col sum is 1700.
compute-fibmod-1✓ pass2m 17s
prompt
Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 5795090167643502 and m = 1000003. Respond with just the integer.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Used fast doubling to compute F(5795090167643502) mod 1000003; also cross-checked a few small n values against naive Fibonacci modulo as a sanity check of the implementation.
compute-words-1✓ pass1m 36s
prompt
Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. baslu Vozan "kafic" titi pelvo volu quibas nixren luqui titru pelvo; zandor baslu pelbas baslu; Pelka Vozan baslu titru Nixren Shavo volu Baslu quibas dorfic baska Shavo baslu Quitru NIXPEL Baslu pelvo nixpel Pelvo "pelvo" titru Pelka nixren Titru baslu; pelka shador! baslu? nixtru quibas titru; kafic shavo? vopel baska Shavo dorfic dorfic shavo, Pelka Shavo shavo shavo Baslu baslu! titru Dorfic truzan quitru shavo Titi kati kafic. voti kafic vozan titru baslu dorren shador vozan volu titru baslu QUIBAS; shavo voti Pelvo Pelvo renfic Dorfic "vopel" vozan volu Pelvo volu! vopel voti "voti" Baslu. dorvo Dorvo PELVO basqui; pelvo dorren. nixpel shavo Vozan "pelvo" dorfic! pelvo titru! quitru nixren baslu "basqui" kafic voti nixren baslu; baska. vozan Vozan volu! pelvo Basqui dorvo Nixdor baslu quitru baslu Titru baslu pelbas shavo baslu shador? volu kafic, pelbas Renfic Baska vozan dorfic dorfic basqui; truzan baslu volu! basqui pelka nixren truzan kati, pelvo renfic trusha pelvo; "Pelka" Kati "quitru" nixtru nixdor Baslu BASLU nixdor titru titru quitru pelvo quibas. Dorvo voti volu Kafic Shavo Dorren kafic dorfic nixtru Dorren nixren Kafic Titru! "pelka" titru nixtru basqui Shavo, baslu pelka, nixren shavo renfic! dorren volu Vopel luqui shavo quibas! nixdor vozan nixpel? vopel baslu Quibas "Basqui" nixdor vozan BASQUI, Quitru kafic baslu baslu quitru basqui Pelvo nixpel nixtru Baslu nixpel pelka; kafic? nixpel truzan kafic nixdor dorren baska shavo quitru Basqui luqui? kafic nixren quibas "quitru" "quibas" pelvo "volu" shador kati nixtru shador nixtru baslu renfic Quitru! quibas QUIBAS baslu pelbas, nixdor Baslu titru nixren truzan titru pelka titi Vozan, titru BASLU titi nixtru nixdor titru baslu! Kafic quitru Truzan BASLU, nixdor; pelka pelvo, BASLU baslu volu Pelka baslu Kafic baska voti! shador dorren vopel Vopel, shador vozan vopel shavo Kafic kati PELKA, Truzan vopel titru kafic Volu "shavo" basqui shavo Pelka kati dorren Baslu shavo voti dorfic Vozan TITRU truzan nixren vopel. pelbas renfic baslu? titru Quibas pelka luqui. BASLU kafic Pelvo titru; Baslu pelbas luqui kafic pelvo Kafic titru baslu Dorfic truzan volu pelka baslu. renfic kafic nixren baslu baslu trusha Voti titi kafic pelvo? vozan Kati nixpel Dorfic baslu titru nixpel? nixpel pelbas titi, nixtru renfic luqui "truzan" shavo QUITRU pelbas dorren "nixpel" baslu shavo shador kati Nixtru Pelvo titi nixdor pelvo baslu? pelvo dorvo titru baslu basqui shavo vozan kafic Shador basqui Quitru baslu Baslu ZANDOR baska pelbas shador. voti pelvo pelvo volu? nixren nixren Nixren trusha Shavo luqui baslu basqui kafic; baslu zandor shavo trusha Volu; titru Baska Quitru baslu nixpel Nixren kati pelbas trusha Baslu titi Basluanswer
answer hidden on shared reportsgrader · score 1
agent's debrief
Lowercased the text, stripped attached punctuation/quotes per token, counted with a Counter; shavo and titru tied at 26 so shavo wins the 3rd slot alphabetically.
trace-1✓ pass2m 09s
prompt
What exactly does this JavaScript program print? Respond with just the printed output. const v1arr = [9, 7]; v1arr[7] = 6; const v1 = v1arr.length + ":" + v1arr.filter(() => true).length; const v2 = ["2", "98", "10"].map(parseInt).join(","); const v3 = [28, 5, 522, 1406].sort().join(","); const v4 = "2" + 5 - 9 + "9"; console.log(v1, v2, v3, v4);answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Ran the exact program in Node.js v22: sparse array keeps length 8 but filter yields 3 elements; map(parseInt) gives 2,98,NaN (10 parsed as decimal in the 2nd slot... index 2 gets radix 10 -> 10? no: parseInt("10",10)=10 is wrong — actually the third call is parseInt("10",2)=NaN); sort is lexicographic; string coercion gives 169.
fix-1✓ pass1m 46s
prompt
Our shipping quote function below has one bug. Bug report: this order is quoted 1757 cents, but the correct quote is 380: {"country":"AU","items":[{"grams":469,"qty":1,"price":19900,"fragile":false}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 393, 756, 1377, 1693]; // cents, by zone const PER_STEP = [0, 86, 123, 190, 297]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 5200, 9000, 19900, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams * item.qty; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value <= FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"MX","items":[{"grams":628,"qty":1,"price":5400,"fragile":true},{"grams":342,"qty":4,"price":7402,"fragile":false},{"grams":967,"qty":4,"price":2689,"fragile":false}]} {"country":"JP","items":[{"grams":269,"qty":4,"price":3308,"fragile":false},{"grams":1449,"qty":1,"price":6932,"fragile":true},{"grams":821,"qty":1,"price":5687,"fragile":false}],"coupon":"SHIP10"} {"country":"ES","items":[{"grams":390,"qty":3,"price":8921,"fragile":true},{"grams":249,"qty":1,"price":1574,"fragile":true}]} {"country":"GB","items":[{"grams":1657,"qty":4,"price":4134,"fragile":false},{"grams":398,"qty":3,"price":2868,"fragile":false},{"grams":572,"qty":2,"price":2407,"fragile":false}]} {"country":"ES","items":[{"grams":1515,"qty":1,"price":5200,"fragile":false}]} {"country":"JP","items":[{"grams":289,"qty":1,"price":19900,"fragile":false}]} {"country":"ES","items":[{"grams":501,"qty":3,"price":7707,"fragile":false}],"express":true} {"country":"GB","items":[{"grams":813,"qty":1,"price":9000,"fragile":false}]} {"country":"AU","items":[{"grams":123,"qty":1,"price":19900,"fragile":false}]} {"country":"DE","items":[{"grams":1585,"qty":1,"price":8584,"fragile":true}]} {"country":"CA","items":[{"grams":92,"qty":5,"price":2985,"fragile":false},{"grams":414,"qty":2,"price":6813,"fragile":false},{"grams":829,"qty":3,"price":4713,"fragile":false},{"grams":1496,"qty":1,"price":8832,"fragile":false}]} {"country":"FR","items":[{"grams":1717,"qty":1,"price":5200,"fragile":false}]} {"country":"GB","items":[{"grams":1178,"qty":1,"price":5377,"fragile":false}],"express":true,"coupon":"SHIP10"} {"country":"GB","items":[{"grams":363,"qty":2,"price":7734,"fragile":false},{"grams":602,"qty":5,"price":6352,"fragile":false},{"grams":863,"qty":3,"price":8980,"fragile":true}],"express":true,"coupon":"SHIP10"} {"country":"MX","items":[{"grams":135,"qty":1,"price":7838,"fragile":false},{"grams":524,"qty":1,"price":7072,"fragile":false},{"grams":431,"qty":1,"price":5555,"fragile":false},{"grams":889,"qty":1,"price":3546,"fragile":true}],"coupon":"SHIP10"} {"country":"AU","items":[{"grams":371,"qty":1,"price":5966,"fragile":true},{"grams":612,"qty":1,"price":5032,"fragile":false}],"express":true} {"country":"ES","items":[{"grams":1067,"qty":2,"price":7758,"fragile":false},{"grams":498,"qty":3,"price":7360,"fragile":false}],"coupon":"SHIP10"} {"country":"US","items":[{"grams":1376,"qty":1,"price":9000,"fragile":false}]} {"country":"AU","items":[{"grams":330,"qty":1,"price":19900,"fragile":false}]} {"country":"GB","items":[{"grams":255,"qty":5,"price":6840,"fragile":false}]}answer
answer hidden on shared reportsgrader · score 1
20/20 outputs match
agent's debrief
The bug was the free-base threshold comparison: value <= FREE_BASE_OVER[zone] wrongly charged the base fee exactly AT the threshold (test order 19900 <= 19900 -> 1757 instead of 380). Changed <= to < and ran the corrected quote() over all 20 orders, hand-verifying several results by hand.
implement-1✓ pass42s
prompt
Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[12,17],[0,5],[31,38],[14,15]] [[22,25],[30,37],[22,25]] [[30,35],[24,25],[28,31],[31,34],[25,26]] [[11,15],[21,28],[38,39],[23,29],[39,42]] [[31,39],[10,15],[6,13],[4,8],[5,5],[21,25]] [[10,10],[6,14],[19,19],[32,37]] [[23,25],[38,46],[8,11]] [[9,16],[35,42],[36,41],[20,26],[14,19],[0,5],[1,4],[20,22]]answer
answer hidden on shared reportsgrader · score 1
12/12 lines match
agent's debrief
Implemented mergeIntervals (sort by start, merge when next.start <= current.end so touching intervals merge) and ran all 12 inputs; hand-verified the tricky cases (touching [1,3]+[3,5], non-touching [1,2]+[3,4], containment, duplicates).
repo-1✓ pass14m 33s
prompt
Download airbench.ai/f/1078bcecb5ca4c00fa1a97f6149322c0.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Bug: fees.py used cents > 50000 for the 1% withdrawal fee, but the README and the unit test require withdrawals of exactly 50000 to also pay 1%; changed > to >=. All 7 tests pass and sample.csv prints the documented 4e742d95.
repo-2✓ pass8m 59s
prompt
Download airbench.ai/f/cba19d180cd8db3e3f5b52910d429421.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Two bugs: fees.py computed the 1% fee with ceiling division -(-cents//100) instead of floor cents//100 (caught by the rounds-down unit test), and parse.py sorted by (date, amount) instead of date alone, breaking the documented file-order tie-break for same-date transactions (caught by the sample checksum mismatch). Both fixed; all 7 tests pass and sample.csv prints the documented 0f244af5.
how this agent was configured
Hardware: NVIDIA RTX 5090 32 GB (SM120), x86_64, driver 595.84, Ubuntu 24.04 (host hal5090). Model server: Checkpoint QUASAR-QAT/Qwen3.8-27B-QUASAR-NVFP4 (quantization-aware-trained NVFP4, compressed-tensors, MTP head kept). vLLM 0.27.1 (vllm/vllm-openai:v0.27.1): --kv-cache-dtype fp8 --trust-remote-code --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_xml --max-model-len 131072 --max-num-seqs 4 --gpu-memory-utilization 0.95 --speculative-config '{"method":"mtp","num_speculative_tokens":2}'. Harness: dsh 0.2.0-rc.2 (DeepSeek Harness, @deepseek-ai/dsh) in a container (node:22): `dsh --profile headless --patch <file> --json <prompt>` with DSH_PERMISSION_MODE=danger-full-access and DSH_TELEMETRY_MODE=DISABLED; the patch adds one pi-ai openai-completions route; context 131072, max output 32768 tokens and makes it the default model, same overlay as gx10's dsh agent; everything else, compaction included, is the headless profile's default. Orchestrator: github.com/dh7/agent-checkup-benchmark @ 2a84999; the checkup instructions are passed verbatim as a single prompt, one attempt, no time limit unless noted. Model requests pass through the orchestrator's local proxy, which only intervenes when the server rejects prompt + max_tokens as over the context window: it then retries once with max_tokens lowered to fit.
conclusion
Result: 26 passed, 0 failed, 23 unanswered (math 9/9, vision 6/19, mail 0/6, purchasing 0/4, coding 11/11). Wall time 120 min. Stopped by the operator at the 120-minute cap (the benchmark's time limit) while still working. On the 26 challenges it did answer, 26 were right (100%); the rest were left unanswered. Written by the orchestrator (agent-checkup-benchmark) from the run's logs.
discussion
Sign in to join the discussion
No messages yet.