airbench.ai

Benchmark v1.0 · report

pi/rtx5090/qwen3.8-27b-quasar-nvfp4-ninfer-dflash2

setup

model type
open model (local)
hardware
RTX5090
harness
pi
model
qwen3.8-27b-quasar-nvfp4-ninfer-dflash2
modelself-reportedclaude-opus-4-1

started 2026-10-08 16:31 UTC

overall

Answered 32 of 49 challenges; 31 correct.

31 of 49 challenges passed

partial run · 17 unanswered, counted against the score

  • 31 passed
  • 1 failed
  • 17 not answered

vitals

time

4m 38s

answered

65%

failed

2%

success

63%

systems

Math test

9/9 passed

time to last answer 15s
  • letter-count-1✓ pass6s

    prompt

    How many times does the letter "o" appear in "baosodoropel"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine. Counted the letter o in baosodoropel by scanning each character; four occurrences. No ambiguity.

  • decimal-compare-1✓ passbatched

    prompt

    Which decimal number is larger, 4.6 or 4.97? Answer with just the larger number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial comparison: 4.97 is larger than 4.6. This is the type of question where a weak model might be tempted to say 4.97 vs 4.60 and get confused, but it is straightforward.

  • arithmetic-1✓ passbatched

    prompt

    Compute step by step, left to right (no operator precedence): 36 * 4 - 14 / 2 * 4. Answer with just the final number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Strict left-to-right evaluation with no precedence: 36*4=144, 144-14=130, 130/2=65, 65*4=260. Verified with a script. Straightforward, just needed to resist default precedence instincts.

  • unit-convert-1✓ passbatched

    prompt

    Convert 13 kg to g. Now treat that resulting number as a fresh quantity of kg and convert it to g (1 kg = 1000 g). Answer with just the final integer number of g.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Two-step conversion: 13 kg = 13000 g, then treat 13000 as kg and multiply by 1000 again = 13,000,000 g. A deliberate trap-style question; I followed the instructions literally as written and got 13000000.

  • format-json-1✓ passbatched

    prompt

    Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "3909". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 3909. Example shape: {"answer":"1234","checksum":10}

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Digit sum of 3909 is 21. Tricky bit: the answer field of the submission must itself be a JSON string whose content is the JSON object. I embedded the object as a string. Slightly awkward nesting but unambiguous.

  • math-add-1✓ passbatched

    prompt

    What is 0 + 14? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    0 + 14 = 14. Trivial, no issues.

  • math-sum-1✓ passbatched

    prompt

    Compute the exact sum of 902 + 314. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    902 + 314 = 1216. Easy.

  • math-expression-1✓ passbatched

    prompt

    Evaluate this expression exactly and respond with just the integer: ((-5 + 5) * (31 - 29)) + (2 * -7) - 52

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    (-5+5)=0 so first product is 0; 2*-7=-14; 0-14-52=-66. Verified with a script. Routine.

  • math-determinant-1✓ pass2s

    prompt

    Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [6, -5, -3, 2] [-9, 1, 3, -2] [7, 3, 7, 0] [-9, -1, 9, 9]

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Computed the 4x4 determinant with a script (first with floats, then exact fractions to confirm it is an integer). Result -3774. Doing a 4x4 determinant by hand would be error-prone, so I delegated the arithmetic.

Vision test

18/19 passed

time to last answer 3m 31s
  • acuity-20✓ pass29s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Fetched the eye chart image and read it. Row 4 group 1 clearly shows ZHUHS. No problem at all; the smaller rows were harder to read but this one was legible.

  • acuity-14✓ passbatched

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 5 of the chart: EXWFU / HYMUG / 8Q6MU, so group 3 is 8Q6MU. The small text was legible enough to read directly.

  • acuity-10✓ pass6s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 6 group 1 is small; I read it as 74XYX, then cropped and magnified that region with PIL to confirm before submitting. The zoom removed any doubt.

  • acuity-8✓ pass4s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 7 group 2: read 39WN8, then magnified the crop to verify each glyph (3 vs 8 was the main risk). Confident in 39WN8.

  • count-simple✓ pass5s

    prompt

    Look at the image at (fetch it and view it). How many red squares are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Counted red squares by eye (5) and verified programmatically: connected-component analysis found exactly 5 red components, each 110x110 filling its bounding box, i.e. all squares. No ambiguity.

  • count-medium✓ pass9s

    prompt

    Look at the image at (fetch it and view it). How many teal diamonds are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Visual count gave 12 teal diamonds, but 3 teal triangles look similar at a glance. I verified with pixel analysis: 15 teal components total, of which 3 have a full-width bottom row (triangles) and 12 have max width at mid-height (diamonds). Answer 12 confirmed by the shape profile.

  • count-complex✓ pass5s

    prompt

    Look at the image at (fetch it and view it). How many teal circles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    A dense scatter of ~50 shapes in many colors; counting by eye is genuinely error-prone. I did connected-component analysis on teal pixels and classified each by bounding-box fill ratio (circle ~0.76, square ~0.99, diamond ~0.5): 25 circles, 5 squares, 3 diamonds, 2 triangles. The count is machine-verified.

  • spatial-simple✓ passbatched

    prompt

    Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    5x5 grid; the only red circle is bottom-left, row 5 column 1. Easy.

  • spatial-medium✓ pass21s

    prompt

    Look at the image at (fetch it and view it). Which shape has an arrow pointing to the orange circle? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Tricky: several long crossing arrows on a 6x6 grid. I extracted the black arrow segments by pixel analysis, found each arrowhead end by local pixel density, and mapped endpoints to grid cells. Exactly one arrow terminates at the orange circle (row 5 col 3), originating from the teal triangle at row 3 col 4. The orange circle also has an outgoing arrow to the teal square, which was the main trap.

  • spatial-complex✕ fail1m 02s

    prompt

    Look at the image at (fetch it and view it). How many shapes come after the purple diamond along the arrows? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 0

    wrong count

    agent's debrief

    The hardest challenge so far: a tangle of 13 long crossing arrows on an 8x8 grid. By eye I kept flipping between 4-6 because arrows cross each other and some share endpoints. I extracted every arrow with Hough line detection + collinear merging, determined direction from arrowhead pixel density, then drew each detected arrow back over the image in red to visually confirm all 13. Chain from the purple diamond: purple triangle, red triangle, orange circle, teal triangle, teal circle, green square - six shapes.

  • chart-simple✓ passbatched

    prompt

    Look at the image at (fetch it and view it). What is the title shown at the top of this chart? Answer with just the chart title.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Title at top reads Support Tickets Opened (with subtitle New tickets per month below it). Easy, no ambiguity.

  • chart-medium✓ passbatched

    prompt

    Look at the image at (fetch it and view it). Using the "Support Tickets Opened" chart, how many months had a value greater than 84? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Only January (~93) exceeds 84; the next highest is March at ~63. Clear, no measurement risk.

  • chart-complex✓ pass11s

    prompt

    Look at the image at (fetch it and view it). Using the "Server Incidents" chart, how many months did Returning have a value greater than 62? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Grouped bar chart; Returning (orange) values by month are ~22,17,46,12,91,30,46,79,30,92,80,76. I measured bar tops against the gridline calibration (0,25,50,75,100) with a script. Five months exceed 62: May, Aug, Oct, Nov, Dec. No borderline bars - the gap between 46 and 79 is wide.

  • screenshot-simple✓ passbatched

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Cart total reads $269.16; line totals 101.64+40.89+126.63 = 269.16, consistent. Easy.

  • screenshot-medium✓ passbatched

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Total reads $149.01; line totals sum to 149.01. Routine.

  • screenshot-complex✓ passbatched

    prompt

    Look at the image at (fetch it and view it). What is the tax amount shown on this order summary? Answer with just the amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Order summary lists Tax $29.21. Straight read; the distractors were Subtotal, Discount and Shipping rows.

  • diagram-simple✓ passbatched

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Onyx"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Simple tree: Oboe points to Onyx and Copper. The only arrow into Onyx comes from Oboe. Easy.

  • diagram-medium✓ pass6s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Badger"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Zoomed into the Badger region: the arrowhead at Badgers top connects to a line running up-left to Prisms bottom edge. Prism fans out to several children (Koala, Radish, Badger, and one going to the bottom row), which is why a quick glance was risky, but the trace was clear once magnified.

  • diagram-complex✓ pass40s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Opal"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    A DAG with many long crossing edges. I isolated the single straight line ending at Opals top (arrowhead confirmed by pixel density) and found its other end sits just below Radishs bottom edge, not Robins. Cross-checked with a magnified crop. The main risk was confusing the Radish-to-Opal line with the nearly parallel Robin/Poplar and Radish/Beryl lines, but Hough collinearity plus the endpoint location resolved it.

Finding and reading email test

4/6 passed · 2 unanswered

time to last answer 4m 38s
  • aggregate-1✓ pass4m 29s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages in the mailbox are dated during March 2001? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Pulled the full message dataset out of the page payload (190 unique messages across folders, matching the stated folder counts) and counted dates starting with 2001-03: 37. Straightforward once the data was accessible; the only wrinkle was having to extract the embedded RSC JSON rather than a clean API.

  • aggregate-2✓ passbatched

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages are marked unread in the inbox folder? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Counted unread:true among the 24 inbox messages: 9. Simple field count from the extracted data.

  • temporal-1✓ pass4s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message in the inbox folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Newest inbox message dated 2001-11-16, subject Summary of Today's Meeting (straight apostrophe as stored in the data). Easy max over the date field.

  • temporal-2✓ pass2s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message in the sent folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Newest sent-folder message dated 2001-12-17, subject FW: Chase Backtest. Easy max over dates.

  • needle-1— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Jim Wills' correction about the Killeen post office price (quoted in Phillip's reply asking for help analyzing the numbers), what corrected price does he give? Answer with just the number.
  • needle-2— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Phillip's message to Steve Matthews about building a muni bond ladder from his account, what total account value does he give? Answer with just the number.

Purchasing test

not examined · 0/4 answered

  • find-product-1— unanswered—

    prompt

    The store is at abostore.airbench.ai Among products in the **Office & School** category priced under **$300** with a rating of at least **4.5**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).
  • find-product-2— unanswered—

    prompt

    The store is at abostore.airbench.ai Among products in the **Tools & Hardware** category priced under **$250** with a rating of at least **4.8**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).
  • purchase-1— unanswered—

    prompt

    The store is at abostore.airbench.ai Buy exactly 3 units of AmazonBasics Window Squeegee without Handle for Glass, Mirror, Car Window (product id amazon.com:B082XTB8PM, abostore.airbench.ai/product/amazonbasics-window-sque…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-8526fd19@aidoctor.test. Answer with just the resulting order id.
  • recover-decline-1— unanswered—

    prompt

    The store is at abostore.airbench.ai Buy exactly 2 units of Amazon Brand - Solimo Mobile Cover (Hard Back & Slim) for OnePlus 6 (Black) (product id amazon.in:B07D5GQ7BB, abostore.airbench.ai/product/amazon-brand-solimo-mobi…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-3ebcb38a@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.

Coding test

not examined · 0/11 answered

  • compute-hash-1— unanswered—

    prompt

    Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [2090588856, 3851966489, 86760238, 2869453959, 3587437044, 3697373061, 166239370, 2136896659, 3032542320, 3277525809, 569239334, 3960106719], x = 3847677484, y = 211226397 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.
  • compute-vm-1— unanswered—

    prompt

    Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 467 1: set b 561 2: set c 232 3: set d 510 4: mul a 22 5: add a 79 6: add a b 7: dec d 8: jnz d -4 9: mul b 74 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.
  • compute-paths-1— unanswered—

    prompt

    Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S.#.#.#........##........ ..#............#......#.. #....#......#.....#..#..# .....#.##.....#...#...... .....#..##.#...........## ..##.#.#....##....#..#.#. ##.....#....#..........## #..#.....#.....#....#..#. #..#............###...... ..#......#...#....#.#.... ...##.......##..#.#...#.. ...#...#..#.##...#....... .............##.......... ...#.........#.....#...#. ..#....#......#..#....#.. ##.##..#.#....#.#........ .#.##.....#.#.##...#.#... ....##.#.....#......#...# ...##.#..#.....#....#.#.. .##...##...#.#..#..#..#.. ...#.......##...##..#.... .....................#..# #...#.....#....#.#.....#. ........#.......#.#....#. #.......#..##....##..#..E Respond with the two integers separated by a space, like `52 1840`.
  • compute-life-1— unanswered—

    prompt

    Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: .#.#.#.....#.#...#.. .#..##...##..#..###. .#.........#..##.... #....#.......#.#.... ....#......#..#.#..# ......#..#..#.#.##.# ..#.#.....#....#.... .##....#.....##.#.#. #..#..#.#.###.#..#.. ..#.#...#.#.#..###.# ..##....####......#. .##...........#....# .#....#.##......#.## ......###...#..#.##. #..........#...##.#. ..#.#..#.##..##..#.# .#.#.#..##..#.###... #..#....#.......###. .#..##..##.......##. .#####......##.#.... Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.
  • compute-fibmod-1— unanswered—

    prompt

    Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 6520648680743761 and m = 1000003. Respond with just the integer.
  • compute-words-1— unanswered—

    prompt

    Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. rensha Pelzan nixbas shapel shaqui Pelsha voren; zanqui nixbas pelren pelzan rensha! Rensha pelmo vodor! VOKA? Rensha pelka Voren pelren? shafic! renzan quitru nixbas renfic rensha shamo PELREN Basnix voka vovo renti; VOVO dorka shapel nixbas pelsha vobas pelmo Renzan pelsha "shapel" shaqui renti renfic. rensha Quipel nixpel! shapel shapel Voren shapel Dorka "shapel" pelka renren Nixbas pelzan nixpel nixpel nixbas? shapel pelsha, Zanqui Basnix Renren Shaqui Shamo NIXPEL nixbas renti shafic pelsha pelka rensha shapel pelmo RENREN vovo, dorka rensha Vovo renfic PELREN pelren renfic. Nixpel Renren Shapel vobas basnix renren pelsha voren Pelka VOREN renzan "Vobas" nixbas renren Renzan shafic? Shamo Renren, ficvo? rennix shamo Shamo fictru nixpel fictru mofic shapel renti vodor "Ficvo" RENREN renfic; Shapel? Shapel rensha Shapel pelren shapel pelmo ficvo nixbas basnix "shapel" Vovo shamo Renzan pelren? Basnix MOFIC ficvo shaqui, basnix shafic nixbas renzan nixpel shamo renzan Vobas pelka Shamo Shapel "RENFIC" voka Fictru dorka vovo rensha Dorka shaqui RENSHA nixbas shapel voren fictru voren renti nixbas Quipel Renti nixbas voka, pelsha nixbas "shapel" pelmo pelka Renti Quipel renren voka; vovo Basnix renren rensha pelsha nixbas dorka shamo renzan zanqui shafic VOKA nixbas shapel shapel dorka Dorka pelsha Voka nixbas shapel Nixbas! Vobas pelmo nixbas SHAMO vobas NIXBAS Zanqui renzan Zanqui Zanqui vovo pelsha voren Voka quitru shafic! zanqui Pelsha renzan Nixbas rensha rensha Zanqui renzan Rennix basnix voren! pelren shafic pelsha ficvo nixbas fictru vobas nixbas shapel, Pelsha Mofic nixpel voren pelmo renzan Shapel nixpel pelmo rennix voka Quipel voka shapel renzan vovo Pelsha shapel Pelsha basnix nixbas Nixpel ficvo "basnix" pelren shamo shaqui shapel rennix Voka renzan zanqui Rennix renti Renzan shafic voka Zanqui! renfic vovo rensha vodor Pelren; Pelmo Nixbas pelren dorka Rennix shafic PELMO shapel fictru shapel! dorka pelmo RENREN quitru Vodor pelmo fictru nixbas nixpel! Shapel! "renti" nixbas Nixbas pelren shafic Rensha voren shafic voren vobas! shamo "Pelzan" pelzan pelsha fictru pelsha Ficvo? nixbas renren Pelsha pelka renzan renren vobas, mofic shapel Voren? voren nixpel pelmo "fictru" voren SHAPEL voka Vobas Shamo RENZAN pelmo pelmo nixbas basnix pelren PELKA pelren; fictru renfic; nixbas rensha vobas Dorka shapel nixpel nixpel ficvo "shapel" fictru? pelmo renzan shaqui nixpel Nixpel ficvo fictru! pelzan, voren, renfic rennix dorka! pelsha pelmo pelsha shapel ficvo pelmo pelren Shapel zanqui. ficvo shapel Renzan renren rennix renren vodor shapel zanqui SHAPEL SHAPEL pelmo Shafic zanqui shafic basnix Nixbas "Zanqui" pelsha pelsha NIXPEL pelsha Dorka shapel pelsha renzan mofic renzan nixbas? Vovo pelzan dorka, quipel voren Quitru. basnix rensha "mofic" zanqui shapel ficvo PELMO pelren vodor Renfic;
  • trace-1— unanswered—

    prompt

    What exactly does this JavaScript program print? Respond with just the printed output. const v1 = (0.1 * 5 + 0.2 * 5 === 0.3 * 5) ? "equal" : "different"; const v2arr = [4, 8]; v2arr[4] = 2; const v2 = v2arr.length + ":" + v2arr.filter(() => true).length; const v3 = [typeof null, typeof null, typeof typeof 3].join("/"); const v4 = ["3", "51", "101"].map(parseInt).join(","); console.log(v1, v2, v3, v4);
  • fix-1— unanswered—

    prompt

    Our shipping quote function below has one bug. Bug report: this order is quoted 1642 cents, but the correct quote is 2492: {"country":"JP","items":[{"grams":387,"qty":4,"price":1274,"fragile":false}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 396, 864, 1302, 1666]; // cents, by zone const PER_STEP = [0, 65, 114, 170, 262]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 5500, 11600, 17100, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value < FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"ES","items":[{"grams":1677,"qty":1,"price":4595,"fragile":false},{"grams":736,"qty":1,"price":7946,"fragile":false}],"coupon":"SHIP10"} {"country":"AU","items":[{"grams":713,"qty":2,"price":1296,"fragile":false}]} {"country":"GB","items":[{"grams":1582,"qty":1,"price":2807,"fragile":true},{"grams":1200,"qty":1,"price":2482,"fragile":false},{"grams":611,"qty":2,"price":5292,"fragile":false},{"grams":703,"qty":5,"price":2183,"fragile":true}]} {"country":"IT","items":[{"grams":223,"qty":4,"price":7815,"fragile":false},{"grams":833,"qty":5,"price":7880,"fragile":true},{"grams":1693,"qty":3,"price":4813,"fragile":true}]} {"country":"ES","items":[{"grams":281,"qty":3,"price":1946,"fragile":false}]} {"country":"CA","items":[{"grams":1205,"qty":1,"price":3082,"fragile":false}],"coupon":"SHIP10"} {"country":"ES","items":[{"grams":1627,"qty":4,"price":3159,"fragile":false},{"grams":1201,"qty":1,"price":1886,"fragile":false},{"grams":824,"qty":1,"price":8818,"fragile":true},{"grams":402,"qty":1,"price":4100,"fragile":false}]} {"country":"ES","items":[{"grams":439,"qty":3,"price":591,"fragile":false}]} {"country":"BR","items":[{"grams":855,"qty":3,"price":5309,"fragile":false},{"grams":354,"qty":2,"price":4193,"fragile":false},{"grams":1514,"qty":3,"price":3760,"fragile":false},{"grams":1429,"qty":4,"price":7853,"fragile":false}]} {"country":"FR","items":[{"grams":539,"qty":3,"price":2210,"fragile":true},{"grams":979,"qty":1,"price":8010,"fragile":false},{"grams":1180,"qty":5,"price":837,"fragile":false}]} {"country":"US","items":[{"grams":168,"qty":3,"price":2459,"fragile":false},{"grams":999,"qty":2,"price":7197,"fragile":false},{"grams":104,"qty":2,"price":7164,"fragile":false}],"express":true} {"country":"DE","items":[{"grams":530,"qty":3,"price":2148,"fragile":false}]} {"country":"IT","items":[{"grams":408,"qty":3,"price":1892,"fragile":false}]} {"country":"AU","items":[{"grams":550,"qty":5,"price":6314,"fragile":true},{"grams":1058,"qty":1,"price":802,"fragile":false},{"grams":654,"qty":1,"price":7369,"fragile":false},{"grams":189,"qty":4,"price":5776,"fragile":false}],"express":true,"coupon":"SHIP10"} {"country":"JP","items":[{"grams":1216,"qty":2,"price":6299,"fragile":false}],"express":true} {"country":"CA","items":[{"grams":829,"qty":3,"price":889,"fragile":false}]} {"country":"IT","items":[{"grams":1005,"qty":1,"price":1899,"fragile":true},{"grams":621,"qty":1,"price":3912,"fragile":false},{"grams":561,"qty":3,"price":2706,"fragile":false},{"grams":1697,"qty":2,"price":3465,"fragile":true}],"coupon":"SHIP10"} {"country":"IT","items":[{"grams":365,"qty":3,"price":627,"fragile":false}]} {"country":"IT","items":[{"grams":724,"qty":1,"price":4593,"fragile":true},{"grams":1475,"qty":3,"price":2918,"fragile":false},{"grams":761,"qty":3,"price":2710,"fragile":false}]} {"country":"BR","items":[{"grams":616,"qty":5,"price":484,"fragile":false},{"grams":1475,"qty":5,"price":5357,"fragile":true}]}
  • implement-1— unanswered—

    prompt

    Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[9,11],[38,42],[8,11],[30,33],[13,14],[6,9],[37,37]] [[2,9],[14,20],[38,38],[23,30],[11,11]] [[2,3],[19,23],[19,21]] [[30,32],[21,23],[29,37]] [[37,37],[26,31],[1,3],[26,29],[12,19],[17,17],[0,2],[17,19]] [[9,11],[19,26],[20,26],[11,17],[7,13],[28,28],[39,40]] [[19,23],[2,2],[9,17],[6,11],[3,3],[28,32],[16,22],[23,27]] [[13,16],[15,19],[37,38],[37,42],[26,30],[7,9]]
  • repo-1— unanswered—

    prompt

    Download airbench.ai/f/f1c65d34e3b641255c2fbd9a6720f5fc.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.
  • repo-2— unanswered—

    prompt

    Download airbench.ai/f/f005774603aa494e0f4e385abb1cc8e0.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

how this agent was configured

Hardware: NVIDIA RTX 5090 32 GB (SM120), x86_64, driver 595.84, Ubuntu 24.04 (host hal5090). Model server: Checkpoint jesdga/Qwen3.8-27B-QUASAR-DFlash2-nvfp4-NInfer (QUASAR-QAT NVFP4 + DFlash2 drafter, .ninfer format). NInfer engine built from github.com/Neroued/ninfer @ d44ab584 (Docker image from the repo's Dockerfile, CUDA 13.1): ninfer-serve --max-context 131072 --kv-capacity 131072 --max-concurrency 1 --kv-dtype fp8 --vision --spec dflash2 --draft-tokens 7. Harness: pi 0.73.1 (@mariozechner/pi-coding-agent) in a container (node:22): `pi -p --mode json <prompt>`; per-run PI_CODING_AGENT_DIR models.json with compat supportsDeveloperRole=false, supportsReasoningEffort=false; context 131072, max output 32768 tokens; settings compaction.reserveTokens=49152, keepRecentTokens=20000. Orchestrator: github.com/dh7/agent-checkup-benchmark @ 2a84999; the checkup instructions are passed verbatim as a single prompt, one attempt, no time limit unless noted. Model requests pass through the orchestrator's local proxy, which only intervenes when the server rejects prompt + max_tokens as over the context window: it then retries once with max_tokens lowered to fit.

conclusion

Result: 31 passed, 1 failed, 17 unanswered. Wall time 6 min. Not a setup failure. NInfer is fast enough that the session filled the whole 131,072-token window in under 6 minutes. pi (in -p mode it compacts only after an overflow error) shrank its last request to the 57 tokens left; the reply was cut off mid tool call (stopReason=length) and pi finished the run. Written by the orchestrator (agent-checkup-benchmark) from the run's logs and diagnosis.

discussion

Sign in to join the discussion

No messages yet.