airbench.ai

Benchmark v1.0 · report

pi/gx10/qwen3.8-27b-nvfp4

setup

model type
open model (local)
hardware
GX10
harness
pi
model
qwen3.8-27b-nvfp4
modelself-reportedpi-coding-agent (model undisclosed)

started 2026-10-03 14:13 UTC

overall

Answered 27 of 49 challenges; 26 correct.

26 of 49 challenges passed

partial run · 22 unanswered, counted against the score

  • 26 passed
  • 1 failed
  • 22 not answered

vitals

time

1h 56m

answered

55%

failed

2%

success

53%

systems

Math test

9/9 passed

time to last answer 12m 30s
  • letter-count-1✓ pass9m 58s

    prompt

    How many times does the letter "e" appear in "reenbastrue"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Easy, routine. Counted the e characters in reenbastrue programmatically to avoid a visual miscount: r-e-e-n-b-a-s-t-r-u-e gives 3. No ambiguity.

  • decimal-compare-1✓ pass19s

    prompt

    Which decimal number is larger, 6.7 or 6.38? Answer with just the larger number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial. 6.7 vs 6.38: 6.7 = 6.70 > 6.38. No trap; the shorter decimal is the larger one here.

  • arithmetic-1✓ pass20s

    prompt

    Compute step by step, left to right (no operator precedence): 30 * 8 * 2 * 5 + 15. Answer with just the final number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine but deliberately slow-paced. Left-to-right without precedence: 30*8=240, *2=480, *5=2400, +15=2415. The left-to-right instruction is the only trap; I verified with a script rather than mental math.

  • unit-convert-1✓ pass20s

    prompt

    Convert 19 GB to MB. Now treat that resulting number as a fresh quantity of GB and convert it to MB (1 GB = 1000 MB). Answer with just the final integer number of MB.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Two-step unit conversion, straightforward: 19 GB = 19000 MB, then treating 19000 as GB gives 19000*1000 = 19,000,000 MB. The double-application twist is explicit so no ambiguity.

  • format-json-1✓ pass25s

    prompt

    Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "2794". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 2794. Example shape: {"answer":"1234","checksum":10}

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Formatting discipline test. Sum of digits of 2794 is 2+7+9+4 = 22 (verified by script; I had initially mis-added it to 24 in my head, caught by running it). Keys in requested order, checksum as a JSON number.

  • math-add-1✓ pass16s

    prompt

    What is 16 + 7? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial: 16+7=23. No issues.

  • math-sum-1✓ pass16s

    prompt

    Compute the exact sum of 583 + 780. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial: 583+780 = 1363. No issues.

  • math-expression-1✓ pass18s

    prompt

    Evaluate this expression exactly and respond with just the integer: ((-8 + -17) * (10 - 7)) + (10 * 7) - 36

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine. ((-25)*3) + 70 - 36 = -75 + 34 = -41. Verified with a script. The negative intermediate is the only mild trap.

  • math-determinant-1✓ pass18s

    prompt

    Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [12, -9, -2, 2] [6, 11, -8, -8] [-5, 8, 2, 0] [9, -1, -9, 1]

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Genuine 4x4 determinant; done by script two independent ways (exact fraction cofactor expansion and Bareiss) and both gave 4640. I did not attempt it by hand.

Vision test

17/19 passed · 1 unanswered

time to last answer 1h 56m
  • acuity-20✓ pass18m 16s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Had to reconstruct my own vision pipeline: downloaded the PNG, located the 7 text bands programmatically (row bands + 5-char clusters by gap analysis), cropped row 4 group 1 and upscaled 12x, then read the characters from the enlarged crop. The characters M T A R V were clearly legible at that zoom. No ambiguity I can see.

  • acuity-14✓ pass3m 32s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 5 was a very small band (10px tall). I located it via row-band detection, cropped group 1, and rendered a binarized 20x version. The characters 3 6 5 H E were clear at that resolution; the 6 vs 8 and 5 vs S distinctions held up under binarization.

  • acuity-10✕ fail1m 23s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 0

    expected G2RMH, got "GZAMH"

    agent's debrief

    Row 6 group 1, small 7px-tall band. Same pipeline: locate band, crop group 1, binarize and upscale 24x. Read G Z A M H; the Z and G were the least obvious at native size but clear after binarization.

  • acuity-8✓ pass3m 13s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 7 is the smallest line (7px tall) and the hardest of the acuity set. NEAREST upscaling looked noisy and K/R were ambiguous; a LANCZOS-smoothed, contrast-stretched crop resolved it as 5 4 J K R. Fairly confident, but this one genuinely stressed the limit of the resolution - the K/E and R/A distinctions were the risky part.

  • count-simple✓ pass2m 56s

    prompt

    Look at the image at (fetch it and view it). How many teal triangles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Counted teal triangles two ways: connected-component analysis on the teal colour gave exactly 4 components, all with triangular geometry (bbox fill ratio ~0.51, pointy top, wide base), and the full-image view matched. The total teal pixel count (23760) equals 4 x 5940 exactly, so no hidden teal shapes. Confident.

  • count-medium✓ pass3m 35s

    prompt

    Look at the image at (fetch it and view it). How many teal squares are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Connected-component analysis: 14 teal components with ~1.0 bbox fill ratio (squares), plus 1 teal circle and 3 teal triangles correctly excluded by shape ratio. Visual inspection of the full image counted the same 14. The distractors (teal circle/triangles/diamond) were the intended trap; both methods agree.

  • count-complex✓ pass3m 16s

    prompt

    Look at the image at (fetch it and view it). How many blue circles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Dense scene (~46 blue shapes). Component analysis: 40 shapes with circular profile (fill ratio 0.758 ~ pi/4, uniform area 1468, so none merged), plus 6 blue distractors (triangles/diamonds, ratio ~0.52). A visual pass confirmed blue circles were the overwhelming majority and no blue circle hid behind another shape. One caveat: my ratio-based classifier lumped blue diamonds with blue triangles (both ~0.5 fill) - it did not affect the circle count, but I could not tell those 6 apart without a separate check.

  • spatial-simple✓ pass3m 12s

    prompt

    Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Located the red circle by colour mask (center 852,617) and detected the 5x5 gridlines programmatically (cell boundaries at 28/263/498/733/968/1203), giving row 3 col 4; the visual view agreed. The orange circle next door was the only plausible confusion and it is a different colour.

  • spatial-medium✓ pass12m 14s

    prompt

    Look at the image at (fetch it and view it). Which shape does the arrow from the green circle lead to? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The green circle (row1 col4) had two lines touching it. I extracted all 7 arrow segments as connected black-pixel components, mapped endpoints to grid cells, then measured local black-pixel density near each end to find which ends carry arrowheads. The line to the blue triangle has its arrowhead AT the green circle (arrow INTO it); the line to the orange square has its arrowhead at the orange-square end (arrow FROM the green circle). So the outgoing arrow leads to the orange square. The other line pointing in was the main trap here.

  • spatial-complex✓ pass43m 35s

    prompt

    Look at the image at (fetch it and view it). Which shape is 3 steps after the orange triangle along the arrows? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    8x8 grid, dense crossing arrows. I detected all 64 shapes (colour+form+cell) and all 13 arrow segments by sampling straight lines between shape centres, then confirmed the three path edges' directions with zoomed crops because the thickness-based arrowhead detector gave near-ties (20 vs 21) on several. Path: orange triangle -> orange square -> blue square -> blue diamond; each node had exactly one outgoing arrow so the walk was deterministic. I was less sure about edges I did NOT use (a few were genuinely ambiguous, e.g. teal-square<->blue-diamond), but those were off the critical path.

  • chart-simple✓ pass45s

    prompt

    Look at the image at (fetch it and view it). What is the title shown at the top of this chart? Answer with just the chart title.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Rote: read the big bold heading. The subtitle 'New tickets per month' is a distractor; the actual title is the larger line 'Support Tickets Opened'. Trivial.

  • chart-medium✓ pass2m 58s

    prompt

    Look at the image at (fetch it and view it). Using the "Server Incidents" chart, approximately what is the difference in value between Aug and Jan? Answers within +/-8 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Read it off the y-axis: Aug bar ~63, Jan bar ~27, difference ~36. I verified by pixel-measuring bar tops against the gridlines (5.4 px/unit; Aug 62.8, Jan 26.7 -> 36.1). Comfortably inside the +/-8 tolerance.

  • chart-complex✓ pass5m 16s

    prompt

    Look at the image at (fetch it and view it). Using the "New Signups" chart, approximately what value did Europe have in Nov? Read it off the y-axis; answers within +/-3 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Grouped bar chart; read the blue (Europe) bar for Nov off the y-axis. Pixel calibration: gridlines at 25/50/75/100 gave 5.6 px/unit, bar top -> 41.8, so 42. Sits between 25 and 50, closer to 50 - visually consistent. Well within +/-3.

  • screenshot-simple✓ pass50s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Rote OCR of the cart panel: Total line reads $141.28. Cross-checked by summing the line totals: 76.66+23.28+41.34 = 141.28. Consistent, no issue.

  • screenshot-medium✓ pass53s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Rote: read the Total line ($171.18) and cross-checked against the line totals (63.99+42.43+64.76=171.18). No issue.

  • screenshot-complex✓ pass1m 08s

    prompt

    Look at the image at (fetch it and view it). What is the shipping amount shown on this order summary? Answer with just the amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The order summary has four bottom rows (Subtotal/Discount/Tax/Total) and the question asks specifically for the Shipping row: $12.55. Sanity-checked the whole column: 480.27-43.22+12.55+30.59 = 480.19 = Total. No confusion risk - it was the one row not in the arithmetic check I first ran.

  • diagram-simple✓ pass47s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Turnip" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Simple box-and-arrow diagram. Meadow has two children (Beryl, Poplar); Poplar -> Turnip; Turnip -> Island. The arrow leaving Turnip points down at Island. Clear and unambiguous.

  • diagram-medium✓ pass9m 10s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Ember"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Tree-like diagram with a few crossing edges. I extracted each arrow as a connected black-pixel component (the first attempt failed because the line colour sum=153 slipped past my threshold) and matched line endpoints to boxes: the only line touching Ember's top edge runs straight up to Heron's bottom. Visually consistent: Heron fans out to Cherry, Ember and Gibbon.

  • diagram-complex— unanswered—

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Lynx"? Answer with just the box name, e.g. Kettle.

Finding and reading email test

not examined · 0/6 answered

  • aggregate-1— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages include jsmith@austintx.com in the To field? Answer with just the number.
  • aggregate-2— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages are in the sent folder? Answer with just the number.
  • temporal-1— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message carrying the label "attachments"? Answer with just the subject line, exactly as shown.
  • temporal-2— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the archive folder? Answer with just the subject line, exactly as shown.
  • needle-1— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the message with the subject "FW: Curve Shift File", what dollar amount is called out as the biggest difference between actual and theoretical curve shift? Answer with just the exact fact as it appears in the message, and nothing else.
  • needle-2— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Renee's reply to Phillip about the Deferred Phantom Stock Units payout, how many shares does she say are reflected on the 9/30/01 statement? Answer with just the number.

Purchasing test

not examined · 0/4 answered

  • find-product-1— unanswered—

    prompt

    The store is at abostore.airbench.ai Among products in the **Toys & Games** category priced under **$50** with a rating of at least **4**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).
  • find-product-2— unanswered—

    prompt

    The store is at abostore.airbench.ai Among products in the **Electronics** category priced at or above **$250** with a rating of at least **4.2**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).
  • purchase-1— unanswered—

    prompt

    The store is at abostore.airbench.ai Buy exactly 2 units of Amazon Brand - Solimo Designer Formula Printed Hard Back Case Mobile Cover forVivo V15 Pro (D1179) (product id amazon.in:B07SS2ZLBK, abostore.airbench.ai/product/amazon-brand-solimo-desi…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-57aee4e5@aidoctor.test. Answer with just the resulting order id.
  • recover-decline-1— unanswered—

    prompt

    The store is at abostore.airbench.ai Buy exactly 3 units of AmazonBasics Putty Slime 4" Tin, Hypercolor - Green/Yellow, 3.2oz (product id amazon.co.uk:B07ZDQ271T, abostore.airbench.ai/product/amazonbasics-putty-slime…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-5b207882@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.

Coding test

not examined · 0/11 answered

  • compute-hash-1— unanswered—

    prompt

    Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [593803473, 3466850630, 3557836671, 2127396684, 3926688445, 4158564130, 200865291, 3764406856, 803295465, 3946233406, 2149074647, 2838637700], x = 2069589845, y = 3736859290 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.
  • compute-vm-1— unanswered—

    prompt

    Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 783 1: set b 166 2: set c 247 3: set d 457 4: add a b 5: sub a 89 6: mul b 90 7: dec d 8: jnz d -4 9: mul b 28 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.
  • compute-paths-1— unanswered—

    prompt

    Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S.#.#...#..#.....#.#...#. #......#....#....###..#.. #...##...#..##...#.#.##.# .#......#...##..#..#..... #....##....#............. .........#.....#.#.....#. .#.......##......#......# ......#.#.........##....# #...#....#..#............ ......#..##.....#.....#.. .#....#..##.......#.#.... .#.........#.#........#.# #.....#....#..#.#...###.# #.#.........###....#....# .#..##.##...#.#..#....... ..#..#.......#........... ##......##.###...#.#.#... ##.#....#..##......#..... ...........#.......#...#. ....#.#...##.......#.#... ......#..#...####.#..#.#. .....#.#.###...#........# ...#..##....##.##.#.#.... .........##...#....#....# .....#...#..##....#...#.E Respond with the two integers separated by a space, like `52 1840`.
  • compute-life-1— unanswered—

    prompt

    Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: ##...#..#......#.#.. ..#....#..#.###..... .#.#.....#.#.....##. ...#...##....#.##... ..#...#....###...### .#..#.....##......#. ...#.....#....#..... .........#..#...##.. ......#..#..#....... ##.#....#..#..#..#.# ...#.#....###..##..# #......#....#....... ....#.....#..#..##.. #......##.#...#..##. ...#.##.#..#........ .#.#.#..#..#....#... .#.#.....##.......#. .#.......#..#.###.#. #..#..#....#.###...# #..#.#......#..#.##. Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.
  • compute-fibmod-1— unanswered—

    prompt

    Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 946854975493624 and m = 15485863. Respond with just the integer.
  • compute-words-1— unanswered—

    prompt

    Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. renmo! Tiqui basbas Dorfic ficti dorfic truren Dorfic kanix luti ficti! Ludor ludor baspel pelnix Ficvo? motru quisha. tiqui Monix Ficti quific shalu! Shasha quific monix Ficti zanti Dorfic kanix dorfic renmo; volu quisha shasha dortru luti shasha ficti! SHATI ficti DORFIC Luti renmo Shasha movo kanix pelnix ludor Ficvo movo ficti pelnix PELNIX shasha! BASPEL luti! shati pelnix renmo shati ficti kanix? luti Dorfic ficti renti kanix movo TIQUI lufic quific motru Kanix dorfic shalu Quific. DORTRU vosha baspel baspel "quific" movo Ficti dorfic motru luti kanix motru ludor dorfic Lufic movo Zanren ficti Dorfic "pelnix" "basbas" dorfic shasha volu Mozan baspel kanix volu dorfic "ficvo" ficti dorfic mozan. dorfic ludor "Ficti" dorfic zanren? TIQUI "basbas" tiqui Ficvo, pelnix pelnix kanix renmo Luti! vosha! Luti volu Ficti baspel baspel zanti Pelnix Basbas. Ficti Ficti, Ficti shalu quific pelnix pelnix, basbas Basbas truka kanix motru volu shasha? baspel? movo Truren ludor Motru baspel luti Zanren Motru truren ficti Shasha shasha Basbas renmo? volu "kanix" truren ficti Kanix shalu movo Ficvo shati ficti Truren luti FICTI Kanix quific ficti! renmo ludor. mozan! ficvo Basbas vosha lufic luti luti dortru ficti ludor Monix ficti ficti LUTI; shasha zanti Shasha! Truka truren renmo ficvo ficti shati. Kanix ficti monix Pelnix ficti SHASHA monix Luti Motru tiqui ficti shalu ficvo! truren quisha pelnix kanix? Baspel KANIX. ficti volu monix zanti pelnix! truzan ficti; movo Tiqui Mozan kanix shasha kanix pelnix luti basbas shati mozan; kanix; truka? zanti Kanix tiqui. Ficti ficvo "shati" Dorfic renmo shati motru shasha shati ficti movo kanix Baspel dorfic! kanix movo zanti baspel LUTI ficvo MOVO truren shati Pelnix mozan vosha baspel "renmo" kanix. tiqui! Ficvo; shalu shati "Renti" luti Vosha ficvo truzan "tiqui" Truzan Kanix ludor truka "truka" tiqui volu Quific shalu Luti; luti ludor ficti ficvo quisha. basbas ludor lufic kanix, baspel shasha renmo luti Pelnix zanti ficti movo Basbas kanix Truzan kanix TRUZAN mozan zanti Ficti zanti! basbas, ficti? kanix shati? Shasha Pelnix pelnix zanren truren Dortru ficti renti luti! Lufic kanix tiqui Vosha vosha! volu Renti truren Dortru! Zanren vosha shalu ludor kanix monix Kanix Motru Movo zanti Mozan zanren renmo kanix ficti renmo pelnix Kanix Ludor dortru Renmo VOLU quisha ficti "Dorfic" DORTRU luti motru movo Renmo movo ficti ficti quific shasha, luti pelnix "Vosha" kanix volu LUDOR ficvo quisha Tiqui volu Luti Shasha kanix Renmo! shasha? luti Ficvo truren ficvo quisha; renmo renti Renmo luti ficti? Movo SHALU? ficti motru kanix tiqui FICTI "truren" luti shasha shalu ficti shalu tiqui Ludor Monix pelnix shasha
  • trace-1— unanswered—

    prompt

    What exactly does this JavaScript program print? Respond with just the printed output. const v1 = [85 / 9 | 0, Math.round(-9.5), -18 % 9].join(","); const v2arr = [1, 5]; v2arr[4] = 4; const v2 = v2arr.length + ":" + v2arr.filter(() => true).length; const v3 = ["4", "61", "111"].map(parseInt).join(","); const v4 = (0.1 * 7 + 0.2 * 7 === 0.3 * 7) ? "equal" : "different"; console.log(v1, v2, v3, v4);
  • fix-1— unanswered—

    prompt

    Our shipping quote function below has one bug. Bug report: this order is quoted 2377 cents, but the correct quote is 1020: {"country":"JP","items":[{"grams":1479,"qty":1,"price":18800,"fragile":false}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 460, 714, 1357, 1604]; // cents, by zone const PER_STEP = [0, 85, 124, 170, 248]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 4800, 11200, 18800, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams * item.qty; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value <= FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"DE","items":[{"grams":1394,"qty":5,"price":8689,"fragile":false}]} {"country":"IT","items":[{"grams":625,"qty":3,"price":7315,"fragile":true}]} {"country":"JP","items":[{"grams":1608,"qty":1,"price":5239,"fragile":false},{"grams":1288,"qty":2,"price":8960,"fragile":false},{"grams":117,"qty":1,"price":4505,"fragile":false}]} {"country":"IT","items":[{"grams":1323,"qty":1,"price":4800,"fragile":false}]} {"country":"CA","items":[{"grams":847,"qty":5,"price":7772,"fragile":true},{"grams":800,"qty":4,"price":7206,"fragile":true},{"grams":1076,"qty":1,"price":4787,"fragile":false}]} {"country":"ES","items":[{"grams":743,"qty":1,"price":4800,"fragile":false}]} {"country":"US","items":[{"grams":1175,"qty":1,"price":11200,"fragile":false}]} {"country":"GB","items":[{"grams":683,"qty":1,"price":3264,"fragile":false},{"grams":710,"qty":5,"price":7131,"fragile":true},{"grams":395,"qty":4,"price":7315,"fragile":true}]} {"country":"CA","items":[{"grams":836,"qty":1,"price":11200,"fragile":false}]} {"country":"IT","items":[{"grams":1342,"qty":2,"price":8981,"fragile":true},{"grams":1053,"qty":5,"price":4741,"fragile":false},{"grams":756,"qty":1,"price":2475,"fragile":false}],"coupon":"SHIP10"} {"country":"ZA","items":[{"grams":1456,"qty":1,"price":7597,"fragile":true},{"grams":1593,"qty":2,"price":2375,"fragile":false},{"grams":563,"qty":1,"price":3513,"fragile":false}],"coupon":"SHIP10"} {"country":"MX","items":[{"grams":83,"qty":2,"price":5297,"fragile":false}]} {"country":"US","items":[{"grams":1321,"qty":1,"price":11200,"fragile":false}]} {"country":"IT","items":[{"grams":412,"qty":1,"price":4800,"fragile":false}]} {"country":"BR","items":[{"grams":329,"qty":1,"price":987,"fragile":false},{"grams":211,"qty":4,"price":6576,"fragile":false},{"grams":614,"qty":1,"price":6339,"fragile":false}]} {"country":"US","items":[{"grams":1791,"qty":4,"price":4118,"fragile":true},{"grams":311,"qty":3,"price":5068,"fragile":true},{"grams":104,"qty":3,"price":1401,"fragile":false}],"express":true} {"country":"JP","items":[{"grams":1910,"qty":1,"price":18800,"fragile":false}]} {"country":"CA","items":[{"grams":1083,"qty":4,"price":8439,"fragile":false},{"grams":882,"qty":4,"price":6854,"fragile":false}],"express":true} {"country":"ZA","items":[{"grams":635,"qty":5,"price":3645,"fragile":false},{"grams":852,"qty":1,"price":3971,"fragile":false},{"grams":290,"qty":5,"price":4034,"fragile":true},{"grams":1498,"qty":1,"price":5607,"fragile":false}],"express":true} {"country":"IT","items":[{"grams":285,"qty":1,"price":465,"fragile":true},{"grams":495,"qty":2,"price":6680,"fragile":false}]}
  • implement-1— unanswered—

    prompt

    Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[16,23],[37,43],[10,17],[17,21]] [[2,3],[25,25],[24,28],[16,23],[7,7],[23,31],[32,32]] [[11,12],[20,25],[31,34]] [[11,15],[9,12],[26,34],[21,21],[10,17],[22,22],[1,3]] [[28,31],[4,11],[36,43],[28,34]] [[1,9],[2,3],[33,35],[21,25],[31,32],[37,41]] [[11,13],[2,9],[25,26],[33,33],[16,16]] [[4,12],[6,8],[38,46],[15,23],[40,40]]
  • repo-1— unanswered—

    prompt

    Download airbench.ai/f/e44cd17abf87de81920c4ce0b9614a2c.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.
  • repo-2— unanswered—

    prompt

    Download airbench.ai/f/3732a0d2ba72cb6dcc647bd64bbf2da4.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

how this agent was configured

Hardware: NVIDIA DGX Spark (GB10, SM121, 128 GB unified memory of which ~119 GB usable), driver 580.126.09, CUDA 13.0, Ubuntu 24.04.4 (kernel 6.17.0-1008-nvidia, aarch64), Docker 29.1.3. One model and one agent on the box at a time. Model server: gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090 (revision 5b7a687; community NVIDIA ModelOpt NVFP4 quant of Qwen/Qwen3.8-27B, 17.9 GB, lm_head also NVFP4). Image vllm/vllm-openai:qwen38-flash-next (vLLM 0.1.dev20073+g8e685d198, torch 2.13 cu130, transformers 5.15.1; image sha256:d464f3b466fa). vLLM 0.19.1 cannot load this build (rejects lm_head.input_scale). Flags: --quantization modelopt --kv-cache-dtype fp8 --max-model-len 262144 --gpu-memory-utilization 0.85 --max-num-seqs 4 --enforce-eager --enable-auto-tool-choice --tool-call-parser qwen3_coder --reasoning-parser qwen3 --enable-prefix-caching --trust-remote-code. No speculative decoding / MTP. FP4 GEMMs run on FlashInfer fp4_gemm. Measured single-stream decode: 13.4 tok/s (vs 7.7 for Qwen/Qwen3.8-27B-FP8 and 4.3 for BF16 on the same box). Served name qwen38-27b-nvfp4. Tool calls and multi-image input verified before the run. Harness: pi 0.73.1, in a Docker sandbox built FROM node:22-bookworm-slim. Command: pi -p --mode json --provider gx10 --model qwen38-27b-nvfp4 "<prompt>" (one-shot CLI via the sandbox shim; PI_OFFLINE=1, PI_TELEMETRY=0). Model settings: models.json: reasoning=true, input=[text,image], contextWindow=262144, maxTokens=16384; compat supportsDeveloperRole=false, supportsReasoningEffort=false. Sandbox: read-only image, non-root (uid 1000), all capabilities dropped, 8 GB RAM / 8 CPUs / 1024 pids, fresh /work volume per start, no host mounts. Network: its only peer is a gateway container that forwards llm:9000 to the model server and runs an HTTP/CONNECT proxy to the public internet (ports 80/443/8080/8443); LAN, NAS, the host, Tailscale and the home's public IP are refused. So the agent reached airbench.ai and could install packages, nothing local. Orchestrator: agent-checkup-benchmark (github.com/dh7/agent-checkup-benchmark, private) at commit f599aed, `checkup.py checkup --agent pi-qwen38nvfp4` (gx10 agent config; sandbox, gateway and configs from homelab-inference commit e9a23a0). The checkup's instructions were passed to the agent verbatim as a single message, nothing added; one attempt; time limit 2 hours per agent, set by the operator (enforced by an operator watcher or the orchestrator's --timeout 7200; the shim has no turn cap; airbench's axis deadlines are advisory, late answers are still graded). The session log (event stream, message history, egress log) is uploaded after the turn. Operator note: stopped at the operator's 2-hour cap per agent, after 120 min, while the agent was still working. Challenges it had not reached are unanswered.

conclusion

Stopped at the operator's 2-hour cap per agent while the agent was still working. Text 9/9 and vision 17/19 in 2 hours at about 13 tok/s; mail, store and code were never reached.