airbench.ai

Benchmark v1.0 · report

opencode/openrouter/qwen3.8-max-0902

sharedairbench.ai/checkup/443efa09-d477-4a1d-8973-43a5aba2ea81/report

setup

model type
open model (cloud)
inference provider
openrouter
harness
opencode
model
qwen3.8-max-0902
modelself-reportedqwen3.8-max

started 2026-09-27 21:47 UTC · shared 2026-09-28 05:45 UTC

overall

Answered 36 of 49 challenges; 34 correct.

34 of 49 challenges passed

partial run · 13 unanswered, counted against the score

  • 34 passed
  • 2 failed
  • 13 not answered

vitals

time

18m 42s

answered

73%

failed

4%

success

69%

systems

Math test

9/9 passed

time to last answer 46s
  • letter-count-1✓ pass34s

    prompt

    How many times does the letter "d" appear in "ddordddorti"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial; I counted programmatically to be safe.

  • decimal-compare-1✓ passbatched

    prompt

    Which decimal number is larger, 8.97 or 8.4? Answer with just the larger number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine decimal comparison, no difficulty.

  • arithmetic-1✓ passbatched

    prompt

    Compute step by step, left to right (no operator precedence): 44 * 2 - 6 * 9 / 2. Answer with just the final number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The left-to-right no-precedence rule is unusual but clearly stated; verified with Fraction arithmetic.

  • unit-convert-1✓ passbatched

    prompt

    Convert 19 hours to minutes. Now treat that resulting number as a fresh quantity of hours and convert it to minutes (1 hours = 60 minutes). Answer with just the final integer number of minutes.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Easy; the double conversion wording was clear (19*60 then *60 again).

  • format-json-1✓ passbatched

    prompt

    Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "2438". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 2438. Example shape: {"answer":"1234","checksum":10}

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straightforward formatting task; digit sum 2+4+3+8=17.

  • math-add-1✓ pass10s

    prompt

    What is 1 + 3? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial.

  • math-sum-1✓ passbatched

    prompt

    Compute the exact sum of 712 + 784. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine addition.

  • math-expression-1✓ passbatched

    prompt

    Evaluate this expression exactly and respond with just the integer: ((-10 + -8) * (28 - 11)) + (3 * -9) - 54

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Careful sign handling; verified in Python.

  • math-determinant-1✓ passbatched

    prompt

    Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [2, 5, 2, 4] [1, 10, -6, -1] [2, -4, -1, 2] [-2, -7, 8, 3]

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    My first hand cofactor expansion gave 244, but exact-fraction Gaussian elimination gave -377; I trust the verified computation over my mental arithmetic.

Vision test

18/19 passed

time to last answer 10m 11s
  • acuity-20✓ pass1m 25s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Image rendered clearly; row 4 group 2 was legible at full size. Easy.

  • acuity-14✓ pass7s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 5 group 3 still legible; read QT9VK without strain.

  • acuity-10✓ pass10s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 6 group 3 small but readable: EXQX8. Slight uncertainty on last char (8 vs B) but shape looked like 8.

  • acuity-8✕ fail36s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 0

    expected UM6C5, got "UM6CS"

    agent's debrief

    Row 7 was too small at full view, so I cropped and upscaled 8x with Pillow; confirmed UM6CS.

  • count-simple✓ pass19s

    prompt

    Look at the image at (fetch it and view it). How many green triangles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Counted 4 green triangles visually; scene was sparse and unambiguous.

  • count-medium✓ pass56s

    prompt

    Look at the image at (fetch it and view it). How many green squares are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Visual count matched my connected-component pixel analysis exactly (13 green squares); cross-checking removed doubt.

  • count-complex✓ pass11s

    prompt

    Look at the image at (fetch it and view it). How many blue squares are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    44 blue shapes total; connected-component analysis classified 35 as squares (fill ratio ~0.99), excluding 4 diamonds, 3 triangles, 2 circles. Too dense to trust eyes alone.

  • spatial-simple✓ pass36s

    prompt

    Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Only one red circle in the grid, at row 5 col 3; easy to spot.

  • spatial-medium✓ pass16s

    prompt

    Look at the image at (fetch it and view it). Which shape has an arrow pointing to the purple circle? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Traced arrowheads carefully: the arrow ending at the purple circle starts at the purple square (row1,col3). Several crossing arrows made me double-check head vs tail.

  • spatial-complex✓ pass45s

    prompt

    Look at the image at (fetch it and view it). How many shapes come after the green diamond along the arrows? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Chain from green diamond: -> blue triangle -> green square, then no outgoing arrow. Many crossing arrows elsewhere made me re-trace twice, but the diamond had only one outgoing edge.

  • chart-simple✓ pass28s

    prompt

    Look at the image at (fetch it and view it). What is the title shown at the top of this chart? Answer with just the chart title.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Title text was large and unambiguous: New Signups.

  • chart-medium✓ pass13s

    prompt

    Look at the image at (fetch it and view it). Using the "Server Incidents" chart, approximately what value did May have? Read it off the y-axis; answers within +/-5 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    May bar sits a bit above the midpoint between 20 and 40 gridlines; read ~36. Within +/-5 tolerance either way.

  • chart-complex✓ pass13s

    prompt

    Look at the image at (fetch it and view it). Using the "Server Incidents" chart, how many months did Free have a value greater than 28? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Read Free (blue) bars: only Apr (~22) is below 28; Dec ~34 and Jul ~39 clear it, so 11 of 12 months. Legend check confirmed blue=Free.

  • screenshot-simple✓ pass16s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Total printed clearly and line items sum to it exactly (23.21+121.47+28.78=173.46), nice internal check.

  • screenshot-medium✓ pass11s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Line totals sum exactly to the printed total 419.70, so no ambiguity.

  • screenshot-complex✓ pass14s

    prompt

    Look at the image at (fetch it and view it). What is the tax amount shown on this order summary? Answer with just the amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Tax line read as $34.15; small font but crisp. Cross-check: 632.44-63.24+12.78+34.15=616.13 matches Total.

  • diagram-simple✓ pass12s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Mica"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Simple fan-out diagram; only Lagoon points to Mica.

  • diagram-medium✓ pass40s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Summit" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Three edges converge on Juniper and Silver-to-Pepper crosses Summit-to-Juniper; a 4x zoom crop settled it: Summit edge ends at Juniper lowest arrowhead.

  • diagram-complex✓ pass2m 23s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Osprey" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Dense diagram with two X-crossings near Nebula/Flint; a 4x crop showed Osprey edge goes under Zenith-to-Spruce and ends at Flint, while Tapir feeds Nebula.

Finding and reading email test

5/6 passed

time to last answer 15m 59s
  • aggregate-1✓ pass14m 46s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages in the mailbox are dated during March 2001? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Scraped all 178 message pages (the app only resolves ?id= with the right page param, which cost me a refetch) and counted dates matching Mar 2001 across all folders.

  • aggregate-2✕ fail16s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages include jwills3@swbell.net in the To field? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 0

    agent's debrief

    Three Sent messages address James Wills <jwills3@swbell.net>; parsed To fields from all 178 messages programmatically.

  • temporal-1✓ pass30s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the inbox folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Parsed dates of all 24 inbox messages; oldest is Apr 24 2001. Straightforward once mailbox was scraped.

  • temporal-2✓ pass10s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message in the sent folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Newest of 56 sent messages by parsed timestamp (Dec 17 2001 10:57 PM).

  • needle-1✓ pass9s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the message with the subject "FW: Curve Shift File", what dollar amount is called out as the biggest difference between actual and theoretical curve shift? Answer with just the exact fact as it appears in the message, and nothing else.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Message states the biggest difference is 9/14 ($500 million); quoted verbatim.

  • needle-2✓ pass7s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Phillip's message to gthorse@keyad.com about the Colonial Oaks property, what is the actual NOI for 2001 given for the property? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Colonial Oaks message to gthorse says actual NOI for 2001 is around 305,000 (before his adjustments to 280K/240K); picked the actual figure as asked.

Purchasing test

2/4 passed · 2 unanswered

time to last answer 18m 42s
  • find-product-1✓ pass18m 34s

    prompt

    The store is at abostore.airbench.ai Among products in the **Tools & Hardware** category priced under **$800** with a rating of at least **4.5**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Scraped all 39 Tools & Hardware pages (973 items) pairing RSC-payload ids with card prices/ratings; cheapest qualifying is $9.46 @4.8.

  • find-product-2✓ pass7s

    prompt

    The store is at abostore.airbench.ai Among products in the **Fashion** category priced at or above **$400** with a rating of at least **4**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Scraped all 974 Fashion items; lowest price at or above $400 with rating >=4 is $400.46 @4.0 (a boundary case, but 4.0 meets the threshold).

  • purchase-1— unanswered—

    prompt

    The store is at abostore.airbench.ai Buy exactly 2 units of AmazonBasics College Ruled Wirebound Spiral Notebook 70-Sheet (product id amazon.ae:B07D2M7CD8, abostore.airbench.ai/product/amazonbasics-college-rul…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-549924b8@aidoctor.test. Answer with just the resulting order id.
  • recover-decline-1— unanswered—

    prompt

    The store is at abostore.airbench.ai Buy exactly 2 units of AmazonBasics Assorted Size and Color Rubber Bands, 0.5 lb. (product id amazon.ca:B074B1KCXD, abostore.airbench.ai/product/amazonbasics-assorted-si…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-37934e0f@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.

Coding test

not examined · 0/11 answered

  • compute-hash-1— unanswered—

    prompt

    Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [2819261789, 1048703554, 4125836203, 225516648, 2745340297, 472506206, 2076566135, 4291699876, 1249873397, 650840506, 174305155, 3173728288], x = 46007969, y = 1828089174 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.
  • compute-vm-1— unanswered—

    prompt

    Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 75 1: set b 172 2: set c 285 3: set d 436 4: add b a 5: mul a 27 6: add a b 7: dec d 8: jnz d -4 9: mul a 50 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.
  • compute-paths-1— unanswered—

    prompt

    Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S..........#..#...#..###. ...##....#....#.#..#...#. .#...#.....#..#......#..# ..........#.#..#...#...#. ..##.###.........#.#....# .......#..#....#.......#. ...###...###.#.........#. ##...#...###.##........## ##.#.....#..#..#.#...#... ..#..........#..####....# ....####.#....####.....#. .#.#.#.....#.#..#.#...... ..#.#.#......#......#...# #.#...#...#...#..##.....# #....#......##.#...#...#. #.........###.....#.##... ..#.#......#....###.....# .###.#.#......#.......... ..##.#.###..#...##..#...# ..#.##..#......###.#...#. .#..##.#..#.#.......#...# .#.#..#....#...#.#..#...# ....##.##...#........#... ...##.##..#........#..#.. #...##..............###.E Respond with the two integers separated by a space, like `52 1840`.
  • compute-life-1— unanswered—

    prompt

    Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: ...###........##...# .##.#.#........#.#.. .##.##..##..#.#...#. .##.#...##.##....#.# .#..#........##..### .#..##.#..##......#. #.##.....##.#....#.# ....#....#.....##... .#...#.###.#..####.# .#...##....#..###... .##..#.#...#......#. .#..#..#............ ......#...#.#.##...# #...#..#.....##...#. #..#...#.####....#.# .#..##....#..#.#.... #..##.....#.#.#..#.# .#...###.#.#.#...#.. ##...##.#.#....#.#.# .###..#.........##.# Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.
  • compute-fibmod-1— unanswered—

    prompt

    Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 1652917953347787 and m = 1299709. Respond with just the integer.
  • compute-words-1— unanswered—

    prompt

    Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. kabas shati kasha basnix kabas "Moti" trulu? Kasha trulu quivo PELZAN shador! quivo quivo Kanix Vovo, moti! Quivo movo luvo kasha Luvo dorren tisha Moqui luvo trulu shaka. quivo tidor pelzan tibas Shaka "Basbas" modor; tidor quisha kanix TISHA kanix tisha Quivo LUQUI dorren kasha basbas, tisha Kabas Shaka tibas kasha! Ficti trulu Kasha dorren ficti quivo Tidor "trulu" Trudor, tisha luvo? kasha? shador trulu kasha modor modor kasha moqui kabas shaka quisha zansha moqui quivo kasha MODOR movo vovo modor Luvo. kabas LUVO zansha Moqui quivo kabas quivo. quivo tibas basnix Quisha TISHA? vodor shaka kasha trulu kabas shador DORLU tidor luka? luqui Dorqui shador kasha QUIVO Vodor! tibas quivo kasha, moti kabas Luka moqui dorqui dorren. modor kasha quisha quivo ficti Tibas. kabas dorlu kanix luvo basbas moqui shador Zansha Shati quivo trudor Luka modor zansha Shador basnix zansha shaka luqui kanix quivo moti; trudor Kabas Kasha moti kanix kabas Shador trulu. Vovo trulu moqui luka vodor ficti tibas Trulu Quisha! "Dorlu" kabas. kabas tidor luqui luqui basnix shaka Luka "kasha" kasha? Tisha Trulu dorlu. luka tidor movo? quisha SHADOR kasha; tisha Kabas ficti moti quivo KASHA movo kabas Dorqui Movo kanix LUKA! basbas TIBAS trulu kabas kabas? Trulu Dorlu Quisha tibas vodor movo Kabas modor! quisha TRULU Trulu dorqui quivo kasha. modor tidor Pelzan trulu moti kasha trulu zansha? BASBAS tidor quivo quivo luka tibas Basnix kasha kabas moti shati Luvo, quisha Luka modor moqui KABAS Kasha Trudor basbas luqui luvo. shati; tisha trudor luka kasha ficti Dorlu Tibas movo trulu zansha quivo moti? basbas "luka" kasha moti! Quivo quisha moti TIBAS Luvo Trulu pelzan moqui quisha ZANSHA "moqui" quivo! pelzan moti Pelzan dorren quivo vovo quisha ficti; movo moqui vodor kabas Tibas trudor kasha kasha luka basnix! pelzan tisha basnix tibas trulu luvo luvo Quisha kasha kabas pelzan! quivo dorlu quivo luvo luka kasha? Quisha Ficti Kasha, kasha; luka KABAS luka quivo trudor. dorren quivo Tisha Dorlu kasha Quivo luka zansha; kanix? shati LUVO dorqui Kabas moti Kanix quisha kasha kabas basbas trulu Trudor dorren kanix moqui "trudor" LUKA dorren modor Shati, shati kabas quivo luka Quivo modor Ficti luvo dorren Quivo? kanix zansha pelzan trudor! Tisha trulu vodor kabas quivo tisha zansha zansha moti quisha vodor! movo tisha Basbas basnix; kasha luqui moti ficti Shador Kasha Shador TISHA dorlu "shador" luka quivo modor Shador tidor dorren dorqui; vovo moti. quisha Trudor FICTI ficti Luvo! Kasha quivo kabas vovo dorlu luka basbas dorlu Shati Ficti zansha. basnix luqui modor "Quisha" trulu Quivo Luvo? quisha, basnix movo
  • trace-1— unanswered—

    prompt

    What exactly does this JavaScript program print? Respond with just the printed output. const v1 = "5" + 4 - 4 + "4"; const v2 = ["2", "97", "111"].map(parseInt).join(","); const v3arr = [6, 6]; v3arr[4] = 9; const v3 = v3arr.length + ":" + v3arr.filter(() => true).length; const v4 = ["30" < "4", [] == false, null >= 0].map(Number).join(""); console.log(v1, v2, v3, v4);
  • fix-1— unanswered—

    prompt

    Our shipping quote function below has one bug. Bug report: this order is quoted 178 cents, but the correct quote is 890: {"country":"DE","items":[{"grams":469,"qty":5,"price":1643,"fragile":false}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 401, 799, 1169, 1832]; // cents, by zone const PER_STEP = [0, 89, 124, 230, 259]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 4700, 11000, 17300, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value < FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"ES","items":[{"grams":1595,"qty":1,"price":1870,"fragile":false}],"coupon":"SHIP10"} {"country":"US","items":[{"grams":1440,"qty":4,"price":1678,"fragile":false},{"grams":806,"qty":5,"price":539,"fragile":false},{"grams":951,"qty":1,"price":6405,"fragile":false}],"express":true,"coupon":"SHIP10"} {"country":"MX","items":[{"grams":1701,"qty":2,"price":6576,"fragile":false},{"grams":1136,"qty":3,"price":7946,"fragile":false}],"express":true,"coupon":"SHIP10"} {"country":"AU","items":[{"grams":827,"qty":2,"price":1093,"fragile":false}]} {"country":"NZ","items":[{"grams":843,"qty":1,"price":6135,"fragile":false}]} {"country":"CA","items":[{"grams":552,"qty":1,"price":3610,"fragile":false},{"grams":1211,"qty":1,"price":4451,"fragile":true}]} {"country":"IT","items":[{"grams":1288,"qty":2,"price":1313,"fragile":false}]} {"country":"FR","items":[{"grams":1224,"qty":4,"price":941,"fragile":false},{"grams":1713,"qty":2,"price":5668,"fragile":true}]} {"country":"FR","items":[{"grams":729,"qty":2,"price":545,"fragile":false},{"grams":1477,"qty":5,"price":1028,"fragile":false}],"coupon":"SHIP10"} {"country":"JP","items":[{"grams":537,"qty":2,"price":640,"fragile":false}]} {"country":"FR","items":[{"grams":427,"qty":5,"price":419,"fragile":false}]} {"country":"BR","items":[{"grams":879,"qty":2,"price":2063,"fragile":false}]} {"country":"ES","items":[{"grams":1406,"qty":3,"price":5780,"fragile":false}],"coupon":"SHIP10"} {"country":"CA","items":[{"grams":276,"qty":4,"price":8301,"fragile":false}]} {"country":"JP","items":[{"grams":260,"qty":5,"price":885,"fragile":false}]} {"country":"FR","items":[{"grams":372,"qty":1,"price":1749,"fragile":true},{"grams":1107,"qty":5,"price":3013,"fragile":false},{"grams":1322,"qty":4,"price":306,"fragile":false}]} {"country":"ZA","items":[{"grams":1327,"qty":1,"price":4363,"fragile":true},{"grams":680,"qty":4,"price":3774,"fragile":false}]} {"country":"IT","items":[{"grams":412,"qty":5,"price":2951,"fragile":false}]} {"country":"FR","items":[{"grams":443,"qty":2,"price":437,"fragile":false}]} {"country":"AU","items":[{"grams":525,"qty":1,"price":8444,"fragile":false}],"coupon":"SHIP10"}
  • implement-1— unanswered—

    prompt

    Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[36,38],[21,28],[31,34],[27,28],[24,28],[4,10]] [[39,45],[21,26],[27,31],[1,8],[6,9],[6,7],[26,32],[19,24]] [[3,5],[29,31],[17,25],[35,36],[31,38],[15,22],[39,44],[26,31]] [[26,30],[0,0],[23,27],[19,20]] [[30,33],[16,17],[32,37],[17,17]] [[31,39],[15,19],[2,7],[22,24],[14,19],[9,16],[22,28]] [[39,44],[32,33],[25,30],[5,5],[37,42],[22,22]] [[7,12],[4,5],[25,25],[16,24]]
  • repo-1— unanswered—

    prompt

    Download airbench.ai/f/2983dcf46c3aca6e5da147b8c79cf9ac.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.
  • repo-2— unanswered—

    prompt

    Download airbench.ai/f/59edd168abc8459a7331f7be784296b1.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

how this agent was configured

Hardware: the model runs on OpenRouter; the harness and its sandbox run on a NVIDIA DGX Spark (GB10, SM121, 128 GB unified memory of which ~119 GB usable), driver 580.126.09, CUDA 13.0, Ubuntu 24.04.4 (kernel 6.17.0-1008-nvidia, aarch64), Docker 29.1.3. Model: qwen/qwen3.8-max-0902 on OpenRouter ($2.00/$6.00 per M tokens, 1M context, tools + image input), default OpenRouter provider routing, no sampling or reasoning parameters overridden (each harness sends its own defaults). Reached through the sandbox gateway's LLM forward on llm:9000 (served name qwen3.8-max): it accepts only the sandbox, forwards only /v1/chat/completions and /v1/completions, rewrites every request's model to qwen/qwen3.8-max-0902, adds the OpenRouter key itself (the agent never sees it), and drops a stream that sends nothing for 300 s so the harness can retry. Context window declared to the harnesses: 1,000,000. Harness: opencode 1.18.23, in a Docker sandbox built FROM node:22-bookworm-slim. Command: opencode serve --hostname 0.0.0.0 --port 4096 --pure, driven over its HTTP API (POST /session/{id}/prompt_async, the whole prompt as one turn). Model settings: provider gx10 (@ai-sdk/openai-compatible, baseURL http://llm:9000/v1); model declared attachment=true, modalities.input=[text,image]; permissions edit/bash/webfetch/external_directory = allow; no explicit context or output cap (opencode defaults). Sandbox: read-only image, non-root (uid 1000), all capabilities dropped, 8 GB RAM / 8 CPUs / 1024 pids, fresh /work volume per start, no host mounts. Network: its only peer is a gateway container that forwards llm:9000 to the model server and runs an HTTP/CONNECT proxy to the public internet (ports 80/443/8080/8443); LAN, NAS, the host, Tailscale and the home's public IP are refused. So the agent reached airbench.ai and could install packages, nothing local. Orchestrator: agent-checkup-benchmark (github.com/dh7/agent-checkup-benchmark, private) at commit 9d74217, `checkup.py checkup --agent opencode-qwen38max` (gx10 agent config; sandbox, gateway and configs from homelab-inference commit 79af47a). The checkup's instructions were passed to the agent verbatim as a single message, nothing added; one attempt; time limit 2 hours per agent, set by the operator (enforced by an operator watcher or the orchestrator's --timeout 7200; the shim has no turn cap; airbench's axis deadlines are advisory, late answers are still graded). The session log (event stream, message history, egress log) is uploaded after the turn.