airbench.ai

Benchmark v1.0 · report

opencode/rtx5090/qwen3.8-27b-nvfp4-65k

sharedairbench.ai/checkup/dd5860bd-85c4-44e5-ad39-e4ee7f41c037/report

setup

model type
open model (local)
hardware
RTX5090
harness
opencode
model
qwen3.8-27b-nvfp4-65k
modelself-reportedhal/qwen38-27b-nvfp4

started 2026-09-29 21:03 UTC · shared 2026-09-30 06:33 UTC

overall

Answered 25 of 49 challenges; 22 correct.

22 of 49 challenges passed

partial run · 24 unanswered, counted against the score

  • 22 passed
  • 3 failed
  • 24 not answered

vitals

time

39m 25s

answered

51%

failed

6%

success

45%

systems

Math test

8/9 passed

time to last answer 1m 56s
  • letter-count-1✓ pass21s

    prompt

    How many times does the letter "c" appear in "tcicficcczan"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straightforward character counting. I counted each character one by one and got 5. No ambiguity.

  • decimal-compare-1✓ pass5s

    prompt

    Which decimal number is larger, 5.81 or 5.8? Answer with just the larger number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial. 5.81 has a nonzero hundredths digit so it exceeds 5.8. Easy and routine.

  • arithmetic-1✓ pass17s

    prompt

    Compute step by step, left to right (no operator precedence): 36 + 19 * 7 + 10 / 5. Answer with just the final number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The no-precedence rule is the trap: standard math gives 171, but strict left-to-right gives 55, 385, 395, 79. I followed the stated rule.

  • unit-convert-1✕ fail9s

    prompt

    Convert 18 GB to MB. Now treat that resulting number as a fresh quantity of km and convert it to m (1 km = 1000 m). Answer with just the final integer number of m.

    answer

    answer hidden on shared reports

    grader · score 0

    agent's debrief

    Used binary GB=1024 MB (18432), then 18432 km to m is 18432000. If they meant decimal 1000-based GB it would be 18000000, but GB in this kind of puzzle usually means 1024 MB, and I followed the most common convention.

  • format-json-1✓ pass7s

    prompt

    Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "8802". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 8802. Example shape: {"answer":"1234","checksum":10}

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Key order matters here; I kept answer first, checksum second, checksum as a number. Digit sum is 18. Simple but the ordering constraint is easy to fumble.

  • math-add-1✓ pass4s

    prompt

    What is 0 + 16? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial. Routine.

  • math-sum-1✓ pass5s

    prompt

    Compute the exact sum of 881 + 167. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    881+167: 881+100=981, +67=1048. Routine.

  • math-expression-1✓ pass7s

    prompt

    Evaluate this expression exactly and respond with just the integer: ((-17 + -4) * (33 - 24)) + (-1 * -1) - 23

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    (-17+-4)=-21; (33-24)=9; -21*9=-189; -1*-1=1; -189+1-23=-211. Double-checked the signs. Confident.

  • math-determinant-1✓ pass43s

    prompt

    Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [-1, -3, 2, 3] [-4, 7, -4, -7] [-2, -3, 4, 0] [9, 3, -8, 4]

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    This one needed care. I expanded along row 1 by hand and made an arithmetic slip (sign on one cofactor) that gave -134, then verified with a Python determinant which confirmed 42. The 42 answer feels almost too neat, but both methods agree.

Vision test

14/19 passed · 3 unanswered

time to last answer 39m 25s
  • acuity-20✓ pass5m 24s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I have no vision, so I fetched the PNG and processed it programmatically: located the text rows, segmented the 5 characters, and template-matched each glyph against rendered DejaVu fonts. Tesseract alone was unreliable on the small row-4 glyphs, but per-character template matching gave clear scores. Confident.

  • acuity-14✓ pass17s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Same pipeline as acuity-20: downloaded the chart, found row 5, segmented 5 glyphs, template-matched against DejaVu renders. All five glyphs scored above 0.83 with big gaps to the runner-up, so I trust HXDBM.

  • acuity-10✓ pass1m 54s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row-6 glyphs are only 10px tall so tesseract was blank on three of them. I fell back to per-glyph template matching plus hand-drawn ASCII inspection of the ink masks: char2 showed a left stem with three horizontal bars (B, not 8) and char4 showed a filled waist (8, not 0). Tesseract confirmed chars 1 and 3.

  • acuity-8✓ pass1m 28s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 7 glyphs are 6x4 px, so I inspected raw grayscale values per glyph: char1 has a solid left stem, full top/bottom bars and a central bar = B (rules out R and D); tesseract confirmed 3, Z, V on the upscaled crops.

  • count-simple✓ pass21s

    prompt

    Look at the image at (fetch it and view it). How many orange circles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Connected-component analysis of the orange colour mask (scipy.ndimage.label): 6 components of identical area 9452 px (110x110 bbox, fill 0.78 ~ pi/4, consistent with filled circles), no overlaps.

  • count-medium✕ fail1m 41s

    prompt

    Look at the image at (fetch it and view it). How many blue squares are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 0

    expected 11, got "12"

    agent's debrief

    scipy.ndimage.label on blue mask -> 17 components. Classified each by bbox fill ratio: fill~1.0 = square (12), fill~0.5 = triangle (4), fill~0.79 = circle (1). Corner-fill and midline checks confirmed. Answer = 12 blue squares.

  • count-complex✓ pass1m 18s

    prompt

    Look at the image at (fetch it and view it). How many green diamonds are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Green mask + ndimage.label gave 37 components: 7 squares (fill 1.0), 1 circle (0.79), 22 diamonds and 7 triangles both at area 968. Distinguished diamonds from triangles via row-width profiles: diamonds narrow-wide-narrow (1,43,1); triangles bottom-heavy (2,22,44). ASCII renderings confirm. Answer = 22.

  • spatial-simple✓ pass38s

    prompt

    Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Grid detected via dark-line row/column profiles: 6 vertical and 6 horizontal lines at x/y = 29, 264, 499, 734, 969, 1204 (5x5 grid). Red circle component (fill 0.79) centered at (x=147, y=617) -> x in [29,264) = col 1, y in [499,734) = row 3.

  • spatial-medium✕ fail7m 08s

    prompt

    Look at the image at (fetch it and view it). Which shape does the arrow from the green triangle lead to? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 0

    wrong shape

  • spatial-complex✓ pass13m 28s

    prompt

    Look at the image at (fetch it and view it). Which shape is 3 steps after the orange diamond along the arrows? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

  • chart-simple✓ pass5m 06s

    prompt

    Look at the image at (fetch it and view it). What is the title shown at the top of this chart? Answer with just the chart title.

    answer

    answer hidden on shared reports

    grader · score 1

  • chart-medium✓ passbatched

    prompt

    Look at the image at (fetch it and view it). Using the "Server Incidents" chart, approximately what value did Apr have? Read it off the y-axis; answers within +/-5 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

  • chart-complex✓ passbatched

    prompt

    Look at the image at (fetch it and view it). Using the "Monthly Active Users" chart, approximately what is the difference between New and Returning in Jan? Answers within +/-4 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

  • screenshot-simple✓ pass41s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

  • screenshot-medium✓ passbatched

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

  • screenshot-complex✓ passbatched

    prompt

    Look at the image at (fetch it and view it). What is the tax amount shown on this order summary? Answer with just the amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

  • diagram-simple— unanswered—

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Gecko"? Answer with just the box name, e.g. Kettle.
  • diagram-medium— unanswered—

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Jackal" point to? Answer with just the box name, e.g. Kettle.
  • diagram-complex— unanswered—

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Maple"? Answer with just the box name, e.g. Kettle.

Finding and reading email test

not examined · 0/6 answered

  • aggregate-1— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages carry the label "attachments"? Answer with just the number.
  • aggregate-2— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages in the mailbox are dated during October 2001? Answer with just the number.
  • temporal-1— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message in the archive folder? Answer with just the subject line, exactly as shown.
  • temporal-2— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message in the inbox folder? Answer with just the subject line, exactly as shown.
  • needle-1— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the message with the subject "FW: Curve Shift File", what dollar amount is called out as the biggest difference between actual and theoretical curve shift? Answer with just the exact fact as it appears in the message, and nothing else.
  • needle-2— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Phillip's reply about Bob Huntley's request for a survey of the lot, what fax number does Bob give for receiving faxed documents? Answer with just the exact fact as it appears in the message, and nothing else.

Purchasing test

not examined · 0/4 answered

  • find-product-1— unanswered—

    prompt

    The store is at abostore.airbench.ai Among products in the **Home & Furniture** category priced under **$650** with a rating of at least **4.8**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).
  • find-product-2— unanswered—

    prompt

    The store is at abostore.airbench.ai Among products in the **Sports & Outdoors** category priced at or above **$400** with a rating of at least **4.5**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).
  • purchase-1— unanswered—

    prompt

    The store is at abostore.airbench.ai Buy exactly 3 units of Amazon Brand - 206 Collective Women's Cedar Casual Double Band Sandal, White Leather, 9.5 B US (product id amazon.ae:B078GRH9RH, abostore.airbench.ai/product/amazon-brand-206-collect…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-9a06bdcb@aidoctor.test. Answer with just the resulting order id.
  • recover-decline-1— unanswered—

    prompt

    The store is at abostore.airbench.ai Buy exactly 1 unit of AmazonBasics Vinyl Coated Cast Iron Kettlebell Weight (product id amazon.ae:B07F1NPWCR, abostore.airbench.ai/product/amazonbasics-vinyl-coate…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-dee7ab3c@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.

Coding test

not examined · 0/11 answered

  • compute-hash-1— unanswered—

    prompt

    Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [2250746228, 3802098949, 3606779914, 2161313299, 758647792, 2459801777, 3193997990, 949438559, 2680483244, 1319607453, 110969474, 747004139], x = 653765288, y = 777356489 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.
  • compute-vm-1— unanswered—

    prompt

    Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 603 1: set b 248 2: set c 297 3: set d 517 4: mul a 29 5: mul a 43 6: mul a 9 7: dec d 8: jnz d -4 9: mul a 59 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.
  • compute-paths-1— unanswered—

    prompt

    Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S.....#....#...#......... .#....#.....#.#..#....... #....#.##.......#..#....# ..#.......#..##......#... ....#....#.#..#.#.##....# ....####......##.#..#..## #.##.#.....#...#.#.#...#. .........#.#....###....#. ###.....#...........#...# .#..##.#..#.#...#..#....# ...#..#...#.#..#..#...... ....#....#.......##..#... ...........##............ ##...#.###..#...#.#...#.# ...##...#....#.#...#..### .....#..#....#...##...... ...#..##......#....#..#.. ....##....#....#..#...#.. ##....##..........#..#.## ......#....####.#....#... ..#....#..............#.. ...#..........#.......... ..##.###......##.#####.#. .#........#...#.##....#.. #.#.#.###.....#..####..#E Respond with the two integers separated by a space, like `52 1840`.
  • compute-life-1— unanswered—

    prompt

    Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: ....#......#####...# ...#.#....##.....#.# ..###...######.#.### #.#...#..###.....#.# ......#....##....##. .#..#####..#...#.#.# .##.##....#......##. ##..##..#.#......#.. #.#....##...#.....## .......#.##..##..#.. ..###...#....##...#. ..####.##.##....#... ..#.#.#....###.#.##. ..#....##.#.###.###. ........##.....###.# #.#..##.#.#.....#.#. ..#..#.##..#...##.#. ##.....#....##...#.# ####..##.#..#.#.#.## ....#..........##### Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.
  • compute-fibmod-1— unanswered—

    prompt

    Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 8484331735363230 and m = 1000003. Respond with just the integer.
  • compute-words-1— unanswered—

    prompt

    Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. shamo kamo Basqui Shazan. kamo TRUREN Zantru zantru luqui. kador quimo. zantru Quimo basnix; nixsha nixsha motru vovo quibas quibas "momo" Quimo basqui karen basqui, lutru kamo lubas Mopel quimo Truren shazan basqui dorfic; shazan renvo shamo dorfic bastru truren ficvo shanix renvo ficvo mopel kamo momo Ficvo quibas lubas quibas "quibas" truren dorfic Mopel kamo? Renvo quibas zanzan Shamo zanzan basdor; quibas! renvo, titru zanzan vovo! renvo renvo motru basqui vovo MOTRU mopel zantru lutru pelzan pelzan Luqui kador shanix lubas RENVO shazan bastru shanix! mopel quibas, nixbas pelzan ficvo quibas bastru Dorfic pelzan kador zanzan "kador" QUIMO Dorfic lutru kamo; MOTRU ficvo quibas lutru Pelzan? basqui? Shazan karen Kamo shanix vovo karen shamo dorfic. Quimo Zantru Motru ficvo, basnix quibas quibas BASQUI "quibas" pelzan Motru quibas kador renvo mopel Basqui; renvo vovo renvo truren shazan nixbas "basnix" Quibas quibas BASTRU zanzan Renvo basqui LUBAS; quibas QUIBAS nixbas, shazan, truren renvo shanix, kamo Nixsha Nixsha quibas bastru LUTRU quimo basqui MOMO lubas kamo, Renvo basqui karen zantru? truren VOVO KADOR luqui quibas? mopel bastru mopel zanzan dorfic dorfic NIXBAS nixbas renvo basqui! ficvo; Quibas kamo KAMO Ficvo shazan lubas Ficvo kamo shamo; zantru "luqui" Shamo bastru shamo lubas quibas lubas Karen lubas kamo titru Zantru luqui truren; quibas Dorfic Renvo shazan, ficvo vovo "BASQUI" "Vovo" bastru mopel; shazan modor shazan renvo quibas basqui modor renvo Kamo Momo basqui shanix luqui renvo lubas vovo QUIBAS Quibas shazan Kamo renvo, Momo Quibas motru shamo Vovo quibas bastru renvo mopel quibas basdor nixbas; ficvo Nixbas Zanzan TRUREN Lubas quimo lutru bastru. Bastru quibas zantru quibas Titru? basqui "ficvo" lubas? ficvo zanzan! Luqui quibas ficvo dorfic lutru vovo "Zanzan" motru quimo nixbas; renvo luqui lutru basnix ficvo; quibas basqui nixbas Basqui Momo Dorfic; QUIMO SHAMO karen shazan shamo renvo; luqui Vovo BASDOR shazan bastru quibas karen quibas kador Shamo zantru KAMO mopel Shamo nixsha; ficvo "Quibas" shanix MODOR ficvo shamo dorfic zantru Zantru Mopel ficvo lutru? Momo Ficvo BASQUI zanzan, quibas quibas karen Basqui dorfic lubas dorfic modor, dorfic quibas kamo nixbas basnix Renvo vovo basdor "shanix" lubas shazan; shazan shamo renvo basqui shazan, Renvo momo bastru "nixsha" Shamo Luqui "shazan" quibas basdor shanix renvo lutru quimo karen dorfic "zanzan" vovo shanix "lutru" basdor! Ficvo basqui basnix. "ficvo" motru basqui nixsha shamo shazan Quibas! Basqui ficvo kador Basqui quibas nixsha mopel renvo Shanix basdor "quibas" shazan kamo shanix "shamo" pelzan. motru renvo dorfic shazan nixbas zanzan zanzan pelzan vovo Nixsha "quimo" TITRU Basqui momo dorfic nixbas; Shamo shanix quibas zanzan dorfic nixsha dorfic
  • trace-1— unanswered—

    prompt

    What exactly does this JavaScript program print? Respond with just the printed output. const v1 = [71 / 7 | 0, Math.round(-6.5), -96 % 7].join(","); const v2 = [77, 7, 530, 1277].sort().join(","); const v3 = [typeof null, typeof [], typeof typeof 2].join("/"); const v4 = [null >= 0, "40" < "5", null == 0].map(Number).join(""); console.log(v1, v2, v3, v4);
  • fix-1— unanswered—

    prompt

    Our shipping quote function below has one bug. Bug report: this order is quoted 1789 cents, but the correct quote is 959: {"country":"GB","items":[{"grams":1544,"qty":1,"price":11100,"fragile":false}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 483, 830, 1338, 1875]; // cents, by zone const PER_STEP = [0, 79, 137, 186, 265]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 4500, 11100, 18200, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams * item.qty; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value <= FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"CA","items":[{"grams":1227,"qty":4,"price":8784,"fragile":false},{"grams":1241,"qty":1,"price":2115,"fragile":true},{"grams":668,"qty":1,"price":3676,"fragile":false}],"express":true} {"country":"AU","items":[{"grams":356,"qty":1,"price":6847,"fragile":true},{"grams":682,"qty":4,"price":6920,"fragile":false},{"grams":149,"qty":1,"price":3404,"fragile":false}],"coupon":"SHIP10"} {"country":"FR","items":[{"grams":177,"qty":3,"price":5040,"fragile":false},{"grams":1776,"qty":5,"price":4552,"fragile":false},{"grams":792,"qty":1,"price":6633,"fragile":false}],"express":true} {"country":"FR","items":[{"grams":1996,"qty":1,"price":4500,"fragile":false}]} {"country":"CA","items":[{"grams":159,"qty":1,"price":1736,"fragile":true},{"grams":986,"qty":4,"price":2249,"fragile":false}],"coupon":"SHIP10"} {"country":"IT","items":[{"grams":358,"qty":5,"price":474,"fragile":false},{"grams":504,"qty":2,"price":443,"fragile":false},{"grams":697,"qty":3,"price":5599,"fragile":false}]} {"country":"AU","items":[{"grams":560,"qty":1,"price":18200,"fragile":false}]} {"country":"CA","items":[{"grams":343,"qty":1,"price":11100,"fragile":false}]} {"country":"ES","items":[{"grams":119,"qty":1,"price":4500,"fragile":false}]} {"country":"AU","items":[{"grams":688,"qty":3,"price":8880,"fragile":false},{"grams":207,"qty":1,"price":5572,"fragile":false},{"grams":602,"qty":5,"price":7595,"fragile":false}],"express":true} {"country":"AU","items":[{"grams":1310,"qty":1,"price":18200,"fragile":false}]} {"country":"AU","items":[{"grams":111,"qty":5,"price":5830,"fragile":false},{"grams":102,"qty":1,"price":7320,"fragile":true},{"grams":934,"qty":1,"price":6736,"fragile":true},{"grams":803,"qty":5,"price":6312,"fragile":false}],"coupon":"SHIP10"} {"country":"IT","items":[{"grams":124,"qty":1,"price":4500,"fragile":false}]} {"country":"ZA","items":[{"grams":885,"qty":5,"price":3431,"fragile":true},{"grams":1619,"qty":5,"price":1464,"fragile":false},{"grams":1075,"qty":1,"price":7914,"fragile":true},{"grams":941,"qty":3,"price":678,"fragile":false}]} {"country":"ES","items":[{"grams":414,"qty":3,"price":2886,"fragile":false},{"grams":795,"qty":1,"price":4975,"fragile":true},{"grams":153,"qty":1,"price":4962,"fragile":false},{"grams":365,"qty":1,"price":351,"fragile":false}],"express":true} {"country":"ES","items":[{"grams":1448,"qty":1,"price":4870,"fragile":true},{"grams":1340,"qty":4,"price":6127,"fragile":false},{"grams":686,"qty":1,"price":8781,"fragile":false},{"grams":712,"qty":2,"price":5792,"fragile":true}],"coupon":"SHIP10"} {"country":"DE","items":[{"grams":514,"qty":1,"price":7716,"fragile":false},{"grams":1170,"qty":2,"price":2738,"fragile":true},{"grams":179,"qty":1,"price":8023,"fragile":true}]} {"country":"ZA","items":[{"grams":1386,"qty":2,"price":7684,"fragile":false},{"grams":1131,"qty":4,"price":1240,"fragile":false},{"grams":801,"qty":5,"price":6450,"fragile":true}]} {"country":"ES","items":[{"grams":1362,"qty":1,"price":4500,"fragile":false}]} {"country":"IT","items":[{"grams":708,"qty":1,"price":8742,"fragile":false},{"grams":983,"qty":5,"price":3179,"fragile":true},{"grams":924,"qty":1,"price":1444,"fragile":false}]}
  • implement-1— unanswered—

    prompt

    Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[13,15],[1,2],[1,1],[19,19],[32,37],[6,10]] [[31,38],[26,30],[28,35],[30,36],[21,26]] [[39,44],[11,14],[27,28],[28,34],[32,38],[8,16],[40,45],[7,11]] [[2,4],[19,26],[29,30],[33,33],[6,11],[28,36],[31,32]] [[2,10],[22,29],[4,5],[25,30],[16,23],[5,11]] [[18,23],[12,14],[3,3],[36,39],[40,40],[7,8],[28,29],[32,33]] [[8,9],[18,23],[32,35],[6,14],[36,42]] [[33,39],[40,42],[26,28],[1,5],[34,37],[13,15]]
  • repo-1— unanswered—

    prompt

    Download airbench.ai/f/4399e5b1e37ab4b3370d8e22ac7a881d.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.
  • repo-2— unanswered—

    prompt

    Download airbench.ai/f/bec0fc6bec581998b13bf9d165d4b8f2.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

how this agent was configured

Hardware: NVIDIA RTX 5090 32 GB (SM120), x86_64, driver 595.84, Ubuntu 24.04 (host hal5090). Model server: Checkpoint gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090 (NVFP4, no MTP head). vLLM 0.27.1 (vllm/vllm-openai:v0.27.1): --quantization modelopt --kv-cache-dtype fp8 --trust-remote-code --reasoning-parser qwen3 --enable-auto-tool-choice --tool-call-parser qwen3_xml --max-model-len 65536 --max-num-seqs 8 --gpu-memory-utilization 0.95. ~79 tok/s single-stream decode. Harness: opencode 1.18.29 on the host: `opencode run --auto --format json <prompt>`; per-run opencode.json pins the model via @ai-sdk/openai-compatible; context 65536, max output 16384 tokens, image input declared only when the server passed a two-image probe. Orchestrator: github.com/dh7/agent-checkup-benchmark @ 50881e1; the checkup instructions are passed verbatim as a single prompt, one attempt, no time limit unless noted.