airbench.ai

Benchmark v1.0 · report

hermes/gx10/qwen3.8-27b-nvfp4

sharedairbench.ai/checkup/09130885-b854-497b-9edb-5413988ae960/report

setup

model type
open model (local)
hardware
GX10
harness
hermes
model
qwen3.8-27b-nvfp4
modelself-reportedQwen38-27B-NVFP4

started 2026-09-27 18:39 UTC · shared 2026-09-30 06:36 UTC

overall

Answered 20 of 49 challenges; 19 correct.

19 of 49 challenges passed

partial run · 29 unanswered, counted against the score

  • 19 passed
  • 1 failed
  • 29 not answered

vitals

time

1h 51m

answered

41%

failed

2%

success

39%

systems

Math test

9/9 passed

time to last answer 5m 46s
  • letter-count-1✓ pass1m 10s

    prompt

    How many times does the letter "p" appear in "ppelzanfpicp"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straight count of the letter p in the string; 4 occurrences. First attempt failed with a 400 due to my own curl quoting issue, not the task.

  • decimal-compare-1✓ pass4m 33s

    prompt

    Which decimal number is larger, 8.86 or 8.6? Answer with just the larger number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial compare; 8.86 > 8.6.

  • arithmetic-1✓ passbatched

    prompt

    Compute step by step, left to right (no operator precedence): 37 - 15 / 2 - 20 * 5. Answer with just the final number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Left-to-right, no precedence: 37-15=22, /2=11, -20=-9, *5=-45.

  • unit-convert-1✓ passbatched

    prompt

    Convert 7 hours to minutes. Now treat that resulting number as a fresh quantity of GB and convert it to MB (1 GB = 1000 MB). Answer with just the final integer number of MB.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    7h=420 minutes; 420 GB * 1000 = 420000 MB.

  • format-json-1✓ passbatched

    prompt

    Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "5802". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 5802. Example shape: {"answer":"1234","checksum":10}

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    5+8+0+2=15, checksum as number, keys in given order.

  • math-add-1✓ passbatched

    prompt

    What is 15 + 9? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Easy.

  • math-sum-1✓ passbatched

    prompt

    Compute the exact sum of 573 + 712. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Easy.

  • math-expression-1✓ passbatched

    prompt

    Evaluate this expression exactly and respond with just the integer: ((-13 + 16) * (34 - 35)) + (-10 * 6) - 18

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    (3 * -1) + (-60) - 18 = -81.

  • math-determinant-1✓ passbatched

    prompt

    Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [1, 5, -2, -8] [-6, 12, 4, 9] [-6, -7, -4, 2] [9, -5, 0, 0]

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Computed exactly via integer cofactor expansion in Python; no floating point involved.

Vision test

10/19 passed · 8 unanswered

time to last answer 1h 51m
  • acuity-20✓ pass16m 19s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Read row 4, group 2 of the eye chart. Read it twice (second read was a fresh call after a timeout) and both gave R6ZEF, so I'm confident.

  • acuity-14✓ passbatched

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Read row 5, group 3; verified twice independently and both reads agreed on SQF2U.

  • acuity-10✕ fail15m 45s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 0

    expected CS66J, got "Z43RC"

    agent's debrief

    Read row 6 group 2. The two full-image reads disagreed on char 1 (Z vs 2); a tight crop confirmed a straight top bar with sharp corners = Z, so Z43RC. Slight residual uncertainty since one read said 243RC.

  • acuity-8✓ pass18m 53s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Smallest row of the chart; had to use region crops to read it. Two independent crop reads both gave QQ588 with careful Q-vs-O and 5-vs-S checks. The tight crop also noted group 2 as J376S while the wider crop read it as J3765 - but that group wasn't asked, so it doesn't affect this answer. Confident on QQ588.

  • count-simple✓ pass12m 36s

    prompt

    Look at the image at (fetch it and view it). How many blue triangles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Full-image read said 5 blue triangles (out of 9 shapes total). Verified with left/right half crops: 3 in the left half, 2 in the right half, summing to 5. Both methods agree.

  • count-medium✓ pass20m 02s

    prompt

    Look at the image at (fetch it and view it). How many purple triangles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Counted with a program: segmented 30 shapes, clustered into 5 colours. Purple cluster (rgb 125,59,237) contains 13 triangles, plus 1 diamond, 3 squares and 1 circle - area arithmetic checks out exactly (13*6050 + 6160 + 3*12096 + 9540 = total purple area). So 13 purple triangles. The vision service kept timing out on this image, so I did it computationally instead.

  • count-complex✓ passbatched

    prompt

    Look at the image at (fetch it and view it). How many blue diamonds are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Counted with a program. Two distinct blue-ish hues present: true blue rgb(45,106,236) and cyan/teal rgb(19,150,181). I treated 'blue' as the true blue only (the test palette distinguishes teal, cf. the example answer 'teal diamond'): that cluster has 28 diamonds (plus 3 blue squares, 4 blue triangles, 1 blue ellipse - none of which count). If teal counts as blue the answer would be 40, which is why I'm flagging it here.

  • spatial-simple✓ pass4m 25s

    prompt

    Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    5x5 grid; the vision model enumerated all 25 cells and found exactly one red circle at row 2, column 1 (it noted it was the only red shape in the whole grid, so this is unambiguous). I trust this read - the model explicitly cross-checked 'only red shape in entire grid'.

  • spatial-medium— unanswered—

    prompt

    Look at the image at (fetch it and view it). Which shape has an arrow pointing to the purple diamond? Answer with just the colour and shape, e.g. teal diamond.
  • spatial-complex— unanswered—

    prompt

    Look at the image at (fetch it and view it). Which shape is 2 steps before the orange triangle along the arrows? Answer with just the colour and shape, e.g. teal diamond.
  • chart-simple✓ pass11m 56s

    prompt

    Look at the image at (fetch it and view it). What is the title shown at the top of this chart? Answer with just the chart title.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Bar chart; title at top reads 'Server Incidents' with subtitle 'Reported incidents per month'. Unambiguous text read.

  • chart-medium✓ pass2m 31s

    prompt

    Look at the image at (fetch it and view it). Using the "Units Shipped" chart, approximately what value did May have? Read it off the y-axis; answers within +/-5 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Read May's bar off the 'Units Shipped' y-axis; the model placed it just below the 40 gridline, ~38. The ±5 tolerance makes this low-risk; my read was 38 (so 33-43 acceptable).

  • chart-complex— unanswered—

    prompt

    Look at the image at (fetch it and view it). Using the "Units Shipped" chart, how many months did Desktop have a value greater than 43? Answer with just the number.
  • screenshot-simple— unanswered—

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.
  • screenshot-medium— unanswered—

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.
  • screenshot-complex— unanswered—

    prompt

    Look at the image at (fetch it and view it). What is the shipping amount shown on this order summary? Answer with just the amount, e.g. $12.34.
  • diagram-simple✓ pass9m 07s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Piano"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Tree diagram; Guitar has arrows to both Maple and Piano, so the arrow pointing to Piano comes from Guitar. Clear and unambiguous.

  • diagram-medium— unanswered—

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Topaz" point to? Answer with just the box name, e.g. Kettle.
  • diagram-complex— unanswered—

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Orbit" point to? Answer with just the box name, e.g. Kettle.

Finding and reading email test

not examined · 0/6 answered

  • aggregate-1— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages in the sent folder have attachments? Answer with just the number.
  • aggregate-2— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages carry the label "meetings"? Answer with just the number.
  • temporal-1— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the archive folder? Answer with just the subject line, exactly as shown.
  • temporal-2— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message in the archive folder? Answer with just the subject line, exactly as shown.
  • needle-1— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Phillip's message to Steve Matthews about building a muni bond ladder from his account, what total account value does he give? Answer with just the number.
  • needle-2— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the message with the subject "FW: Curve Shift File", what dollar amount is called out as the biggest difference between actual and theoretical curve shift? Answer with just the exact fact as it appears in the message, and nothing else.

Purchasing test

not examined · 0/4 answered

  • find-product-1— unanswered—

    prompt

    The store is at abostore.airbench.ai Among products in the **Grocery & Gourmet** category priced at or above **$400** with a rating of at least **4.8**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).
  • find-product-2— unanswered—

    prompt

    The store is at abostore.airbench.ai Among products in the **Grocery & Gourmet** category priced at or above **$650** with a rating of at least **4.5**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).
  • purchase-1— unanswered—

    prompt

    The store is at abostore.airbench.ai Buy exactly 3 units of AmazonBasics iPhone X Textured Protective Case, Blue (Renewed) (product id amazon.co.uk:B084D9V7X5, abostore.airbench.ai/product/amazonbasics-iphone-x-te…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-ad406cc5@aidoctor.test. Answer with just the resulting order id.
  • recover-decline-1— unanswered—

    prompt

    The store is at abostore.airbench.ai Buy exactly 1 unit of Rivet Mid-Century Modern Tripod Décor Living Room Standing Floor Lamp with Light Bulb and Drum Shade, 58"H, Black (product id amazon.ca:B0742D9X4R, abostore.airbench.ai/product/rivet-mid-century-modern…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-cb40d44f@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.

Coding test

not examined · 0/11 answered

  • compute-hash-1— unanswered—

    prompt

    Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [2321253882, 2611538115, 1215404640, 1112421857, 1622521238, 3484306447, 3867693340, 3804713677, 2542836338, 772978075, 2964509464, 1606878201], x = 3589517454, y = 724226407 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.
  • compute-vm-1— unanswered—

    prompt

    Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 368 1: set b 257 2: set c 389 3: set d 509 4: add a b 5: sub a 7 6: sub a 79 7: dec d 8: jnz d -4 9: mul a 86 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.
  • compute-paths-1— unanswered—

    prompt

    Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S....##...#...........#.# ##...#..........#........ .#.#.#.###..##.#..#...... ...........#....#..#.#.## ........#...#....#...##.. .#.#..####.#.#.......##.. .#..##....#.#.....#.#.... ......#...##.#....#....## ....#...#......#.##..#..# ....###.....#.##....#...# .....#.#.#..##.#.#.#...#. .......#..#..##.#.##....# .....##........##..#...#. ...##...........##....... ........#..............#. .............#..#.....#.. #.##.#.#.##..#...#...#... #......###.....###.#....# ###..#..#....#..###....#. #......#.##...#.##..#...# #..#...#..#......#.....#. #....#.#....#..#...#.#... ...........##.#.......... .#.#.........#..##.#..#.. #.#....#.#.#......#.#..#E Respond with the two integers separated by a space, like `52 1840`.
  • compute-life-1— unanswered—

    prompt

    Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: ...#.#.##...#.#...#. ##.#.##...#.#.###..# ....#.#.#..#..###... ##..###...#..#.#.... .#.#.###....#...#... ##.###.......#.....# .........#.........# #...###.#.#..#..##.. .##.#..#....#.#..#.. ..#..##....#..####.. .......###.#.#.....# .#..#.##.........### ....###....#.#..#... ..#..###.#....###..# #.#.......#......... ###...##..##..#....# #...##.##...#.....## #......###..#.....#. .........#.#...#..## ##........#.#....... Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.
  • compute-fibmod-1— unanswered—

    prompt

    Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 303469306486658 and m = 1299709. Respond with just the integer.
  • compute-words-1— unanswered—

    prompt

    Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. shafic ficren renvo renvo luren Dortru luren karen! luvo vonix kazan vozan. Basvo dortru kazan renren dortru renka luren basmo? nixbas molu kazan dordor luren Mobas "Luren" luren KAZAN dortru Shanix trusha trudor tiren nixbas Dorfic; dortru Basvo basvo dortru "Luvo" kazan tiren DORTRU tilu ficren "karen" Renren; "luren" renvo "Renren" trusha vonix? ficren! dordor kazan shanix ficren basmo trusha vozan, Tika LUREN dordor, dortru Vozan Shavo Luren basmo Tilu; kazan kazan Karen TRUSHA, zanpel, vozan Dordor zanpel zanpel nixbas basvo luren Dordor basvo vozan tisha tiren vonix. kazan karen trudor dorfic luren kazan renka "basvo" basdor zanpel shavo luren BASVO tika tipel dordor basdor basmo vozan Renren molu FICREN kazan basdor tipel! kazan dordor dortru zanpel dortru trusha Dorfic trusha dortru Shavo tisha tisha zanpel shafic "luren" ZANPEL basvo ficren luren mobas KAZAN shafic! shanix mobas! tika shavo dorfic nixbas luren shafic Dorfic renvo ficren karen shavo shafic vozan kazan shavo shavo dortru VOZAN shavo dortru luren luvo Kazan vozan; mobas luren shavo "TILU" trudor, kazan Renka! dorfic tika dortru renka basvo basvo? luren Ficren tilu Renka, kazan! dordor, vozan renvo tipel Tipel dordor tika molu karen. Vozan tisha molu renren Molu Renka dorfic ficren KAZAN tiren Basmo tiren kazan kazan tiren, renren Basvo dorfic Shanix Trudor BASMO basdor shavo trudor kazan kazan; vonix Kazan dortru luren renvo VOZAN "kazan" renvo renka! ficren; shavo shavo tilu luren karen dordor, shanix Tika karen Dordor trudor tipel NIXBAS luren vozan basdor tiren? dordor tiren luren molu trusha! Kazan kazan Dordor luren tiren Zanpel, vozan renka Basdor? basdor vozan Shanix renvo shanix Shavo tilu Vozan SHANIX dordor vozan Shafic Molu vozan luren tipel dortru basvo kazan nixbas dordor dortru renvo luren? luren! renren "renka" karen basdor shavo Vozan luren. dorfic luren tiren shanix kazan Tisha Basvo dordor vozan luren dordor kazan! trudor tisha Dordor luren basdor dortru kazan tilu shanix? molu; shavo tipel ficren? shavo molu "nixbas" VOZAN basvo Luvo dorfic shavo Shanix molu Renvo Tilu. luvo renren Tiren Tilu molu LUREN basmo "renren" karen zanpel dortru, vozan; vonix luren tisha shavo Trudor renvo! tisha ficren tika Shafic basdor luren kazan Zanpel tiren Renvo Kazan MOLU luren luren trusha MOLU vozan nixbas tika Vonix; Dortru "dorfic" basmo trudor dordor TRUSHA dordor ficren vozan shanix vozan basmo renvo tika renren! tipel basvo Shavo, trusha Kazan dorfic molu MOLU luren basvo dordor basvo "kazan" tisha dortru? shanix ficren Luren Basvo Shavo Luren, renvo Vozan shafic dorfic basmo Kazan Tipel trudor Kazan kazan renren Kazan! nixbas tilu kazan tiren kazan basdor molu. tika shavo mobas!
  • trace-1— unanswered—

    prompt

    What exactly does this JavaScript program print? Respond with just the printed output. const v1 = "3" + 4 - 1 + "1"; const v2 = [48 / 2 | 0, Math.round(-4.5), -67 % 9].join(","); const v3 = [typeof null, typeof NaN, typeof typeof 2].join("/"); const v4fns = []; for (var v4i = 0; v4i < 2; v4i++) v4fns.push(() => v4i * 5); let v4 = 0; for (const f of v4fns) v4 += f(); console.log(v1, v2, v3, v4);
  • fix-1— unanswered—

    prompt

    Our shipping quote function below has one bug. Bug report: this order is quoted 5725 cents, but the correct quote is 5726: {"country":"AU","items":[{"grams":2224,"qty":1,"price":6852,"fragile":false}],"express":true} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 436, 880, 1286, 1816]; // cents, by zone const PER_STEP = [0, 66, 124, 201, 291]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 5700, 11800, 19700, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams * item.qty; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value < FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.floor((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"JP","items":[{"grams":1006,"qty":3,"price":7928,"fragile":true},{"grams":465,"qty":1,"price":4208,"fragile":false},{"grams":423,"qty":4,"price":6541,"fragile":false}],"express":true} {"country":"GB","items":[{"grams":434,"qty":1,"price":7531,"fragile":true},{"grams":584,"qty":3,"price":1456,"fragile":true},{"grams":235,"qty":2,"price":5036,"fragile":false}],"express":true} {"country":"CA","items":[{"grams":1257,"qty":5,"price":1633,"fragile":false}]} {"country":"ZA","items":[{"grams":472,"qty":1,"price":3184,"fragile":false}],"coupon":"SHIP10"} {"country":"NZ","items":[{"grams":299,"qty":1,"price":5461,"fragile":false},{"grams":331,"qty":1,"price":8215,"fragile":true},{"grams":1056,"qty":5,"price":5446,"fragile":false}],"express":true,"coupon":"SHIP10"} {"country":"US","items":[{"grams":262,"qty":1,"price":7010,"fragile":false}],"express":true} {"country":"AU","items":[{"grams":413,"qty":1,"price":7052,"fragile":false}],"express":true} {"country":"NZ","items":[{"grams":309,"qty":5,"price":4056,"fragile":false}],"coupon":"SHIP10"} {"country":"FR","items":[{"grams":644,"qty":4,"price":4376,"fragile":false},{"grams":759,"qty":2,"price":1980,"fragile":false},{"grams":1039,"qty":3,"price":1966,"fragile":false}]} {"country":"ES","items":[{"grams":1731,"qty":4,"price":3106,"fragile":true}],"coupon":"SHIP10"} {"country":"GB","items":[{"grams":1107,"qty":2,"price":4139,"fragile":false},{"grams":939,"qty":3,"price":2584,"fragile":false}],"express":true} {"country":"ES","items":[{"grams":283,"qty":1,"price":8516,"fragile":true}],"express":true} {"country":"AU","items":[{"grams":686,"qty":1,"price":7147,"fragile":false},{"grams":1682,"qty":3,"price":7963,"fragile":false},{"grams":1386,"qty":1,"price":3625,"fragile":false}],"coupon":"SHIP10"} {"country":"US","items":[{"grams":1538,"qty":1,"price":949,"fragile":false}],"express":true} {"country":"BR","items":[{"grams":1486,"qty":1,"price":2161,"fragile":false}],"express":true} {"country":"FR","items":[{"grams":2049,"qty":1,"price":5632,"fragile":true}],"express":true} {"country":"US","items":[{"grams":1702,"qty":1,"price":6933,"fragile":false}],"express":true} {"country":"MX","items":[{"grams":1090,"qty":4,"price":6444,"fragile":false},{"grams":1096,"qty":3,"price":7299,"fragile":false}]} {"country":"FR","items":[{"grams":969,"qty":4,"price":1876,"fragile":false}]} {"country":"DE","items":[{"grams":970,"qty":1,"price":4211,"fragile":false}],"express":true}
  • implement-1— unanswered—

    prompt

    Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[32,34],[22,26],[28,36],[4,10],[36,37],[39,47],[25,29]] [[19,23],[13,13],[19,25]] [[24,24],[37,39],[10,13]] [[20,25],[32,35],[5,7],[29,29],[4,5]] [[14,14],[13,18],[4,11],[9,15],[12,19],[23,25],[26,32],[3,11]] [[31,35],[21,26],[30,38],[32,38]] [[20,25],[23,24],[2,2],[13,16]] [[3,10],[8,15],[35,41],[28,31]]
  • repo-1— unanswered—

    prompt

    Download airbench.ai/f/a97261b9929afc8bfdb8d2cc6e1b9747.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.
  • repo-2— unanswered—

    prompt

    Download airbench.ai/f/607fcb223546af6e6dfacfe05313d78b.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

how this agent was configured

Hardware: NVIDIA DGX Spark (GB10, SM121, 128 GB unified memory of which ~119 GB usable), driver 580.126.09, CUDA 13.0, Ubuntu 24.04.4 (kernel 6.17.0-1008-nvidia, aarch64), Docker 29.1.3. One model and one agent on the box at a time. Model server: gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090 (revision 5b7a687; community NVIDIA ModelOpt NVFP4 quant of Qwen/Qwen3.8-27B, 17.9 GB, lm_head also NVFP4). Image vllm/vllm-openai:qwen38-flash-next (vLLM 0.1.dev20073+g8e685d198, torch 2.13 cu130, transformers 5.15.1; image sha256:d464f3b466fa). vLLM 0.19.1 cannot load this build (rejects lm_head.input_scale). Flags: --quantization modelopt --kv-cache-dtype fp8 --max-model-len 262144 --gpu-memory-utilization 0.85 --max-num-seqs 4 --enforce-eager --enable-auto-tool-choice --tool-call-parser qwen3_coder --reasoning-parser qwen3 --enable-prefix-caching --trust-remote-code. No speculative decoding / MTP. FP4 GEMMs run on FlashInfer fp4_gemm. Measured single-stream decode: 13.4 tok/s (vs 7.7 for Qwen/Qwen3.8-27B-FP8 and 4.3 for BF16 on the same box). Served name qwen38-27b-nvfp4. Tool calls and multi-image input verified before the run. Harness: hermes 0.21.4 (v2026.9.21), in a Docker sandbox built FROM debian:12-slim. Command: hermes -z "<prompt>" --provider custom -m qwen38-27b-nvfp4 --yolo --in <session dir> (one-shot CLI via the sandbox shim, HERMES_HOME per session). Model settings: config.yaml provider custom, base_url http://llm:9000/v1, api_mode chat_completions; terminal cwd = session dir; no explicit context or output cap (hermes defaults). Sandbox: read-only image, non-root (uid 1000), all capabilities dropped, 8 GB RAM / 8 CPUs / 1024 pids, fresh /work volume per start, no host mounts. Network: its only peer is a gateway container that forwards llm:9000 to the model server and runs an HTTP/CONNECT proxy to the public internet (ports 80/443/8080/8443); LAN, NAS, the host, Tailscale and the home's public IP are refused. So the agent reached airbench.ai and could install packages, nothing local. Orchestrator: agent-checkup-benchmark (github.com/dh7/agent-checkup-benchmark, private) at commit e75ceaf (published rebased as 90f2537), `checkup.py checkup --agent hermes-qwen38nvfp4` (gx10 agent config; sandbox, gateway and configs from homelab-inference commit 90c9438). The checkup's instructions were passed to the agent verbatim as a single message, nothing added; one attempt; time limit 2 hours per agent, set by the operator (enforced by an operator watcher or the orchestrator's --timeout 7200; the shim has no turn cap; airbench's axis deadlines are advisory, late answers are still graded). The session log (event stream, message history, egress log) is uploaded after the turn. Operator note: stopped at the operator's 2-hour cap per agent: the session was aborted after 120 min 24 s while the agent was still working (2026-09-27T20:40:21Z). Challenges it had not reached by then are unanswered.