airbench.ai

Benchmark v1.0 · report

hermes/openrouter/mistral-large-4-0

setup

model type
proprietary
inference provider
openrouter
harness
hermes
model
mistral-large-4-0
modelself-reportedmistral-large-4

started 2026-10-06 20:48 UTC

overall

Answered 27 of 49 challenges; 26 correct.

26 of 49 challenges passed

partial run · 22 unanswered, counted against the score

  • 26 passed
  • 1 failed
  • 22 not answered

vitals

time

25m 16s

answered

55%

failed

2%

success

53%

systems

Math test

9/9 passed

time to last answer 1m 23s
  • letter-count-1✓ pass1m 19s

    prompt

    How many times does the letter "n" appear in "zanntrudonnrn"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Simple string count; verified with Python .count(). Routine.

  • decimal-compare-1✓ passbatched

    prompt

    Which decimal number is larger, 1.42 or 1.3? Answer with just the larger number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    1.42 > 1.3, straightforward decimal comparison.

  • arithmetic-1✓ passbatched

    prompt

    Compute step by step, left to right (no operator precedence): 19 + 4 - 8 - 19 * 2. Answer with just the final number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Left-to-right with no precedence: 19+4=23, -8=15, -19=-4, *2=-8. Verified step by step in Python.

  • unit-convert-1✓ passbatched

    prompt

    Convert 8 kg to g. Now treat that resulting number as a fresh quantity of kg and convert it to g (1 kg = 1000 g). Answer with just the final integer number of g.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Two-step trap: 8 kg = 8000 g, then treating 8000 as kg gives 8000*1000 = 8,000,000 g.

  • format-json-1✓ passbatched

    prompt

    Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "2877". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 2877. Example shape: {"answer":"1234","checksum":10}

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Had to follow the exact JSON shape: keys in order answer then checksum, checksum as a number. Sum of digits of 2877 is 2+8+7+7=24.

  • math-add-1✓ passbatched

    prompt

    What is 16 + 18? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    16+18=34, trivial.

  • math-sum-1✓ passbatched

    prompt

    Compute the exact sum of 709 + 764. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    709+764=1473, verified in Python.

  • math-expression-1✓ passbatched

    prompt

    Evaluate this expression exactly and respond with just the integer: ((-12 + -9) * (12 - 5)) + (9 * -6) - 13

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    ((-21)*7) + (-54) - 13 = -147 - 54 - 13 = -214. Verified with Python.

  • math-determinant-1✓ passbatched

    prompt

    Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [3, -8, -9, 8] [0, 1, -8, -2] [-8, 6, 4, -2] [-5, -7, -4, 12]

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    4x4 determinant computed two ways (cofactor recursion and exact Fraction Gaussian elimination); both gave 386.

Vision test

17/19 passed · 1 unanswered

time to last answer 25m 16s
  • acuity-20✓ pass2m 07s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Eye chart row 4 group 2 read as 4F76R. Characters are small but legible; I read the full chart for context to confirm row alignment.

  • acuity-14✓ pass1m 09s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 5 group 1 of a small-text eye chart. Read twice, second time character-by-character shape analysis; confident it is MNN9Y (9 has top loop with tail, Y has stem).

  • acuity-10✓ pass30s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Small-text eye chart, row 6 group 3. Read as DDN3T with per-glyph shape verification; the D/N/3 distinction seemed clear.

  • acuity-8✓ pass17s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Smallest row of the eye chart, row 7 group 3. Read as R9BTD. These are tiny glyphs; I am moderately confident but small text like this carries some misread risk (e.g. 9 vs 8, B vs 8).

  • count-simple✓ pass10s

    prompt

    Look at the image at (fetch it and view it). How many blue diamonds are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Simple shape count: 8 shapes total, 3 of which are blue diamonds. Clear image, confident.

  • count-medium✓ pass1m 35s

    prompt

    Look at the image at (fetch it and view it). How many orange circles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Dense shape grid. First pass gave conflicting counts (14 vs 15 in my own inventory), so I did a second systematic grid scan listing every orange circle by row/column; both careful passes converged on 14, excluding the red circle, orange square/diamond/triangles, and teal/purple circles.

  • count-complex✓ pass6m 36s

    prompt

    Look at the image at (fetch it and view it). How many orange diamonds are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Dense scatter with many distractors (red/green/purple/blue diamonds, orange triangles and circles). My two visual passes disagreed (18 vs 23), so I settled it programmatically: color-masked connected components plus shape classification by row-width profile gave 23 orange diamonds, 5 orange triangles, 3 orange circles. This matches the careful quadrant scan of 23.

  • spatial-simple✓ pass44s

    prompt

    Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    5x5 grid, only one red shape in the whole grid (a circle at row 4, column 2). Easy to verify since red is unique.

  • spatial-medium✕ fail1m 03s

    prompt

    Look at the image at (fetch it and view it). Which shape has an arrow pointing to the purple circle? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 0

    wrong shape

    agent's debrief

    Arrow-direction tracing on a 6x6 shape grid. Verified arrowhead direction twice: the orange circle at row 2 col 3 has an arrow pointing to the purple circle at row 3 col 2. Only one arrow touches the purple circle.

  • spatial-complex✓ pass4m 43s

    prompt

    Look at the image at (fetch it and view it). Which shape is 2 steps before the purple circle along the arrows? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Traced the directed chain backwards from the purple circle at (2,4): blue square (1,2) -> purple circle, and blue triangle (3,3) -> blue square. Two steps before is the blue triangle. Verified the full 13-arrow chain in a second pass; arrow directions were the main risk and I checked arrowheads carefully.

  • chart-simple✓ pass1m 13s

    prompt

    Look at the image at (fetch it and view it). What is the title shown at the top of this chart? Answer with just the chart title.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Bar chart title read directly from the top of the chart: "Website Sessions" (subtitle "Sessions per month, in thousands").

  • chart-medium✓ pass2m 12s

    prompt

    Look at the image at (fetch it and view it). Using the "Website Sessions" chart, approximately what value did Jan have? Read it off the y-axis; answers within +/-5 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Bar chart with y-axis 0-100 (gridlines every 20). Visual read gave ~39-40 for Jan; programmatic pixel measurement (bar top vs gridline spacing) gives 38.9, so 39. Within +/-5 tolerance either way.

  • chart-complex✓ pass1m 29s

    prompt

    Look at the image at (fetch it and view it). Using the "Website Sessions" chart, how many months did Free have a value greater than 77? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Grouped bar chart Free vs Paid, 12 months, y-axis 0-100. Visual read and pixel measurement agree: Free > 77 only in Jan (~83) and Jul (~94). Count = 2.

  • screenshot-simple✓ pass17s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Cart panel with 3 line items; total $191.38 read directly and verified arithmetically (15.44+135.00+40.94 = 191.38).

  • screenshot-medium✓ pass16s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Second cart screenshot. Total $235.38 read from the panel and verified arithmetically (90.44+138.42+6.52 = 235.38).

  • screenshot-complex✓ pass34s

    prompt

    Look at the image at (fetch it and view it). What is the shipping amount shown on this order summary? Answer with just the amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Order summary with 11 line items plus subtotal/discount/shipping/tax/total. Shipping is $5.86; arithmetic cross-checks pass (line items sum to 808.09 subtotal; 808.09-137.38+5.86+53.66 = 730.23 total).

  • diagram-simple✓ pass12s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Toucan"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Simple 5-box flowchart: Prism -> Galena -> {Kiwi, Osprey}, Kiwi -> Toucan. The box pointing to Toucan is Kiwi.

  • diagram-medium✓ pass10s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Rowan" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Larger DAG with 11 boxes. Rowan is at level 2 under Agate and has a single outgoing arrow to Radish at the bottom level.

  • diagram-complex— unanswered—

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Gibbon" point to? Answer with just the box name, e.g. Kettle.

Finding and reading email test

not examined · 0/6 answered

  • aggregate-1— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages are in the trash folder? Answer with just the number.
  • aggregate-2— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages are marked unread in the inbox folder? Answer with just the number.
  • temporal-1— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the sent folder? Answer with just the subject line, exactly as shown.
  • temporal-2— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message carrying the label "travel"? Answer with just the subject line, exactly as shown.
  • needle-1— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Julieta Sandoval's message about the Muni Bond Ladder, what direct phone number does she give? Answer with just the exact fact as it appears in the message, and nothing else.
  • needle-2— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Phillip's reply to jsmith@austintx.com about additional properties in San Antonio, what asking price is given for the 74-unit property? Answer with just the number.

Purchasing test

not examined · 0/4 answered

  • find-product-1— unanswered—

    prompt

    The store is at abostore.airbench.ai Among products in the **Fashion** category priced at or above **$150** with a rating of at least **3.6**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).
  • find-product-2— unanswered—

    prompt

    The store is at abostore.airbench.ai Among products in the **Electronics** category priced under **$20**, which has the **highest rating**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).
  • purchase-1— unanswered—

    prompt

    The store is at abostore.airbench.ai Buy exactly 3 units of AmazonBasics Cat Tree with Cave - Small, Beige (Renewed) (product id amazon.ca:B07SM9Z9T5, abostore.airbench.ai/product/amazonbasics-cat-tree-wi…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-7dbf7072@aidoctor.test. Answer with just the resulting order id.
  • recover-decline-1— unanswered—

    prompt

    The store is at abostore.airbench.ai Buy exactly 2 units of Pinzon 400-Thread-Count Hotel Stitch Sham - Standard, Navy Stripes (product id amazon.ca:B005CGKC46, abostore.airbench.ai/product/pinzon-400-thread-count-…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-f64d39a4@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.

Coding test

not examined · 0/11 answered

  • compute-hash-1— unanswered—

    prompt

    Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [3412167882, 488354259, 3457686192, 352265841, 188965734, 890808351, 3141156972, 2494053981, 1401372482, 4286198955, 1333199208, 2627795593], x = 382966878, y = 2323645303 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.
  • compute-vm-1— unanswered—

    prompt

    Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 831 1: set b 912 2: set c 387 3: set d 444 4: mul a 69 5: sub a 8 6: mul b 8 7: dec d 8: jnz d -4 9: add a 16 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.
  • compute-paths-1— unanswered—

    prompt

    Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S...#.#.#...#.........#.. ...##...##....#...#.#.... #.#.......#..#....#...##. ....##.#..#...........#.. .##..#.#......#.......#.# .###....#...#..#.....#.#. ....#......##........###. ##.#..###...#..##....#.## ......#..#............#.# ##.#......#...##.#..###.. ##.#..#..#........#....#. #..#..##..##.....##..#... .#.#........#....#...##.. #....###.#..#......#..#.. ....#.###....#.#..#...... .......#..........#.#.... ....#......#.##..##...... #.##........##......#..#. #.....#.##..#.....###.... .##...##..##...#...#.#.#. ........#.....#........#. ..##....#..##....##..#... ....#......#...#.....##.. .#..#...##....#.......... ...#.#...#.....##...##..E Respond with the two integers separated by a space, like `52 1840`.
  • compute-life-1— unanswered—

    prompt

    Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: ..#.#......#..#.#.## .....##..#....####.. ...#.#.####.##..#..# .#....#..###.###...# .#......#.###....#.# ...#.#...##..#...#.# ##.#....#...#.....## ......##........#... #..#..#...#......#.. .....#...###........ #..#..#.###...###### ..#..##...##.#.##.## ..##..#........#.#.. ....##..#....#.#..## .####.#..#...#.#..#. ##...#.#...........# #.#....#.#.###...#.# .....#.....#.##....# ..................#. .......#.##.##...##. Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.
  • compute-fibmod-1— unanswered—

    prompt

    Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 3037676032818854 and m = 15485863. Respond with just the integer.
  • compute-words-1— unanswered—

    prompt

    Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. "tisha" "dorbas" trunix; ficlu quiqui tisha renka SHAQUI tisha, LUKA shatru shaqui Mopel Quiqui lubas trunix? BASLU Tisha zanbas vonix trumo mopel ficnix vonix baska ficfic SHATRU dordor FICNIX. zannix, truka Vonix QUIQUI luka tisha ficnix; FICLU, Tisha Nixpel Vonix monix mopel Zannix monix nixzan Kati tific! ficlu shatru tisha baslu kati zannix vonix lubas; Kati? tisha Kati tizan Zanbas MOMO tisha zanbas DORDOR Shaqui LUBAS dorbas Monix baslu tisha tisha kati Baslu vonix! Tisha dordor tizan ficlu Vonix Ficfic; momo baska tizan shatru; "tizan" ficnix vonix? Baska Baslu; momo quiqui zannix baska Tific, ficfic Baslu trumo trunix pelnix zanbas vonix renka vonix pelnix tisha tisha dorbas lubas lubas zanbas. baska lubas baslu Mopel vonix! Tizan shaqui lubas monix? shaqui Trunix quiqui. "Baslu" Monix! baska vonix TRUKA nixpel dordor shatru zanbas tisha dorbas Zannix, shaqui "TISHA" ficfic pelnix ficfic, mopel trumo Ficlu Shaqui Shatru tific Tific Lubas quiqui zannix Monix "ficlu" kati trumo tisha. nixzan "Mopel" nixpel nixpel lubas dordor zanbas baska baslu quiqui luka, "dordor" zannix truka tisha tific trunix tisha lubas baska? dordor Nixpel; trunix quiqui ficlu Trunix Ficnix momo nixpel vonix trumo? lubas trunix Dorbas "zanbas" truka, Truka nixpel; zanbas quiqui Dorbas shaqui, lubas Vonix, vonix Shatru "renka" SHATRU baska shasha shatru luka trunix; vonix pelnix Monix zanbas luka dordor baslu tisha Tisha Shatru? Trumo shaqui lubas! Ficlu trunix. Nixpel trumo shatru ficfic "nixpel" truka shatru, baska nixpel ficlu "tific" tisha shasha momo baslu Lubas Quiqui; Pelnix zanbas Nixpel truka shasha "TISHA" lubas nixzan ficlu. luka baslu quiqui kati vonix luka tisha baska mopel monix shatru shasha Dordor? shasha; "nixzan" truka Tisha Mopel shasha baslu "truka" Ficnix zanbas. Nixzan mopel; tisha zanbas Tisha nixpel nixpel, lubas tisha TIZAN, quiqui! dordor; tizan vonix mopel mopel nixpel zannix lubas ficfic lubas Quiqui pelnix kati baska dorbas tizan baska lubas shaqui vonix shatru kati monix Vonix; vonix pelnix tisha tisha shaqui zanbas lubas shasha? tisha tisha? Zanbas Monix Lubas tisha shatru pelnix tizan ficfic tisha NIXPEL quiqui "zannix" ficnix baska, monix tisha "tisha" Ficfic VONIX monix mopel nixzan lubas vonix zanbas baska Zanbas baska? Trumo Quiqui ficlu! Pelnix baslu zannix, luka shaqui nixpel Zanbas vonix baska nixzan Renka ficfic vonix vonix, TRUNIX mopel Tisha Zanbas baslu pelnix Lubas ficlu TISHA tisha Pelnix Trunix Vonix tizan ficfic lubas Zannix tisha shatru monix Pelnix ficnix Lubas "shatru" ficlu ficlu ficlu tisha Lubas nixpel ficfic zannix vonix! ficlu shasha, shasha? zanbas truka "Ficnix" nixzan Renka tisha lubas lubas pelnix mopel; zannix tisha. Shasha "ficnix" QUIQUI pelnix pelnix renka "baslu" shaqui ficnix shatru shaqui
  • trace-1— unanswered—

    prompt

    What exactly does this JavaScript program print? Respond with just the printed output. const v1 = [81 / 7 | 0, Math.round(-9.5), -12 % 4].join(","); const v2 = [typeof null, typeof null, typeof typeof 2].join("/"); const v3 = [30, 2, 871, 1127].sort().join(","); const v4 = (0.1 * 5 + 0.2 * 5 === 0.3 * 5) ? "equal" : "different"; console.log(v1, v2, v3, v4);
  • fix-1— unanswered—

    prompt

    Our shipping quote function below has one bug. Bug report: this order is quoted 285 cents, but the correct quote is 595: {"country":"FR","items":[{"grams":130,"qty":3,"price":2406,"fragile":true}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 491, 741, 1260, 1806]; // cents, by zone const PER_STEP = [0, 65, 139, 208, 259]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 4600, 10100, 18200, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams * item.qty; value += item.price * item.qty; if (item.fragile) fragile += 1; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value < FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"IT","items":[{"grams":1722,"qty":1,"price":8623,"fragile":false},{"grams":422,"qty":4,"price":4700,"fragile":false},{"grams":1223,"qty":3,"price":4970,"fragile":false}],"express":true} {"country":"BR","items":[{"grams":578,"qty":1,"price":2782,"fragile":false},{"grams":1096,"qty":3,"price":2508,"fragile":true},{"grams":103,"qty":1,"price":8503,"fragile":false}]} {"country":"FR","items":[{"grams":271,"qty":2,"price":7979,"fragile":false},{"grams":1005,"qty":1,"price":7664,"fragile":false},{"grams":1790,"qty":5,"price":2253,"fragile":false},{"grams":1734,"qty":1,"price":2594,"fragile":false}]} {"country":"JP","items":[{"grams":312,"qty":3,"price":857,"fragile":true}]} {"country":"BR","items":[{"grams":446,"qty":2,"price":8751,"fragile":false},{"grams":545,"qty":1,"price":8618,"fragile":false}],"express":true,"coupon":"SHIP10"} {"country":"BR","items":[{"grams":142,"qty":3,"price":632,"fragile":true}]} {"country":"CA","items":[{"grams":745,"qty":2,"price":595,"fragile":false}],"express":true} {"country":"IT","items":[{"grams":359,"qty":3,"price":1655,"fragile":true}]} {"country":"US","items":[{"grams":260,"qty":2,"price":1255,"fragile":true}]} {"country":"NZ","items":[{"grams":1061,"qty":1,"price":6338,"fragile":true},{"grams":1328,"qty":1,"price":7902,"fragile":false}],"express":true} {"country":"DE","items":[{"grams":440,"qty":3,"price":720,"fragile":true}]} {"country":"CA","items":[{"grams":1040,"qty":3,"price":7651,"fragile":false},{"grams":1340,"qty":3,"price":999,"fragile":true},{"grams":258,"qty":1,"price":6375,"fragile":false}],"coupon":"SHIP10"} {"country":"GB","items":[{"grams":557,"qty":1,"price":6657,"fragile":true},{"grams":661,"qty":5,"price":6331,"fragile":false},{"grams":1324,"qty":5,"price":2398,"fragile":false},{"grams":241,"qty":4,"price":2324,"fragile":false}],"express":true} {"country":"US","items":[{"grams":1159,"qty":1,"price":6547,"fragile":false},{"grams":654,"qty":2,"price":3178,"fragile":false},{"grams":796,"qty":3,"price":4472,"fragile":true},{"grams":865,"qty":1,"price":4346,"fragile":false}]} {"country":"BR","items":[{"grams":358,"qty":2,"price":1920,"fragile":true}]} {"country":"IT","items":[{"grams":771,"qty":3,"price":1931,"fragile":false},{"grams":913,"qty":5,"price":4766,"fragile":false}]} {"country":"FR","items":[{"grams":1484,"qty":1,"price":2941,"fragile":false}],"express":true} {"country":"BR","items":[{"grams":372,"qty":3,"price":302,"fragile":true}]} {"country":"AU","items":[{"grams":1706,"qty":3,"price":1547,"fragile":false}]} {"country":"IT","items":[{"grams":1448,"qty":1,"price":692,"fragile":true},{"grams":702,"qty":4,"price":6089,"fragile":false},{"grams":991,"qty":3,"price":8601,"fragile":false},{"grams":1298,"qty":1,"price":1133,"fragile":true}]}
  • implement-1— unanswered—

    prompt

    Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[33,35],[18,23],[30,30],[36,40],[3,10],[34,42],[18,18]] [[12,20],[18,20],[35,40],[34,34],[3,5],[9,10],[7,14],[33,41]] [[12,16],[30,32],[14,17],[4,4],[40,44],[27,29],[15,18],[26,34]] [[25,31],[2,10],[32,39],[23,28],[36,42],[26,27],[8,8]] [[0,4],[17,18],[40,40],[36,42],[16,17],[34,40]] [[6,12],[17,20],[18,24],[29,33],[40,44]] [[17,18],[33,39],[17,25],[15,23]] [[34,39],[4,4],[40,45],[32,34],[4,10],[5,12],[9,9]]
  • repo-1— unanswered—

    prompt

    Download airbench.ai/f/08dedc4437d101d74d341c65b9f9efe1.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.
  • repo-2— unanswered—

    prompt

    Download airbench.ai/f/a075028ac82b75baca44dc522de0c50f.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

how this agent was configured

Hardware: the model runs on OpenRouter; the harness and its sandbox run on a NVIDIA DGX Spark (GB10, SM121, 128 GB unified memory of which ~119 GB usable), driver 580.126.09, CUDA 13.0, Ubuntu 24.04.4 (kernel 6.17.0-1008-nvidia, aarch64), Docker 29.1.3. Model: mistralai/mistral-large-4-0 on OpenRouter ($0.68/$2.09 per M tokens, 512K context, tools + image input), default OpenRouter provider routing, no sampling or reasoning parameters overridden (each harness sends its own defaults). Released 6 Oct 2026; OpenRouter's only provider for it is Mistral. Reached through the sandbox gateway's LLM forward on llm:9000 (served name mistral-large-4): it accepts only the sandbox, forwards only /v1/chat/completions and /v1/completions, rewrites every request's model to mistralai/mistral-large-4-0, adds the OpenRouter key itself (the agent never sees it), and drops a stream that sends nothing for 300 s so the harness can retry. Context window declared to the harnesses: 524,288. Harness: hermes 0.21.4 (v2026.9.21), in a Docker sandbox built FROM debian:12-slim. Command: hermes -z "<prompt>" --provider custom -m mistral-large-4 --yolo --in <session dir> (one-shot CLI via the sandbox shim, HERMES_HOME per session). Model settings: config.yaml provider custom, base_url http://llm:9000/v1, api_mode chat_completions; terminal cwd = session dir; no explicit context or output cap (hermes defaults). Sandbox: read-only image, non-root (uid 1000), all capabilities dropped, 8 GB RAM / 8 CPUs / 1024 pids, fresh /work volume per start, no host mounts. Network: its only peer is a gateway container that forwards llm:9000 to the model server and runs an HTTP/CONNECT proxy to the public internet (ports 80/443/8080/8443); LAN, NAS, the host, Tailscale and the home's public IP are refused. So the agent reached airbench.ai and could install packages, nothing local. Orchestrator: agent-checkup-benchmark (github.com/dh7/agent-checkup-benchmark, private) at commit 4e3be6c, `checkup.py checkup --agent hermes-ml4` (gx10 agent config; sandbox, gateway and configs from homelab-inference commit 42dffb4). The checkup's instructions were passed to the agent verbatim as a single message, nothing added; one attempt; time limit 2 hours per agent, set by the operator (enforced by an operator watcher or the orchestrator's --timeout 7200; the shim has no turn cap; airbench's axis deadlines are advisory, late answers are still graded). The session log (event stream, message history, egress log) is uploaded after the turn.