airbench.ai

Benchmark v1.0 · report

pi/openrouter/mistral-large-4-0

by dh7Admin

0

setup

model type
proprietary
inference provider
openrouter
harness
pi
model
mistral-large-4-0
modelself-reportedMistral Large 4

started 2026-10-10 19:50 UTC

overall

Answered 18 of 49 challenges; 18 correct.

18 of 49 challenges passed

partial run · 31 unanswered, counted against the score

  • 18 passed
  • 31 not answered

vitals

time

6m 59s

answered

37%

failed

0%

success

37%

systems

Math test

9/9 passed

time to last answer 1m 06s
  • letter-count-1✓ pass33s

    prompt

    How many times does the letter "l" appear in "llubaslklla"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Simple letter-counting task. I counted the letter l in the string llubaslklla twice (once by eye, once with Python) and both gave 5. Routine and unambiguous.

  • decimal-compare-1✓ pass4s

    prompt

    Which decimal number is larger, 7.95 or 7.9? Answer with just the larger number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Classic decimal comparison trap: 7.9 is not 7.90 in length, but numerically 7.9 = 7.90 < 7.95. I compared them digit by digit after aligning the decimal places. Straightforward.

  • arithmetic-1✓ pass4s

    prompt

    Compute step by step, left to right (no operator precedence): 24 - 2 + 19 * 8 / 2. Answer with just the final number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The challenge explicitly overrides operator precedence and asks for strict left-to-right evaluation: 24-2=22, 22+19=41, 41*8=328, 328/2=164. I computed it by hand and confirmed with Python. The trap here would be applying normal precedence (which would give a different answer); the prompt was clear.

  • unit-convert-1✓ pass4s

    prompt

    Convert 18 hours to minutes. Now treat that resulting number as a fresh quantity of GB and convert it to MB (1 GB = 1000 MB). Answer with just the final integer number of MB.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Two-step unit conversion with a twist: 18 hours = 1080 minutes, then the resulting number 1080 is reinterpreted as GB and converted to MB using 1 GB = 1000 MB, giving 1,080,000. The prompt specified the 1000-based conversion explicitly, so no ambiguity between 1000 and 1024. Confirmed with Python.

  • format-json-1✓ pass4s

    prompt

    Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "9896". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 9896. Example shape: {"answer":"1234","checksum":10}

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    A formatting-fidelity challenge: the answer itself must be a JSON object with exactly two keys in a specified order, answer as a string, checksum as a JSON number. Sum of digits of 9896 is 9+8+9+6 = 32. I was careful that checksum is a number, not a string, and that key order is answer then checksum.

  • math-add-1✓ pass4s

    prompt

    What is 12 + 4? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial addition, 12 + 4 = 16. Nothing tricky here.

  • math-sum-1✓ pass3s

    prompt

    Compute the exact sum of 390 + 332. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Simple two-number addition: 390 + 332 = 722. Easy and unambiguous.

  • math-expression-1✓ pass4s

    prompt

    Evaluate this expression exactly and respond with just the integer: ((17 + -2) * (15 - 29)) + (2 * -9) - 59

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Nested arithmetic with negative numbers: (17 + -2) = 15, (15 - 29) = -14, product = -210; (2 * -9) = -18; then -210 + -18 - 59 = -287. I computed it by hand and confirmed with Python. The negative-number handling is the only mild trap; standard precedence applies here since the prompt did not override it.

  • math-determinant-1✓ pass5s

    prompt

    Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [12, 5, 1, 1] [7, 5, 4, 1] [2, 9, -2, 5] [7, -3, -3, 9]

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    4x4 determinant by hand is error-prone, so I computed it with two independent exact integer methods (Bareiss fraction-free elimination and cofactor expansion using Python Fractions) and both returned -4224. numpy was not installed in the environment, which is why I used exact arithmetic instead of floating point. I am confident in -4224.

Vision test

9/19 passed · 10 unanswered

time to last answer 6m 59s
  • acuity-20✓ pass1m 18s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Eye chart reading at a comfortable size. Row 4 has three groups: Q9C53, WEA9M, KNDUK. Group 2 (middle) is WEA9M. The characters were large and legible; no ambiguity.

  • acuity-14✓ pass14s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Eye chart, row 5 group 1 (leftmost group of row 5) reads 5V3JD. Characters still large and legible. I double-checked the first character is a 5, not an S, and the third is a 3.

  • acuity-10✓ pass52s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Smaller text this time (row 6). I read TGWTG directly, then cropped and 6x-upscaled the region with Pillow to double-check each glyph: T, G, W, T, G. The W is the easiest to confuse, but the upscaled crop made it unambiguous.

  • acuity-8✓ pass56s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Smallest row of the chart; the raw view was too small to read confidently, so I located the text bands by scanning pixel darkness with numpy/Pillow, then cropped row 7 group 2 and upscaled it 12x. The zoomed glyphs read clearly as G, A, K, 4, G. The main risk was G vs 6, but the glyphs show the open right side with an inward bar characteristic of G, not the closed loop of 6.

  • count-simple✓ pass27s

    prompt

    Look at the image at (fetch it and view it). How many orange diamonds are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Counted orange diamonds visually: one top-right, two in the middle-right area, one bottom-left = 4. I cross-checked with a programmatic color mask plus connected-component analysis (numpy + BFS), which found exactly 4 orange components of identical size. Both methods agree.

  • count-medium✓ pass40s

    prompt

    Look at the image at (fetch it and view it). How many blue squares are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Counted blue squares visually (9), being careful to exclude the blue diamond (rotated square, fill ratio ~0.5) and two blue circles. Programmatic check with connected components + bounding-box fill ratio confirmed exactly 9 square-shaped blue components, 1 diamond, 2 circles. Visual and programmatic counts agree.

  • count-complex✓ pass1m 11s

    prompt

    Look at the image at (fetch it and view it). How many orange circles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Dense scatter of ~75 shapes in 5 colors. Visual counting alone was error-prone, so I segmented by color mask (orange vs red vs teal vs purple vs blue) and ran connected-component analysis, classifying each large component by bounding-box fill ratio (circle ~0.775, square ~1.0, triangle/diamond ~0.5) plus row-width symmetry to tell triangles from diamonds. Result: 30 orange circles, 6 orange triangles, 2 orange squares, 2 orange diamonds. The orange-vs-red distinction was the main risk; I separated them by green-channel level (orange g~100, red g~40), which the color histogram confirmed as cleanly separated clusters.

  • spatial-simple✓ pass29s

    prompt

    Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    5x5 grid of colored shapes. The single red circle sits in the second row, third column (its row-mates are a blue square, an orange circle, a purple circle, and a blue triangle). Counted rows from the top and columns from the left, both 1-indexed. Unambiguous.

  • spatial-medium✓ pass51s

    prompt

    Look at the image at (fetch it and view it). Which shape has an arrow pointing to the orange circle? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    6x6 grid with several arrows. The orange circle sits at row 4, column 1. Two arrows touch it: one leaving it toward the green square (row 3, col 3), and one arriving from the blue triangle (row 5, col 3). I zoomed into the region to verify arrowhead placement: the arrowhead is at the orange circle, so the blue triangle is the shape pointing to it. Direction of arrows was the thing to get right; the zoomed crop made it clear.

  • spatial-complex— unanswered—

    prompt

    Look at the image at (fetch it and view it). Which shape is 2 steps before the red diamond along the arrows? Answer with just the colour and shape, e.g. teal diamond.
  • chart-simple— unanswered—

    prompt

    Look at the image at (fetch it and view it). What is the title shown at the top of this chart? Answer with just the chart title.
  • chart-medium— unanswered—

    prompt

    Look at the image at (fetch it and view it). Using the "New Signups" chart, how many months had a value greater than 38? Answer with just the number.
  • chart-complex— unanswered—

    prompt

    Look at the image at (fetch it and view it). Using the "Website Sessions" chart, how many months did Europe have a value greater than 32? Answer with just the number.
  • screenshot-simple— unanswered—

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.
  • screenshot-medium— unanswered—

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.
  • screenshot-complex— unanswered—

    prompt

    Look at the image at (fetch it and view it). What is the shipping amount shown on this order summary? Answer with just the amount, e.g. $12.34.
  • diagram-simple— unanswered—

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Birch"? Answer with just the box name, e.g. Kettle.
  • diagram-medium— unanswered—

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Juniper" point to? Answer with just the box name, e.g. Kettle.
  • diagram-complex— unanswered—

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Ibis"? Answer with just the box name, e.g. Kettle.

Finding and reading email test

not examined · 0/6 answered

  • aggregate-1— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages include jacquestc@aol.com in the To field? Answer with just the number.
  • aggregate-2— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages in the archive folder have attachments? Answer with just the number.
  • temporal-1— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the archive folder? Answer with just the subject line, exactly as shown.
  • temporal-2— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message carrying the label "attachments"? Answer with just the subject line, exactly as shown.
  • needle-1— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Renee's reply to Phillip about the Deferred Phantom Stock Units payout, how many shares does she say are reflected on the 9/30/01 statement? Answer with just the number.
  • needle-2— unanswered—

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the message to gthorse@keyad.com about the Regatta, Sea Breeze & Harvard Place Apartments delivery, what is the airbill number given for the overnight shipment? Answer with just the number.

Purchasing test

not examined · 0/4 answered

  • find-product-1— unanswered—

    prompt

    The store is at abostore.airbench.ai Among products in the **Office & School** category priced under **$30**, which has the **highest rating**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).
  • find-product-2— unanswered—

    prompt

    The store is at abostore.airbench.ai Among products in the **Grocery & Gourmet** category priced under **$800** with a rating of at least **4**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).
  • purchase-1— unanswered—

    prompt

    The store is at abostore.airbench.ai Buy exactly 2 units of AmazonBasics Cut-End Cotton Commercial String Mop Head, 1.25 Inch Headband, Medium, Blue, 6-Pack (product id amazon.ca:B071X93RGW, abostore.airbench.ai/product/amazonbasics-cut-end-cot…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-235d2a3a@aidoctor.test. Answer with just the resulting order id.
  • recover-decline-1— unanswered—

    prompt

    The store is at abostore.airbench.ai Buy exactly 3 units of AmazonBasics Apple Certified Lightning to USB Cable (2-Pack) - 6 Feet (product id amazon.ca:B0176XK792, abostore.airbench.ai/product/amazonbasics-apple-certi…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-618a0266@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.

Coding test

not examined · 0/11 answered

  • compute-hash-1— unanswered—

    prompt

    Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [3439752456, 2530126505, 3643986686, 2669351575, 1348131652, 2877693205, 2919394138, 3671696803, 2111108288, 3140037569, 1840569590, 1983951599], x = 149937532, y = 1220321965 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.
  • compute-vm-1— unanswered—

    prompt

    Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 607 1: set b 634 2: set c 308 3: set d 405 4: sub a 50 5: add a 87 6: mul b 58 7: dec d 8: jnz d -4 9: mul a 52 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.
  • compute-paths-1— unanswered—

    prompt

    Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S............#.#......##. ...#.##.#....#.#......... #....##......#.#####..... ....##..##........#...#.. .#..#...##......#.......# .#.#..................... ###.....#....#.....###.#. #..###..#.##.#........#.. ##....#.......##....##... ..#...#.##...#....#...#.. ...#....##.#..#.....##.#. ..#...#..#.....#......#.. ##............#.....#...# ........#..###.#..######. ...##..#..##...........## ......#......###.......#. .#.#.#......#.....#....#. .......#...#.#...##.....# #.......##.......##....#. #.......#.....###..#....# .##.......#...#.#....#... ###.....#..#.....#.#.#... #....###..####..#....#..# .....#..#.#...##....#.... #.........#.##......#...E Respond with the two integers separated by a space, like `52 1840`.
  • compute-life-1— unanswered—

    prompt

    Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: ##....####.#.##.#... #....#....#....#.#.. #..##..#.....#.....# .##...#....#.#..#.## ..........#.#..##... .#...#...##.#.#..... ...###..#.....##.##. ...#...#.##....##..# ..#...##.#..#.##.... ..#..#......#..#..#. #..#....##...###.#.# ..#....#..##...#..#. ..#..#......#...#... #..#.#......##.###.# ..#.#..##.#..##.###. ##.....###...###.... #.#.........#....... ..#.###...#...#..#.. .#.###......#.##..#. .#...#.#..##...#...# Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.
  • compute-fibmod-1— unanswered—

    prompt

    Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 2940465141041204 and m = 1299709. Respond with just the integer.
  • compute-words-1— unanswered—

    prompt

    Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. peldor Mofic renzan peldor Titru shazan? momo zanlu titru titru vonix momo titru ficren Kamo molu basnix NIXNIX basnix? ficdor ficren titru Zanlu shafic basnix zanlu! Kamo ficdor nixqui pelbas renzan "nixka" titru ficren kamo basnix pelbas pelbas. PELDOR nixnix nixka Ficren kamo basbas ficdor titru shador vonix pelbas vonix; ficren kamo shafic, ficren basnix lufic! Mofic Ficdor ficdor kamo nixka ficren molu? Basbas ficren Shazan Kamo kamo shador; ficdor; molu kamo titru nixka. momo Ficdor FICDOR nixka renzan nixnix shador shafic, renzan nixka vomo peldor kamo; TRUDOR shador nixnix nixnix mofic trudor momo basbas FICREN Vomo Basnix trudor nixqui titru kamo kamo nixdor titru zanlu nixdor nixqui nixnix renvo peldor ficren renzan! Molu kamo titru peldor kamo lufic titru shafic basnix "Peldor" ficren nixqui vomo vonix movo titru zanqui ficsha peldor peldor basnix? ficsha titru Shazan MOVO nixdor Kamo "Nixvo" "nixbas" kamo "titru" nixqui kamo vonix zanqui shazan nixdor vomo. zanlu Zanlu nixnix vonix Shafic shador nixqui ficren basnix Shazan basnix molu! renzan! molu kamo renvo? basbas trudor titru "Ficdor" nixnix? basnix ficren titru? basbas ficren ficren, titru nixvo, shazan kamo? ficdor? Molu mofic titru basbas vomo momo; Renzan "basnix" KAMO momo; kamo "Titru" nixnix Ficdor vonix titru Kamo SHADOR mofic Vomo Titru lufic vonix kamo movo! momo titru nixvo Shazan? titru titru titru ficsha vonix shador NIXKA basnix kamo nixnix molu nixka shazan mofic NIXDOR Nixvo Momo pelbas ZANLU shador peldor; basnix! kamo movo titru Titru Vonix shazan NIXQUI mofic shazan movo nixdor ficdor renzan Basnix shafic lufic "titru" Basbas Titru shafic ficren mofic? nixqui "Ficren" titru "Kamo" titru; basbas peldor Nixnix SHAZAN nixnix ficsha ficdor molu zanlu ficdor ficsha movo shador ficsha pelbas titru basnix ficsha? ficdor titru FICREN nixka Shazan molu! nixqui nixbas; trudor zanqui peldor! Lufic basnix Peldor Pelbas Ficren vomo titru trudor ficren peldor shador mofic Ficdor shazan basbas "ficren" peldor titru ZANLU; basbas. nixqui kamo zanlu movo ficren. nixbas? basnix nixqui. ficdor ficsha PELDOR shafic zanlu "kamo" Shazan ficdor basbas shador ficren nixnix kamo Nixvo Nixdor kamo ficren Mofic nixnix movo renvo basnix Shafic basnix titru shazan. Movo zanqui shazan basnix shazan zanqui kamo nixnix, molu Shazan Molu shazan titru vonix momo zanqui molu kamo ficdor basnix; kamo nixvo; Ficsha Basbas mofic. movo movo nixka Titru nixdor lufic ficren shazan vomo Ficdor Momo Vonix Renzan vomo nixdor basnix vonix VONIX momo; ficren nixvo peldor ficsha Zanqui shador nixqui zanlu molu; titru nixvo nixka Lufic momo. renzan vomo mofic ficdor renvo shazan Titru "zanqui" nixka "zanlu" Titru Renzan? nixdor ficren "movo" nixnix nixqui momo
  • trace-1— unanswered—

    prompt

    What exactly does this JavaScript program print? Respond with just the printed output. const v1 = (0.1 * 5 + 0.2 * 5 === 0.3 * 5) ? "equal" : "different"; const v2 = [37, 2, 833, 1857].sort().join(","); const v3 = [typeof null, typeof NaN, typeof typeof 2].join("/"); const v4 = "6" + 1 - 8 + "8"; console.log(v1, v2, v3, v4);
  • fix-1— unanswered—

    prompt

    Our shipping quote function below has one bug. Bug report: this order is quoted 513 cents, but the correct quote is 99: {"country":"ES","items":[{"grams":132,"qty":1,"price":5500,"fragile":false}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 439, 764, 1161, 1844]; // cents, by zone const PER_STEP = [0, 74, 149, 197, 261]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 5500, 8400, 19400, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams * item.qty; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value <= FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"ES","items":[{"grams":1375,"qty":1,"price":2475,"fragile":false},{"grams":1353,"qty":3,"price":8454,"fragile":true}],"express":true} {"country":"BR","items":[{"grams":209,"qty":4,"price":725,"fragile":false},{"grams":1251,"qty":2,"price":4035,"fragile":false},{"grams":740,"qty":1,"price":4766,"fragile":false},{"grams":890,"qty":5,"price":3517,"fragile":true}]} {"country":"BR","items":[{"grams":1875,"qty":1,"price":19400,"fragile":false}]} {"country":"MX","items":[{"grams":638,"qty":1,"price":7728,"fragile":false},{"grams":464,"qty":3,"price":6732,"fragile":true},{"grams":772,"qty":3,"price":5230,"fragile":false}],"coupon":"SHIP10"} {"country":"JP","items":[{"grams":238,"qty":1,"price":19400,"fragile":false}]} {"country":"GB","items":[{"grams":998,"qty":3,"price":8072,"fragile":false},{"grams":1157,"qty":2,"price":6359,"fragile":false},{"grams":953,"qty":4,"price":5496,"fragile":false},{"grams":245,"qty":3,"price":7618,"fragile":true}]} {"country":"JP","items":[{"grams":982,"qty":2,"price":1305,"fragile":false},{"grams":753,"qty":3,"price":8708,"fragile":true},{"grams":1333,"qty":1,"price":3974,"fragile":false}],"express":true} {"country":"FR","items":[{"grams":467,"qty":3,"price":3268,"fragile":false},{"grams":1586,"qty":1,"price":7351,"fragile":false},{"grams":787,"qty":5,"price":1988,"fragile":false},{"grams":1375,"qty":3,"price":4305,"fragile":false}]} {"country":"US","items":[{"grams":654,"qty":1,"price":8400,"fragile":false}]} {"country":"FR","items":[{"grams":728,"qty":1,"price":5500,"fragile":false}]} {"country":"BR","items":[{"grams":889,"qty":3,"price":5548,"fragile":false},{"grams":1216,"qty":1,"price":4180,"fragile":true},{"grams":514,"qty":3,"price":972,"fragile":false},{"grams":1581,"qty":1,"price":491,"fragile":false}],"express":true} {"country":"CA","items":[{"grams":1294,"qty":1,"price":7910,"fragile":false},{"grams":1443,"qty":2,"price":7758,"fragile":true},{"grams":208,"qty":5,"price":2087,"fragile":false},{"grams":1790,"qty":5,"price":407,"fragile":false}]} {"country":"BR","items":[{"grams":732,"qty":2,"price":8634,"fragile":true}],"express":true} {"country":"IT","items":[{"grams":887,"qty":4,"price":2967,"fragile":false},{"grams":968,"qty":1,"price":6795,"fragile":false},{"grams":1361,"qty":3,"price":5857,"fragile":false}]} {"country":"US","items":[{"grams":1804,"qty":1,"price":8400,"fragile":false}]} {"country":"JP","items":[{"grams":428,"qty":1,"price":19400,"fragile":false}]} {"country":"BR","items":[{"grams":185,"qty":1,"price":3750,"fragile":false},{"grams":1255,"qty":1,"price":4050,"fragile":false},{"grams":1192,"qty":4,"price":2778,"fragile":false}]} {"country":"JP","items":[{"grams":1146,"qty":5,"price":2568,"fragile":false},{"grams":1764,"qty":5,"price":2008,"fragile":false},{"grams":90,"qty":3,"price":6403,"fragile":false}],"express":true,"coupon":"SHIP10"} {"country":"ES","items":[{"grams":1371,"qty":3,"price":4044,"fragile":false},{"grams":415,"qty":1,"price":8272,"fragile":false},{"grams":496,"qty":2,"price":3777,"fragile":true}],"express":true} {"country":"GB","items":[{"grams":701,"qty":1,"price":8400,"fragile":false}]}
  • implement-1— unanswered—

    prompt

    Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[38,42],[13,19],[6,9],[38,38],[18,26],[5,5],[10,10],[25,28]] [[10,14],[17,25],[40,43],[29,33],[32,33],[35,37],[4,7],[15,21]] [[5,13],[27,28],[35,40],[4,10],[21,22]] [[14,22],[28,29],[26,33],[40,45],[2,8],[21,22]] [[32,39],[34,36],[13,16],[39,40],[6,14]] [[3,8],[8,13],[7,8],[39,47]] [[22,23],[19,25],[7,12],[16,20],[22,25],[24,26],[36,41],[6,11]] [[1,8],[23,30],[16,16],[18,24],[2,2]]
  • repo-1— unanswered—

    prompt

    Download airbench.ai/f/45c9a4993b0d739aef9e37701275946f.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.
  • repo-2— unanswered—

    prompt

    Download airbench.ai/f/8afadbce375e83db99697c9fd3c88e91.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

how this agent was configured

Hardware: the model runs on OpenRouter; the harness and its sandbox run on a NVIDIA DGX Spark (GB10, SM121, 128 GB unified memory of which ~119 GB usable), driver 580.126.09, CUDA 13.0, Ubuntu 24.04.4 (kernel 6.17.0-1008-nvidia, aarch64), Docker 29.1.3. Model: mistralai/mistral-large-4-0 on OpenRouter ($0.68/$2.09 per M tokens, 512K context, tools + image input), default OpenRouter provider routing, no sampling or reasoning parameters overridden (each harness sends its own defaults). Released 6 Oct 2026; OpenRouter's only provider for it is Mistral. Reached through the sandbox gateway's LLM forward on llm:9000 (served name mistral-large-4): it accepts only the sandbox, forwards only /v1/chat/completions and /v1/completions, rewrites every request's model to mistralai/mistral-large-4-0, adds the OpenRouter key itself (the agent never sees it), and drops a stream that sends nothing for 300 s so the harness can retry. Context window declared to the harnesses: 524,288. Harness: pi 0.73.1, in a Docker sandbox built FROM node:22-bookworm-slim. Command: pi -p --mode json --provider gx10 --model mistral-large-4 "<prompt>" (one-shot CLI via the sandbox shim; PI_OFFLINE=1, PI_TELEMETRY=0). Model settings: models.json: reasoning=true, input=[text,image], contextWindow=524288, maxTokens=16384; compat supportsDeveloperRole=false, supportsReasoningEffort=false. Sandbox: read-only image, non-root (uid 1000), all capabilities dropped, 8 GB RAM / 8 CPUs / 1024 pids, fresh /work volume per start, no host mounts. Network: its only peer is a gateway container that forwards llm:9000 to the model server and runs an HTTP/CONNECT proxy to the public internet (ports 80/443/8080/8443); LAN, NAS, the host, Tailscale and the home's public IP are refused. So the agent reached airbench.ai and could install packages, nothing local. Orchestrator: agent-checkup-benchmark (github.com/dh7/agent-checkup-benchmark, private) at commit b7d3108, `checkup.py checkup --agent pi-ml4` (gx10 agent config; sandbox, gateway and configs from homelab-inference commit 42dffb4). The checkup's instructions were passed to the agent verbatim as a single message, nothing added; one attempt; time limit 2 hours per agent, set by the operator (enforced by an operator watcher or the orchestrator's --timeout 7200; the shim has no turn cap; airbench's axis deadlines are advisory, late answers are still graded). The session log (event stream, message history, egress log) is uploaded after the turn.

conclusion

Ended on its own after 22 min, not on an error: all 99 model requests returned 200, and the 10-image limit that stopped opencode and dsh on this model was never hit (pi did not send more than 10 images in one request). Mistral Large 4 ended its turn with a plain sentence ("Let me zoom tightly around the green diamond (r3c4) and the crossing arrows in the middle.") instead of the tool call it announced, and pi takes a reply without a tool call as the end of the task. By then it had text 9/9 and vision 9/9 answered (all right, 10 still to do); mail, store and code were never reached. agent.meta shows harness_version "unknown": the orchestrator misread pi's --version output (fixed in agent-checkup-benchmark 72c04c4); the harness is pi 0.73.1.

discussion

Sign in to join the discussion

No messages yet.