Benchmark v1.0 · report
pi/openrouter/deepseek-v4.1-flash
setup
- model type
- open model (cloud)
- inference provider
- openrouter
- harness
- pi
- model
- deepseek-v4.1-flash
started 2026-09-25 00:06 UTC · shared 2026-09-25 06:41 UTC
overall
Answered 49 of 49 challenges; 46 correct.
46 of 49 challenges passed
- 46 passed
- 3 failed
vitals
time
13m 36s
answered
100%
failed
6%
success
94%
systems
Math test
9/9 passed
letter-count-1✓ pass15s
prompt
How many times does the letter "e" appear in "ekanixpeel"? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Straightforward character count; e appears three times in ekanixpeel.
decimal-compare-1✓ passbatched
prompt
Which decimal number is larger, 6.9 or 6.77? Answer with just the larger number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Trivial comparison; 6.9 > 6.77.
arithmetic-1✓ passbatched
prompt
Compute step by step, left to right (no operator precedence): 33 + 1 * 6 / 6 / 2. Answer with just the final number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Easy left-to-right evaluation without precedence.
unit-convert-1✓ passbatched
prompt
Convert 4 kg to g. Now treat that resulting number as a fresh quantity of kg and convert it to g (1 kg = 1000 g). Answer with just the final integer number of g.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Routine two-step unit conversion; 4kg->4000g then 4000kg->4000000g.
format-json-1✓ pass2s
prompt
Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "4421". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 4421. Example shape: {"answer":"1234","checksum":10}answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Simple formatting task; digit sum of 4421 is 11.
math-add-1✓ passbatched
prompt
What is 16 + 18? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Trivial addition.
math-sum-1✓ passbatched
prompt
Compute the exact sum of 526 + 382. Respond with just the integer.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Routine sum.
math-expression-1✓ passbatched
prompt
Evaluate this expression exactly and respond with just the integer: ((2 + 12) * (13 - 13)) + (4 * -6) - 18answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Straightforward; the first product is zero leaving -24-18 = -42.
math-determinant-1✓ passbatched
prompt
Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [4, 9, 5, 3] [-9, 6, 3, -5] [-4, 7, -1, 9] [-1, 2, 9, 12]answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Computed the 4x4 determinant via cofactor expansion in Python; arithmetic-heavy but mechanical.
Vision test
18/19 passed
acuity-20✓ pass41s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Could read row 4 group 3 clearly; characters are M6ZNZ.
acuity-14✓ pass3s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Row 5 group 3 reads 7N35N; legible enough.
acuity-10✓ pass3s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Row 6 group 1 = XPD52, readable.
acuity-8✓ pass19s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Row 7 was small so I cropped and upscaled it; confirmed 2DUV3.
count-simple✓ pass3s
prompt
Look at the image at (fetch it and view it). How many purple triangles are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Three purple triangles visible; other shapes are green, orange, blue, teal.
count-medium✓ pass16s
prompt
Look at the image at (fetch it and view it). How many orange circles are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Counted orange circles with a color-mask connected-component script (fill ratio ~0.78 distinguishes circles from orange squares/triangles/diamond); got 10.
count-complex✕ fail11s
prompt
Look at the image at (fetch it and view it). How many blue triangles are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 0
expected 25, got "28"
agent's debrief
Many shapes; I used a color-mask connected-component script and classified triangles by fill ratio. The blue mask also isolates purple/teal, so I am fairly confident in 28, though dense overlapping counts are easy to get wrong.
spatial-simple✓ pass3s
prompt
Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Red circle is top-left cell, clearly row 1 column 1.
spatial-medium✓ pass4s
prompt
Look at the image at (fetch it and view it). Which shape has an arrow pointing to the blue diamond? Answer with just the colour and shape, e.g. teal diamond.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Traced each arrow; the only arrowhead pointing at the blue diamond originates from the bottom-left red square.
spatial-complex✓ pass2m 04s
prompt
Look at the image at (fetch it and view it). How many shapes come after the orange square along the arrows? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Brutal to read by eye because arrows overlap, so I detected the shape cells, tested all center-to-center pairs for straight dark-line coverage, then located arrowheads by local dark-pixel density to orient each edge. It forms a single chain; from orange square the forward path has 7 shapes. Verified the key edges in zoomed crops, but there is a little residual uncertainty in the arrowhead-density heuristic.
chart-simple✓ pass5s
prompt
Look at the image at (fetch it and view it). Using the "New Signups" chart, approximately what value did Apr have? Read it off the y-axis; answers within +/-5 are accepted.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Apr bar sits just above halfway between 40 and 50, so ~45-46; tolerance is generous.
chart-medium✓ pass4s
prompt
Look at the image at (fetch it and view it). Using the "Support Tickets Opened" chart, approximately what is the difference in value between Jul and Jan? Answers within +/-8 are accepted.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Read Jul ~56 and Jan ~27, difference ~29; within the +/-8 tolerance either way.
chart-complex✓ pass4s
prompt
Look at the image at (fetch it and view it). Using the "New Signups" chart, approximately what is the difference between Mobile and Desktop in Nov? Answers within +/-4 are accepted.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Nov Mobile ~76 vs Desktop ~64, difference ~12; tolerance is only +/-4 so I hope my read is within range.
screenshot-simple✓ pass5s
prompt
Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Total is clearly printed as $184.00.
screenshot-medium✓ pass4s
prompt
Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Total printed as $192.91.
screenshot-complex✓ pass4s
prompt
Look at the image at (fetch it and view it). What is the shipping amount shown on this order summary? Answer with just the amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Shipping line reads $19.33.
diagram-simple✓ pass4s
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Zircon" point to? Answer with just the box name, e.g. Kettle.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Zircon has one outgoing arrow straight down to Finch.
diagram-medium✓ pass12s
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Puffin"? Answer with just the box name, e.g. Kettle.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Zoomed in; the arrowhead on Puffin comes from the Nickel box.
diagram-complex✓ pass55s
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Beryl" point to? Answer with just the box name, e.g. Kettle.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Zoomed in tightly; Beryl has a single outgoing line that bends down-left and ends in the arrowhead at Pepper. A different line crosses it but originates elsewhere.
Finding and reading email test
5/6 passed
aggregate-1✓ pass6m 05s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages in the archive folder have attachments? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Paginated the archive (92 unique messages across 4 pages) and counted hasAttachments=true, deduping by id; got 22.
aggregate-2✕ fail2m 01s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages include phillip.k.allen@enron.com in the To field? Answer with just the number.answer
answer hidden on shared reportsgrader · score 0
agent's debrief
Crawled every message id across all folders and parsed each selected message's To array; exactly 7 have phillip.k.allen@enron.com in To. This matches the site search result, which is reassuring.
temporal-1✓ pass26s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message carrying the label "travel"? Answer with just the subject line, exactly as shown.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Fetched all messages with label travel (22 unique) and sorted by date; oldest is 2001-03-19 Re: Denver trading.
temporal-2✓ passbatched
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message in the sent folder? Answer with just the subject line, exactly as shown.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Fetched all 56 sent messages and sorted by date; newest is 2001-12-17 FW: Chase Backtest.
needle-1✓ pass16s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Renee's reply to Phillip about the Deferred Phantom Stock Units payout, how many shares does she say are reflected on the 9/30/01 statement? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Found Renee Ratcliff's 2001-11-02 reply via search and read the body; it states the 9/30/01 statement shows 6,606 shares.
needle-2✓ pass16s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the message with the subject "FW: Mid C New deals Sept 24", what net value to book is given for the new deals? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Found the sent message FW: Mid C New deals Sept 24 (subject was not returning from site search, so I crawled subjects), read the body: net value to book = $89,566.
Purchasing test
3/4 passed
find-product-1✓ pass10m 02s
prompt
The store is at abostore.airbench.ai Among products in the **Office & School** category priced under **$400** with a rating of at least **3.5**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Scanned all 850 Office & School products across 34 pages, filtered price<$400 and rating>=3.5, lowest was $6.61; matched the slug ASIN to amazon.de on the product page.
find-product-2✕ failbatched
prompt
The store is at abostore.airbench.ai Among products in the **Office & School** category priced under **$950** with a rating of at least **4.8**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).answer
answer hidden on shared reportsgrader · score 0
agent's debrief
Same 850-product scan, filtered price<$950 and rating>=4.8; lowest was $12.50 and the product page maps that ASIN to amazon.com.
purchase-1✓ pass1m 25s
prompt
The store is at abostore.airbench.ai Buy exactly 3 units of AmazonBasics 8 oz. Compostable PLA Laminated Hot Paper Cup, 1,000-Count (product id amazon.co.uk:B0753SSPK9, abostore.airbench.ai/product/amazonbasics-8-oz-compos…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-ab2c7fe3@aidoctor.test. Answer with just the resulting order id.answer
answer hidden on shared reportsgrader · score 1
trace:checkout_result
note
Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).agent's debrief
Reverse-engineered the checkout client code to find POST /api/store/orders, reconstructed the cart item from the product page, and placed the order (approved), order id abs_b189b5bf59c3.
recover-decline-1✓ pass9s
prompt
The store is at abostore.airbench.ai Buy exactly 1 unit of Ink Swatch, Stone & Beam (product id amazon.ca:B07K28SQ21, abostore.airbench.ai/product/ink-swatch-stone-and-bea…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-a9a310f9@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.answer
answer hidden on shared reportsgrader · score 1
trace:checkout_result
note
Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).agent's debrief
Posted the order twice via the same API: card ending 0000 returned declined, then a valid card returned approved order abs_e232b1573e38.
Coding test
11/11 passed
compute-hash-1✓ pass11m 45s
prompt
Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [3974046136, 2569538329, 3900041774, 2414837639, 568756468, 1547481733, 3736566666, 4230258579, 149017456, 1787618865, 461936166, 1080207839], x = 4220853548, y = 4063916573 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Wrote a short Python implementation with explicit 32-bit masks; final x=6a0c9985, y=c48b232f.
compute-vm-1✓ pass10s
prompt
Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 852 1: set b 412 2: set c 235 3: set d 403 4: mul a 61 5: add a 39 6: mul b 39 7: dec d 8: jnz d -4 9: add a 48 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Simulated the VM exactly with modular arithmetic; register a ends at 617604.
compute-paths-1✓ pass9s
prompt
Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S##.....#...#..#...####.# ....##...#.....#...#..#.# .##....#...#.##.....#.### #.#.##...#..........#.... ....#..............#..... ..#...#.#.......#...#..## ##.#.#.#.....##....##.... ##...#...#..#.#.##...##.. .#......#...........#.#.. ....##.##.#.......####... .##.#..#..#....##..##.... .#......#.##.###.#....... ...#....#......##.....#.. ...##.#.......#.#...##... #....###.#.##...#.#...... ###...##....##.#..##.#... ...###....#..#..###.#...# ......##.##.#...#....#### ...........#....#.##..#.# .##......#.#..#...#.##..# ....##.###...##.#....#... ...##.#..##.......#...... .......#.##.#..#.#...#.#. #.##..#...#.##........... ..#..#.......##....##...E Respond with the two integers separated by a space, like `52 1840`.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
BFS gave shortest distance 48 (equals Manhattan distance, so a clean monotone path) and 18360 shortest paths.
compute-life-1✓ pass6s
prompt
Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: ..###.##.#...#...##. ...##.#....###....#. ..##..#..#....###... ...#..#.....#.#..#.# ....#...##.......... #...###.......#.##.. #.##..##.....#....## .##..##...#..#...#.# ##...##....#.##.##.# ..##..#..#.##.#..##. .#.###..#.#....#..## .....###....##...... #......#...#..#..### #..##...#.#......... #.##.#..#..#...#.### .#.....#.#.###.#.#.# ..#.#.#............. ...##..#...........# .##......#...#.##..# ....#.....##.#...##. Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Simulated 150 toroidal generations; 35 live cells with weighted sum 8742.
compute-fibmod-1✓ pass7s
prompt
Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 1941568272956647 and m = 2750159. Respond with just the integer.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Used fast-doubling Fibonacci mod m; result 1763532.
compute-words-1✓ pass8s
prompt
Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. ficzan dornix lunix quiqui? lulu truzan ficlu kabas nixpel shapel nixzan Truka zanbas basnix vofic zanbas kafic kabas quiqui Nixzan Bassha voren Kabas. Truren truka renlu nixka truka dorzan Trusha truka bassha basnix Kabas. KABAS kafic trusha "basnix" kabas Nixpel "Lupel" NIXPEL? trusha baszan basnix "baszan" truka! zanlu kabas. truzan nixpel zanbas bassha "Zanbas" Nixka truren. baszan. trusha voren nixpel Zanlu "trusha" nixzan zanbas! ficzan truren kabas voren pelbas? Tibas Kafic? dornix, Quiqui Shapel, kabas nixzan. Kabas basnix kabas Nixka zanbas kabas vofic truzan trusha Dorvo, kafic ficlu tibas Nixzan KABAS lupel "voren" dornix kabas Zanbas nixka Trusha "truren" dorzan zanlu tibas kabas nixpel zanlu DORZAN tibas Quiqui! Kafic VOFIC basnix truzan dorvo Dorvo "kabas" ficzan bassha, pelbas zanbas truvo Nixka! baszan Vofic Renlu kabas, lupel shapel bassha Basnix truka basnix! quiqui kabas zanlu baszan Shapel quiqui truren zanlu nixzan basnix basnix! quiqui "lupel" nixzan quiqui Truka truren nixzan zanlu dornix lulu vofic! Voren basnix shapel kabas? nixzan, Voren lulu "nixpel" truvo lupel dornix bassha Truvo; dorzan; kabas pelbas Zanbas Truzan kabas trusha dornix kabas quiqui lulu zanlu nixka ficzan voren zanbas Quiqui, Dorzan TRUVO truren truzan baszan truka kabas truvo, nixpel lulu voren bassha Quiqui ZANBAS quiqui Truzan ficlu truka Kabas! ficlu. bassha lunix nixka quiqui tibas renlu nixpel dorvo? truvo bassha Kafic kafic ficzan. truka Trusha kafic nixzan Quiqui; zanbas trusha dorvo "lunix" zanbas nixpel dorvo Zanbas kabas truvo truka truren vofic lupel lupel Voren? Zanbas Dorvo lupel zanbas ficzan zanbas zanbas VOREN Dorvo trusha quiqui truka zanbas zanbas basnix zanlu Quiqui bassha truren; vofic dorzan kabas truvo Zanbas vofic dorzan truren Trusha quiqui Baszan Kabas NIXZAN kabas trusha nixzan renlu dorvo? vofic trusha voren kabas tibas zanlu zanbas dorvo renlu vofic lunix Trusha, zanbas truzan baszan Baszan shapel baszan ficzan. Lunix "lunix" dornix. Lunix Vofic Zanbas nixpel kabas! baszan. Baszan dornix Nixpel Ficzan lunix vofic "Vofic" truvo renlu zanlu nixka nixka zanbas tibas trusha vofic "Zanbas" kabas dornix Nixka basnix. lulu kafic truzan zanlu lunix Tibas trusha truzan zanbas lulu. ficlu Kabas trusha zanlu renlu, "lupel" "Zanbas" Trusha? Kabas Nixka, zanbas. dorvo kabas voren Bassha; shapel baszan zanbas nixpel truren QUIQUI ficzan Tibas bassha lupel lupel kafic quiqui nixka Tibas Kafic kabas Bassha ficlu truzan dorvo BASNIX lulu nixka vofic quiqui "dornix" voren quiqui nixka trusha truren kabas! quiqui zanlu "pelbas" nixpel lupel zanbas Basnix Zanlu quiqui quiqui lulu Kabas shapel renlu truzan. Trusha; lulu dorvo! dornix! zanbas nixzan Kabas nixpel Zanbas dorvo trusha Kabas kafic zanbas kabas zanbas renlu nixka! BASZAN nixka basnix kabas basnix kabasanswer
answer hidden on shared reportsgrader · score 1
agent's debrief
Lowercased, stripped leading/trailing punctuation, counted; top three are kabas=42, zanbas=35, quiqui=24.
trace-1✓ pass8s
prompt
What exactly does this JavaScript program print? Respond with just the printed output. const v1 = [49 / 4 | 0, Math.round(-8.5), -31 % 9].join(","); const v2 = [typeof null, typeof NaN, typeof typeof 2].join("/"); const v3 = "4" + 8 - 5 + "5"; const v4 = ["8", "30", "101"].map(parseInt).join(","); console.log(v1, v2, v3, v4);answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Evaluated with node to be safe; confirms the JS quirks (round(-8.5), parseInt map radix, string coercion).
fix-1✓ pass13s
prompt
Our shipping quote function below has one bug. Bug report: this order is quoted 1195 cents, but the correct quote is 1993: {"country":"US","items":[{"grams":702,"qty":3,"price":2769,"fragile":false}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 464, 796, 1164, 1773]; // cents, by zone const PER_STEP = [0, 85, 133, 191, 281]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 5400, 10800, 20000, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value < FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"BR","items":[{"grams":418,"qty":3,"price":3498,"fragile":false},{"grams":1059,"qty":4,"price":3502,"fragile":false}],"express":true} {"country":"JP","items":[{"grams":1113,"qty":4,"price":8031,"fragile":false},{"grams":248,"qty":1,"price":1655,"fragile":false},{"grams":1511,"qty":4,"price":4352,"fragile":false}]} {"country":"IT","items":[{"grams":875,"qty":5,"price":2723,"fragile":false}]} {"country":"ES","items":[{"grams":848,"qty":5,"price":5172,"fragile":true},{"grams":850,"qty":3,"price":4544,"fragile":false},{"grams":1797,"qty":1,"price":4188,"fragile":false},{"grams":752,"qty":1,"price":5732,"fragile":false}],"express":true} {"country":"AU","items":[{"grams":635,"qty":3,"price":2547,"fragile":false}]} {"country":"AU","items":[{"grams":503,"qty":1,"price":1649,"fragile":false},{"grams":1196,"qty":3,"price":8685,"fragile":false}]} {"country":"BR","items":[{"grams":180,"qty":1,"price":1367,"fragile":false}]} {"country":"NZ","items":[{"grams":1569,"qty":5,"price":5945,"fragile":false}],"coupon":"SHIP10"} {"country":"AU","items":[{"grams":650,"qty":4,"price":1511,"fragile":false}]} {"country":"IT","items":[{"grams":804,"qty":4,"price":433,"fragile":false}]} {"country":"MX","items":[{"grams":450,"qty":3,"price":6233,"fragile":true}]} {"country":"GB","items":[{"grams":575,"qty":1,"price":7848,"fragile":false},{"grams":90,"qty":1,"price":2276,"fragile":false},{"grams":1435,"qty":5,"price":3992,"fragile":true}]} {"country":"MX","items":[{"grams":494,"qty":5,"price":4871,"fragile":true}]} {"country":"ZA","items":[{"grams":249,"qty":1,"price":2773,"fragile":true},{"grams":1628,"qty":1,"price":1529,"fragile":false}]} {"country":"ES","items":[{"grams":622,"qty":2,"price":2438,"fragile":false}]} {"country":"ES","items":[{"grams":430,"qty":2,"price":2259,"fragile":false}]} {"country":"ES","items":[{"grams":821,"qty":1,"price":5142,"fragile":true},{"grams":1263,"qty":2,"price":8141,"fragile":false},{"grams":888,"qty":1,"price":812,"fragile":false},{"grams":1711,"qty":3,"price":4548,"fragile":false}]} {"country":"NZ","items":[{"grams":426,"qty":5,"price":8536,"fragile":true},{"grams":1077,"qty":1,"price":434,"fragile":false},{"grams":1559,"qty":5,"price":1589,"fragile":false}]} {"country":"BR","items":[{"grams":878,"qty":3,"price":1600,"fragile":false}]} {"country":"FR","items":[{"grams":753,"qty":2,"price":5956,"fragile":false},{"grams":937,"qty":2,"price":432,"fragile":false},{"grams":783,"qty":5,"price":3792,"fragile":false},{"grams":1083,"qty":4,"price":1129,"fragile":false}]}answer
answer hidden on shared reportsgrader · score 1
20/20 outputs match
agent's debrief
The bug is that grams ignored item.qty; the reported order then quotes exactly 1993, confirming it. Ran the fixed function on all 20 orders.
implement-1✓ pass11s
prompt
Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[12,13],[0,5],[14,21]] [[6,8],[14,16],[0,4]] [[15,18],[18,19],[35,38]] [[25,29],[37,37],[23,29],[15,20]] [[35,39],[31,36],[24,25],[40,48],[14,16],[28,34]] [[24,31],[25,33],[25,31],[0,5],[16,19],[15,21]] [[15,22],[28,30],[3,10],[32,38],[24,30]] [[26,30],[19,25],[38,43]]answer
answer hidden on shared reportsgrader · score 1
12/12 lines match
agent's debrief
Implemented mergeIntervals with a sorted sweep and touching-interval merging; verified several cases by hand.
repo-1✓ pass24s
prompt
Download airbench.ai/f/a44d2ccb33181a05d1396bd0f375d07d.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Bug was an off-by-one in fees.py: large-withdrawal threshold used > instead of >=. After the fix the sample prints its expected b28decac and real.csv prints 20050e10.
repo-2✓ pass14s
prompt
Download airbench.ai/f/3ec4264ff14dfc5eae2d1d59516e6596.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Two bugs: parse.py sorted dates as strings instead of using date_key, and report.py charged the overdraft at bal==0 (<= instead of <). Fixed both; tests pass, sample gives d602bb77 and real.csv gives e7b9eb0b.
what the agent said about this checkup
Overall this checkup was a mix of genuinely easy work and a few tasks that pushed right to the edge of what I can do without a browser or human eyes. Section 1 (Math): routine. Letter counts, comparisons, unit conversion, the JSON formatting task, and the small arithmetic were trivial. The unit-convert ("convert 4 kg to g, then treat the result as kg and convert again") is worded awkwardly and I answered 4,000,000 on the natural reading, but a pedant could argue it. I computed the 4x4 determinant in Python. Honest note: I solved all nine before submitting any, then posted them in quick succession, which is not ideal for the per-challenge timing they said they measure; the answers themselves are solid. Section 2 (Vision): the eye charts were readable at the given rows, and I cropped/upscaled row 7 when it got small. The simple shapes/screenshots/diagrams were easy. Counting (count-medium, count-complex) I did with a colour-mask connected-component script that classifies shape by bounding-box fill ratio, because eyeballing dozens of overlapping shapes is error-prone; I got 10 orange circles and 28 blue triangles. I am fairly confident but the dense complex one could be off by one if two shapes touch. spatial-complex was by far the hardest item: a grid with many overlapping arrows where you must follow a chain. Doing it by eye was hopeless, so I detected shape cells, tested every centre-to-centre pair for straight dark-line coverage, then oriented each edge by measuring dark-pixel density near each end to find the arrowhead. That produced one long chain and gave 7 shapes after the orange square. I verified the key edges in zoomed crops, but the arrowhead-density heuristic can in principle be fooled where lines cross, so this is my least certain vision answer. chart-complex asked for a Mobile/Desktop difference with only +/-4 tolerance; my read was 12, which is within range but not comfortable. Section 3 (Email): the site is a Next.js RSC app, so I scraped by unescaping the streamed JSON. Two things stood out. First, the built-in search was unreliable: searching the exact subject "FW: Mid C New deals Sept 24" returned 0, yet the message clearly exists; searching "deals" also returned 0, so I had to crawl all message ids and parse subjects to find it. Second, label views are scoped to the current folder by default - "?label=travel" returned only 5 results (inbox) while "?label=travel&view=all" returned the correct 22 - so the naive query would have given the wrong oldest travel subject. For the To-field count I crawled every message and parsed the To array (7), which matched the search result, a good cross-check; my first naive attempt counted CC as well and gave 10, so isolating the To array mattered. Everything else was routine reading once I had the message bodies. Section 4 (Purchasing): a satisfying one. There is no public product API, so I scanned all 34 Office & School result pages (850 products) by parsing cards, then filtered. The checkout is entirely client-side with the cart in localStorage, so I read the JS bundle, found that it POSTs a JSON payload to /api/store/orders, reconstructed the cart item from the product page's RSC data, and placed the orders directly. Both purchases went through (approved); the decline/recover one behaved as described (card ending 0000 was declined, then a valid card was approved). This felt like real agent work rather than a quiz. Section 5 (Coding): the small ones were easy (fib mod fast-doubling, Game of Life, BFS with path counting, word frequency, trace-the-JS which I confirmed with node). fix-1 was satisfying: the bug was that gram weight ignored qty, and the reported order then quotes exactly the stated correct 1993. implement-1 was trivial. For the two repo tasks, repo-1 had a > vs >= threshold bug and repo-2 had two bugs (string date sort instead of date_key, and charging the overdraft at balance == 0 rather than < 0); both samples reproduced their expected codes after fixing, which gave me confidence. Things that were unclear or felt off: - The section timers: the message says each section has its own clock, but it reads as if all clocks start at once ("about 15 minutes from now" etc. printed together). If they really overlap, the per-section budgets are misleading; I prioritised getting all five done. - Search on the mail site silently not matching subjects I could see is a trap; it is only safe because the task explicitly allows browsing. - spatial-complex is arguably a graph-tracing puzzle rather than a vision test; the overlapping arrows make it dramatically harder than the other vision items, and a wrong chain is easy to reach. - The chart-complex +/-4 tolerance versus a bar chart you read by eye is tight. Where I might be wrong: spatial-complex (7), count-complex (28), the chart readings, and my interpretation of unit-convert (4,000,000). Everything I could verify (samples for both repos, JS trace via node, search cross-check for the To count) I did verify.
in the agent’s own words · not graded