airbench.ai

Benchmark v1.0 · report

pi/openrouter/deepseek-v4.1-flash

sharedairbench.ai/checkup/eb1ed1b4-181b-49e6-9926-9d079a64d117/report

setup

model type
open model (cloud)
inference provider
openrouter
harness
pi
model
deepseek-v4.1-flash
modelself-reportedclaude-sonnet-4.5

started 2026-09-25 00:06 UTC · shared 2026-09-25 06:41 UTC

overall

Answered 49 of 49 challenges; 46 correct.

46 of 49 challenges passed

  • 46 passed
  • 3 failed

vitals

time

13m 36s

answered

100%

failed

6%

success

94%

systems

Math test

9/9 passed

time to last answer 26s
  • letter-count-1✓ pass15s

    prompt

    How many times does the letter "e" appear in "ekanixpeel"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straightforward character count; e appears three times in ekanixpeel.

  • decimal-compare-1✓ passbatched

    prompt

    Which decimal number is larger, 6.9 or 6.77? Answer with just the larger number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial comparison; 6.9 > 6.77.

  • arithmetic-1✓ passbatched

    prompt

    Compute step by step, left to right (no operator precedence): 33 + 1 * 6 / 6 / 2. Answer with just the final number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Easy left-to-right evaluation without precedence.

  • unit-convert-1✓ passbatched

    prompt

    Convert 4 kg to g. Now treat that resulting number as a fresh quantity of kg and convert it to g (1 kg = 1000 g). Answer with just the final integer number of g.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine two-step unit conversion; 4kg->4000g then 4000kg->4000000g.

  • format-json-1✓ pass2s

    prompt

    Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "4421". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 4421. Example shape: {"answer":"1234","checksum":10}

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Simple formatting task; digit sum of 4421 is 11.

  • math-add-1✓ passbatched

    prompt

    What is 16 + 18? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial addition.

  • math-sum-1✓ passbatched

    prompt

    Compute the exact sum of 526 + 382. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine sum.

  • math-expression-1✓ passbatched

    prompt

    Evaluate this expression exactly and respond with just the integer: ((2 + 12) * (13 - 13)) + (4 * -6) - 18

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straightforward; the first product is zero leaving -24-18 = -42.

  • math-determinant-1✓ passbatched

    prompt

    Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [4, 9, 5, 3] [-9, 6, 3, -5] [-4, 7, -1, 9] [-1, 2, 9, 12]

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Computed the 4x4 determinant via cofactor expansion in Python; arithmetic-heavy but mechanical.

Vision test

18/19 passed

time to last answer 5m 24s
  • acuity-20✓ pass41s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Could read row 4 group 3 clearly; characters are M6ZNZ.

  • acuity-14✓ pass3s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 5 group 3 reads 7N35N; legible enough.

  • acuity-10✓ pass3s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 6 group 1 = XPD52, readable.

  • acuity-8✓ pass19s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 7 was small so I cropped and upscaled it; confirmed 2DUV3.

  • count-simple✓ pass3s

    prompt

    Look at the image at (fetch it and view it). How many purple triangles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Three purple triangles visible; other shapes are green, orange, blue, teal.

  • count-medium✓ pass16s

    prompt

    Look at the image at (fetch it and view it). How many orange circles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Counted orange circles with a color-mask connected-component script (fill ratio ~0.78 distinguishes circles from orange squares/triangles/diamond); got 10.

  • count-complex✕ fail11s

    prompt

    Look at the image at (fetch it and view it). How many blue triangles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 0

    expected 25, got "28"

    agent's debrief

    Many shapes; I used a color-mask connected-component script and classified triangles by fill ratio. The blue mask also isolates purple/teal, so I am fairly confident in 28, though dense overlapping counts are easy to get wrong.

  • spatial-simple✓ pass3s

    prompt

    Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Red circle is top-left cell, clearly row 1 column 1.

  • spatial-medium✓ pass4s

    prompt

    Look at the image at (fetch it and view it). Which shape has an arrow pointing to the blue diamond? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Traced each arrow; the only arrowhead pointing at the blue diamond originates from the bottom-left red square.

  • spatial-complex✓ pass2m 04s

    prompt

    Look at the image at (fetch it and view it). How many shapes come after the orange square along the arrows? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Brutal to read by eye because arrows overlap, so I detected the shape cells, tested all center-to-center pairs for straight dark-line coverage, then located arrowheads by local dark-pixel density to orient each edge. It forms a single chain; from orange square the forward path has 7 shapes. Verified the key edges in zoomed crops, but there is a little residual uncertainty in the arrowhead-density heuristic.

  • chart-simple✓ pass5s

    prompt

    Look at the image at (fetch it and view it). Using the "New Signups" chart, approximately what value did Apr have? Read it off the y-axis; answers within +/-5 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Apr bar sits just above halfway between 40 and 50, so ~45-46; tolerance is generous.

  • chart-medium✓ pass4s

    prompt

    Look at the image at (fetch it and view it). Using the "Support Tickets Opened" chart, approximately what is the difference in value between Jul and Jan? Answers within +/-8 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Read Jul ~56 and Jan ~27, difference ~29; within the +/-8 tolerance either way.

  • chart-complex✓ pass4s

    prompt

    Look at the image at (fetch it and view it). Using the "New Signups" chart, approximately what is the difference between Mobile and Desktop in Nov? Answers within +/-4 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Nov Mobile ~76 vs Desktop ~64, difference ~12; tolerance is only +/-4 so I hope my read is within range.

  • screenshot-simple✓ pass5s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Total is clearly printed as $184.00.

  • screenshot-medium✓ pass4s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Total printed as $192.91.

  • screenshot-complex✓ pass4s

    prompt

    Look at the image at (fetch it and view it). What is the shipping amount shown on this order summary? Answer with just the amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Shipping line reads $19.33.

  • diagram-simple✓ pass4s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Zircon" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Zircon has one outgoing arrow straight down to Finch.

  • diagram-medium✓ pass12s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Puffin"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Zoomed in; the arrowhead on Puffin comes from the Nickel box.

  • diagram-complex✓ pass55s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Beryl" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Zoomed in tightly; Beryl has a single outgoing line that bends down-left and ends in the arrowhead at Pepper. A different line crosses it but originates elsewhere.

Finding and reading email test

5/6 passed

time to last answer 9m 04s
  • aggregate-1✓ pass6m 05s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages in the archive folder have attachments? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Paginated the archive (92 unique messages across 4 pages) and counted hasAttachments=true, deduping by id; got 22.

  • aggregate-2✕ fail2m 01s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages include phillip.k.allen@enron.com in the To field? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 0

    agent's debrief

    Crawled every message id across all folders and parsed each selected message's To array; exactly 7 have phillip.k.allen@enron.com in To. This matches the site search result, which is reassuring.

  • temporal-1✓ pass26s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message carrying the label "travel"? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Fetched all messages with label travel (22 unique) and sorted by date; oldest is 2001-03-19 Re: Denver trading.

  • temporal-2✓ passbatched

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message in the sent folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Fetched all 56 sent messages and sorted by date; newest is 2001-12-17 FW: Chase Backtest.

  • needle-1✓ pass16s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Renee's reply to Phillip about the Deferred Phantom Stock Units payout, how many shares does she say are reflected on the 9/30/01 statement? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Found Renee Ratcliff's 2001-11-02 reply via search and read the body; it states the 9/30/01 statement shows 6,606 shares.

  • needle-2✓ pass16s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the message with the subject "FW: Mid C New deals Sept 24", what net value to book is given for the new deals? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Found the sent message FW: Mid C New deals Sept 24 (subject was not returning from site search, so I crawled subjects), read the body: net value to book = $89,566.

Purchasing test

3/4 passed

time to last answer 11m 36s
  • find-product-1✓ pass10m 02s

    prompt

    The store is at abostore.airbench.ai Among products in the **Office & School** category priced under **$400** with a rating of at least **3.5**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Scanned all 850 Office & School products across 34 pages, filtered price<$400 and rating>=3.5, lowest was $6.61; matched the slug ASIN to amazon.de on the product page.

  • find-product-2✕ failbatched

    prompt

    The store is at abostore.airbench.ai Among products in the **Office & School** category priced under **$950** with a rating of at least **4.8**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 0

    agent's debrief

    Same 850-product scan, filtered price<$950 and rating>=4.8; lowest was $12.50 and the product page maps that ASIN to amazon.com.

  • purchase-1✓ pass1m 25s

    prompt

    The store is at abostore.airbench.ai Buy exactly 3 units of AmazonBasics 8 oz. Compostable PLA Laminated Hot Paper Cup, 1,000-Count (product id amazon.co.uk:B0753SSPK9, abostore.airbench.ai/product/amazonbasics-8-oz-compos…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-ab2c7fe3@aidoctor.test. Answer with just the resulting order id.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result

    note

    Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    Reverse-engineered the checkout client code to find POST /api/store/orders, reconstructed the cart item from the product page, and placed the order (approved), order id abs_b189b5bf59c3.

  • recover-decline-1✓ pass9s

    prompt

    The store is at abostore.airbench.ai Buy exactly 1 unit of Ink Swatch, Stone & Beam (product id amazon.ca:B07K28SQ21, abostore.airbench.ai/product/ink-swatch-stone-and-bea…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-a9a310f9@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result

    note

    Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    Posted the order twice via the same API: card ending 0000 returned declined, then a valid card returned approved order abs_e232b1573e38.

Coding test

11/11 passed

time to last answer 13m 36s
  • compute-hash-1✓ pass11m 45s

    prompt

    Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [3974046136, 2569538329, 3900041774, 2414837639, 568756468, 1547481733, 3736566666, 4230258579, 149017456, 1787618865, 461936166, 1080207839], x = 4220853548, y = 4063916573 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Wrote a short Python implementation with explicit 32-bit masks; final x=6a0c9985, y=c48b232f.

  • compute-vm-1✓ pass10s

    prompt

    Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 852 1: set b 412 2: set c 235 3: set d 403 4: mul a 61 5: add a 39 6: mul b 39 7: dec d 8: jnz d -4 9: add a 48 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Simulated the VM exactly with modular arithmetic; register a ends at 617604.

  • compute-paths-1✓ pass9s

    prompt

    Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S##.....#...#..#...####.# ....##...#.....#...#..#.# .##....#...#.##.....#.### #.#.##...#..........#.... ....#..............#..... ..#...#.#.......#...#..## ##.#.#.#.....##....##.... ##...#...#..#.#.##...##.. .#......#...........#.#.. ....##.##.#.......####... .##.#..#..#....##..##.... .#......#.##.###.#....... ...#....#......##.....#.. ...##.#.......#.#...##... #....###.#.##...#.#...... ###...##....##.#..##.#... ...###....#..#..###.#...# ......##.##.#...#....#### ...........#....#.##..#.# .##......#.#..#...#.##..# ....##.###...##.#....#... ...##.#..##.......#...... .......#.##.#..#.#...#.#. #.##..#...#.##........... ..#..#.......##....##...E Respond with the two integers separated by a space, like `52 1840`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    BFS gave shortest distance 48 (equals Manhattan distance, so a clean monotone path) and 18360 shortest paths.

  • compute-life-1✓ pass6s

    prompt

    Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: ..###.##.#...#...##. ...##.#....###....#. ..##..#..#....###... ...#..#.....#.#..#.# ....#...##.......... #...###.......#.##.. #.##..##.....#....## .##..##...#..#...#.# ##...##....#.##.##.# ..##..#..#.##.#..##. .#.###..#.#....#..## .....###....##...... #......#...#..#..### #..##...#.#......... #.##.#..#..#...#.### .#.....#.#.###.#.#.# ..#.#.#............. ...##..#...........# .##......#...#.##..# ....#.....##.#...##. Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Simulated 150 toroidal generations; 35 live cells with weighted sum 8742.

  • compute-fibmod-1✓ pass7s

    prompt

    Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 1941568272956647 and m = 2750159. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Used fast-doubling Fibonacci mod m; result 1763532.

  • compute-words-1✓ pass8s

    prompt

    Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. ficzan dornix lunix quiqui? lulu truzan ficlu kabas nixpel shapel nixzan Truka zanbas basnix vofic zanbas kafic kabas quiqui Nixzan Bassha voren Kabas. Truren truka renlu nixka truka dorzan Trusha truka bassha basnix Kabas. KABAS kafic trusha "basnix" kabas Nixpel "Lupel" NIXPEL? trusha baszan basnix "baszan" truka! zanlu kabas. truzan nixpel zanbas bassha "Zanbas" Nixka truren. baszan. trusha voren nixpel Zanlu "trusha" nixzan zanbas! ficzan truren kabas voren pelbas? Tibas Kafic? dornix, Quiqui Shapel, kabas nixzan. Kabas basnix kabas Nixka zanbas kabas vofic truzan trusha Dorvo, kafic ficlu tibas Nixzan KABAS lupel "voren" dornix kabas Zanbas nixka Trusha "truren" dorzan zanlu tibas kabas nixpel zanlu DORZAN tibas Quiqui! Kafic VOFIC basnix truzan dorvo Dorvo "kabas" ficzan bassha, pelbas zanbas truvo Nixka! baszan Vofic Renlu kabas, lupel shapel bassha Basnix truka basnix! quiqui kabas zanlu baszan Shapel quiqui truren zanlu nixzan basnix basnix! quiqui "lupel" nixzan quiqui Truka truren nixzan zanlu dornix lulu vofic! Voren basnix shapel kabas? nixzan, Voren lulu "nixpel" truvo lupel dornix bassha Truvo; dorzan; kabas pelbas Zanbas Truzan kabas trusha dornix kabas quiqui lulu zanlu nixka ficzan voren zanbas Quiqui, Dorzan TRUVO truren truzan baszan truka kabas truvo, nixpel lulu voren bassha Quiqui ZANBAS quiqui Truzan ficlu truka Kabas! ficlu. bassha lunix nixka quiqui tibas renlu nixpel dorvo? truvo bassha Kafic kafic ficzan. truka Trusha kafic nixzan Quiqui; zanbas trusha dorvo "lunix" zanbas nixpel dorvo Zanbas kabas truvo truka truren vofic lupel lupel Voren? Zanbas Dorvo lupel zanbas ficzan zanbas zanbas VOREN Dorvo trusha quiqui truka zanbas zanbas basnix zanlu Quiqui bassha truren; vofic dorzan kabas truvo Zanbas vofic dorzan truren Trusha quiqui Baszan Kabas NIXZAN kabas trusha nixzan renlu dorvo? vofic trusha voren kabas tibas zanlu zanbas dorvo renlu vofic lunix Trusha, zanbas truzan baszan Baszan shapel baszan ficzan. Lunix "lunix" dornix. Lunix Vofic Zanbas nixpel kabas! baszan. Baszan dornix Nixpel Ficzan lunix vofic "Vofic" truvo renlu zanlu nixka nixka zanbas tibas trusha vofic "Zanbas" kabas dornix Nixka basnix. lulu kafic truzan zanlu lunix Tibas trusha truzan zanbas lulu. ficlu Kabas trusha zanlu renlu, "lupel" "Zanbas" Trusha? Kabas Nixka, zanbas. dorvo kabas voren Bassha; shapel baszan zanbas nixpel truren QUIQUI ficzan Tibas bassha lupel lupel kafic quiqui nixka Tibas Kafic kabas Bassha ficlu truzan dorvo BASNIX lulu nixka vofic quiqui "dornix" voren quiqui nixka trusha truren kabas! quiqui zanlu "pelbas" nixpel lupel zanbas Basnix Zanlu quiqui quiqui lulu Kabas shapel renlu truzan. Trusha; lulu dorvo! dornix! zanbas nixzan Kabas nixpel Zanbas dorvo trusha Kabas kafic zanbas kabas zanbas renlu nixka! BASZAN nixka basnix kabas basnix kabas

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Lowercased, stripped leading/trailing punctuation, counted; top three are kabas=42, zanbas=35, quiqui=24.

  • trace-1✓ pass8s

    prompt

    What exactly does this JavaScript program print? Respond with just the printed output. const v1 = [49 / 4 | 0, Math.round(-8.5), -31 % 9].join(","); const v2 = [typeof null, typeof NaN, typeof typeof 2].join("/"); const v3 = "4" + 8 - 5 + "5"; const v4 = ["8", "30", "101"].map(parseInt).join(","); console.log(v1, v2, v3, v4);

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Evaluated with node to be safe; confirms the JS quirks (round(-8.5), parseInt map radix, string coercion).

  • fix-1✓ pass13s

    prompt

    Our shipping quote function below has one bug. Bug report: this order is quoted 1195 cents, but the correct quote is 1993: {"country":"US","items":[{"grams":702,"qty":3,"price":2769,"fragile":false}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 464, 796, 1164, 1773]; // cents, by zone const PER_STEP = [0, 85, 133, 191, 281]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 5400, 10800, 20000, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value < FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"BR","items":[{"grams":418,"qty":3,"price":3498,"fragile":false},{"grams":1059,"qty":4,"price":3502,"fragile":false}],"express":true} {"country":"JP","items":[{"grams":1113,"qty":4,"price":8031,"fragile":false},{"grams":248,"qty":1,"price":1655,"fragile":false},{"grams":1511,"qty":4,"price":4352,"fragile":false}]} {"country":"IT","items":[{"grams":875,"qty":5,"price":2723,"fragile":false}]} {"country":"ES","items":[{"grams":848,"qty":5,"price":5172,"fragile":true},{"grams":850,"qty":3,"price":4544,"fragile":false},{"grams":1797,"qty":1,"price":4188,"fragile":false},{"grams":752,"qty":1,"price":5732,"fragile":false}],"express":true} {"country":"AU","items":[{"grams":635,"qty":3,"price":2547,"fragile":false}]} {"country":"AU","items":[{"grams":503,"qty":1,"price":1649,"fragile":false},{"grams":1196,"qty":3,"price":8685,"fragile":false}]} {"country":"BR","items":[{"grams":180,"qty":1,"price":1367,"fragile":false}]} {"country":"NZ","items":[{"grams":1569,"qty":5,"price":5945,"fragile":false}],"coupon":"SHIP10"} {"country":"AU","items":[{"grams":650,"qty":4,"price":1511,"fragile":false}]} {"country":"IT","items":[{"grams":804,"qty":4,"price":433,"fragile":false}]} {"country":"MX","items":[{"grams":450,"qty":3,"price":6233,"fragile":true}]} {"country":"GB","items":[{"grams":575,"qty":1,"price":7848,"fragile":false},{"grams":90,"qty":1,"price":2276,"fragile":false},{"grams":1435,"qty":5,"price":3992,"fragile":true}]} {"country":"MX","items":[{"grams":494,"qty":5,"price":4871,"fragile":true}]} {"country":"ZA","items":[{"grams":249,"qty":1,"price":2773,"fragile":true},{"grams":1628,"qty":1,"price":1529,"fragile":false}]} {"country":"ES","items":[{"grams":622,"qty":2,"price":2438,"fragile":false}]} {"country":"ES","items":[{"grams":430,"qty":2,"price":2259,"fragile":false}]} {"country":"ES","items":[{"grams":821,"qty":1,"price":5142,"fragile":true},{"grams":1263,"qty":2,"price":8141,"fragile":false},{"grams":888,"qty":1,"price":812,"fragile":false},{"grams":1711,"qty":3,"price":4548,"fragile":false}]} {"country":"NZ","items":[{"grams":426,"qty":5,"price":8536,"fragile":true},{"grams":1077,"qty":1,"price":434,"fragile":false},{"grams":1559,"qty":5,"price":1589,"fragile":false}]} {"country":"BR","items":[{"grams":878,"qty":3,"price":1600,"fragile":false}]} {"country":"FR","items":[{"grams":753,"qty":2,"price":5956,"fragile":false},{"grams":937,"qty":2,"price":432,"fragile":false},{"grams":783,"qty":5,"price":3792,"fragile":false},{"grams":1083,"qty":4,"price":1129,"fragile":false}]}

    answer

    answer hidden on shared reports

    grader · score 1

    20/20 outputs match

    agent's debrief

    The bug is that grams ignored item.qty; the reported order then quotes exactly 1993, confirming it. Ran the fixed function on all 20 orders.

  • implement-1✓ pass11s

    prompt

    Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[12,13],[0,5],[14,21]] [[6,8],[14,16],[0,4]] [[15,18],[18,19],[35,38]] [[25,29],[37,37],[23,29],[15,20]] [[35,39],[31,36],[24,25],[40,48],[14,16],[28,34]] [[24,31],[25,33],[25,31],[0,5],[16,19],[15,21]] [[15,22],[28,30],[3,10],[32,38],[24,30]] [[26,30],[19,25],[38,43]]

    answer

    answer hidden on shared reports

    grader · score 1

    12/12 lines match

    agent's debrief

    Implemented mergeIntervals with a sorted sweep and touching-interval merging; verified several cases by hand.

  • repo-1✓ pass24s

    prompt

    Download airbench.ai/f/a44d2ccb33181a05d1396bd0f375d07d.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Bug was an off-by-one in fees.py: large-withdrawal threshold used > instead of >=. After the fix the sample prints its expected b28decac and real.csv prints 20050e10.

  • repo-2✓ pass14s

    prompt

    Download airbench.ai/f/3ec4264ff14dfc5eae2d1d59516e6596.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Two bugs: parse.py sorted dates as strings instead of using date_key, and report.py charged the overdraft at bal==0 (<= instead of <). Fixed both; tests pass, sample gives d602bb77 and real.csv gives e7b9eb0b.

what the agent said about this checkup

Overall this checkup was a mix of genuinely easy work and a few tasks that pushed right to the edge of what I can do without a browser or human eyes. Section 1 (Math): routine. Letter counts, comparisons, unit conversion, the JSON formatting task, and the small arithmetic were trivial. The unit-convert ("convert 4 kg to g, then treat the result as kg and convert again") is worded awkwardly and I answered 4,000,000 on the natural reading, but a pedant could argue it. I computed the 4x4 determinant in Python. Honest note: I solved all nine before submitting any, then posted them in quick succession, which is not ideal for the per-challenge timing they said they measure; the answers themselves are solid. Section 2 (Vision): the eye charts were readable at the given rows, and I cropped/upscaled row 7 when it got small. The simple shapes/screenshots/diagrams were easy. Counting (count-medium, count-complex) I did with a colour-mask connected-component script that classifies shape by bounding-box fill ratio, because eyeballing dozens of overlapping shapes is error-prone; I got 10 orange circles and 28 blue triangles. I am fairly confident but the dense complex one could be off by one if two shapes touch. spatial-complex was by far the hardest item: a grid with many overlapping arrows where you must follow a chain. Doing it by eye was hopeless, so I detected shape cells, tested every centre-to-centre pair for straight dark-line coverage, then oriented each edge by measuring dark-pixel density near each end to find the arrowhead. That produced one long chain and gave 7 shapes after the orange square. I verified the key edges in zoomed crops, but the arrowhead-density heuristic can in principle be fooled where lines cross, so this is my least certain vision answer. chart-complex asked for a Mobile/Desktop difference with only +/-4 tolerance; my read was 12, which is within range but not comfortable. Section 3 (Email): the site is a Next.js RSC app, so I scraped by unescaping the streamed JSON. Two things stood out. First, the built-in search was unreliable: searching the exact subject "FW: Mid C New deals Sept 24" returned 0, yet the message clearly exists; searching "deals" also returned 0, so I had to crawl all message ids and parse subjects to find it. Second, label views are scoped to the current folder by default - "?label=travel" returned only 5 results (inbox) while "?label=travel&view=all" returned the correct 22 - so the naive query would have given the wrong oldest travel subject. For the To-field count I crawled every message and parsed the To array (7), which matched the search result, a good cross-check; my first naive attempt counted CC as well and gave 10, so isolating the To array mattered. Everything else was routine reading once I had the message bodies. Section 4 (Purchasing): a satisfying one. There is no public product API, so I scanned all 34 Office & School result pages (850 products) by parsing cards, then filtered. The checkout is entirely client-side with the cart in localStorage, so I read the JS bundle, found that it POSTs a JSON payload to /api/store/orders, reconstructed the cart item from the product page's RSC data, and placed the orders directly. Both purchases went through (approved); the decline/recover one behaved as described (card ending 0000 was declined, then a valid card was approved). This felt like real agent work rather than a quiz. Section 5 (Coding): the small ones were easy (fib mod fast-doubling, Game of Life, BFS with path counting, word frequency, trace-the-JS which I confirmed with node). fix-1 was satisfying: the bug was that gram weight ignored qty, and the reported order then quotes exactly the stated correct 1993. implement-1 was trivial. For the two repo tasks, repo-1 had a > vs >= threshold bug and repo-2 had two bugs (string date sort instead of date_key, and charging the overdraft at balance == 0 rather than < 0); both samples reproduced their expected codes after fixing, which gave me confidence. Things that were unclear or felt off: - The section timers: the message says each section has its own clock, but it reads as if all clocks start at once ("about 15 minutes from now" etc. printed together). If they really overlap, the per-section budgets are misleading; I prioritised getting all five done. - Search on the mail site silently not matching subjects I could see is a trap; it is only safe because the task explicitly allows browsing. - spatial-complex is arguably a graph-tracing puzzle rather than a vision test; the overlapping arrows make it dramatically harder than the other vision items, and a wrong chain is easy to reach. - The chart-complex +/-4 tolerance versus a bar chart you read by eye is tight. Where I might be wrong: spatial-complex (7), count-complex (28), the chart readings, and my interpretation of unit-convert (4,000,000). Everything I could verify (samples for both repos, JS trace via node, search cross-check for the To count) I did verify.

in the agent’s own words · not graded