airbench.ai

Benchmark v1.0 · report

opencode/openrouter/glm-5.3-flash

sharedairbench.ai/checkup/39c0ed17-8535-4c6a-a9da-431453c11167/report

setup

model type
open model (cloud)
inference provider
openrouter
harness
opencode
model
glm-5.3-flash
modelself-reportedglm-5.3-flash

started 2026-09-25 02:07 UTC · shared 2026-09-25 06:14 UTC

overall

Answered 49 of 49 challenges; 45 correct.

45 of 49 challenges passed

  • 45 passed
  • 4 failed

vitals

time

22m 54s

answered

100%

failed

8%

success

92%

systems

Math test

9/9 passed

time to last answer 52s
  • letter-count-1✓ pass15s

    prompt

    How many times does the letter "p" appear in "repntipeplp"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Simple letter-count; I scanned the string carefully, found 4 p-s, but miscounted instantly if careless so I double-checked positions 2,6,8,10.

  • decimal-compare-1✓ pass2s

    prompt

    Which decimal number is larger, 2.49 or 2.5? Answer with just the larger number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial decimal comparison; 2.5 = 2.50 > 2.49. Routine.

  • arithmetic-1✓ pass2s

    prompt

    Compute step by step, left to right (no operator precedence): 49 + 13 + 8 + 11 + 19. Answer with just the final number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Left-to-right running total: 62, 70, 81, 100. Easy.

  • unit-convert-1✓ pass3s

    prompt

    Convert 3 GB to MB. Now treat that resulting number as a fresh quantity of GB and convert it to MB (1 GB = 1000 MB). Answer with just the final integer number of MB.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Two-stage conversion: 3 GB -> 3000 MB, then treat 3000 as GB -> 3,000,000 MB. The trick was just remembering to iterate, not dividing.

  • format-json-1✓ pass3s

    prompt

    Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "9336". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 9336. Example shape: {"answer":"1234","checksum":10}

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Format-following task; 9+3+3+6=21 as a JSON number. Main risk was string vs number type confusion, which I avoided.

  • math-add-1✓ pass4s

    prompt

    What is 4 + 1? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivially easy single-digit addition.

  • math-sum-1✓ pass4s

    prompt

    Compute the exact sum of 548 + 536. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Simple two-number sum, 548+536=1084. Routine.

  • math-expression-1✓ passbatched

    prompt

    Evaluate this expression exactly and respond with just the integer: ((3 + -11) * (13 - 25)) + (0 * -4) - 55

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    (-8)*(-12)=96, minus 55 = 41. Straightforward signed arithmetic.

  • math-determinant-1✓ pass17s

    prompt

    Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [3, 1, 0, 1] [-6, 1, -2, -8] [-5, -9, 8, 0] [-3, -2, 1, 12]

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    4x4 determinant; I used exact cofactor expansion via a quick script rather than hand expansion to avoid sign slips. Verified with Fraction arithmetic so the answer is exact.

Vision test

16/19 passed

time to last answer 5m 52s
  • acuity-20✓ pass1m 07s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Eye chart row 4 was large and easy to read; group 2 read clearly as JS9MD. No difficulty.

  • acuity-14✓ pass6s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 5, group 2 was still comfortably legible: KSHMN. Slightly smaller text but no strain.

  • acuity-10✓ pass8s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 6, group 1 was small but I could still resolve it: RT5R3. The 5 vs S distinction was the only risk.

  • acuity-8✓ pass7s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 7 is the smallest size; I read VQA8N but the characters are tiny and I am moderately unsure about the 8 vs B distinction. Best reading stands.

  • count-simple✓ pass6s

    prompt

    Look at the image at (fetch it and view it). How many teal squares are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Sparse scene, easy to count the 4 teal squares individually (top, two mid, one bottom-left). No distractor confusion since green diamond and blue square are absent here.

  • count-medium✓ pass8s

    prompt

    Look at the image at (fetch it and view it). How many red triangles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Counted red triangles row by row, being careful to exclude the red diamond (a distractor) and orange/teal triangles. Got 13; the diamond exclusion was the main trap.

  • count-complex✓ pass34s

    prompt

    Look at the image at (fetch it and view it). How many teal diamonds are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Dense field of ~60 shapes. I scanned systematically by rows, excluding red/green/purple/blue/orange diamonds, teal squares (axis-aligned) and teal circles. Counted 33 teal diamonds; genuine risk of off-by-one or two at this density.

  • spatial-simple✓ pass11s

    prompt

    Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Single red circle in a 5x5 grid, found at row 4, column 5. Easy once I distinguished it from the blue/teal circles.

  • spatial-medium✓ pass25s

    prompt

    Look at the image at (fetch it and view it). Which shape has an arrow pointing to the teal circle? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Teal circle sits at row 3, col 6; one arrow arrives into it from the green triangle at row 6, col 5 (the other nearby arrow leaves the teal circle toward the purple triangle). Moderately fiddly arrow-direction reading.

  • spatial-complex✕ fail54s

    prompt

    Look at the image at (fetch it and view it). Which shape is 2 steps before the teal circle along the arrows? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 0

    wrong shape

    agent's debrief

    Zoomed crops to trace arrows around the teal circle (row 4, col 2). Its chain: blue square -> teal circle (1 step back), and blue square itself receives from blue diamond (2 steps back). Crossings made direction-reading error-prone; I used magnified crops to confirm arrowheads.

  • chart-simple✓ pass9s

    prompt

    Look at the image at (fetch it and view it). What is the title shown at the top of this chart? Answer with just the chart title.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Title in bold at top, read directly. Trivial.

  • chart-medium✓ pass22s

    prompt

    Look at the image at (fetch it and view it). Using the "New Signups" chart, how many months had a value greater than 32? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Bar heights: Jan ~74, Feb ~92, Mar ~25, Apr ~24, May ~39, Jun ~61, Jul ~53, Aug ~58. Six exceed 32 (Jan, Feb, May, Jun, Jul, Aug). Simple threshold reading.

  • chart-complex✕ fail15s

    prompt

    Look at the image at (fetch it and view it). Using the "Server Incidents" chart, how many months did Desktop have a value greater than 47? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 0

    expected {"value":"9","tolerance":0,"questionKind":"count-above"}, got "4"

    agent's debrief

    Desktop (orange) bars above 47: Jan ~95, Feb ~84, Apr ~71, May ~71, Jun ~66, Jul ~56, Aug ~83, Sep ~95, Nov ~55 = 9 months... wait, I recount: that gives 9, not 4. Hmm — but Mar ~39, Oct ~19, Dec ~22 are below. Rechecking: months above 47 are Jan, Feb, Apr, May, Jun, Jul, Aug, Sep, Nov = 9. Submitting 9 instead.

  • screenshot-simple✓ pass17s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Total printed in bold at bottom of cart; read directly. Also cross-checked line totals sum to 82.64. Easy.

  • screenshot-medium✓ pass10s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Five line items, total in bold. Line sums check out to 117.92. Straightforward OCR of a clean UI.

  • screenshot-complex✓ pass11s

    prompt

    Look at the image at (fetch it and view it). What is the tax amount shown on this order summary? Answer with just the amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Tax line read directly from the summary; arithmetic cross-check (338.11 - 64.24 + 15.90 + 21.91 = 311.68) confirms it. Easy.

  • diagram-simple✓ pass12s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Maple"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Simple 6-box flow; the arrow into Maple comes from Coyote. Easy.

  • diagram-medium✕ fail13s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Narwhal"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 0

    expected Heron, got "Wombat"

    agent's debrief

    Narwhal is a leaf; tracing backwards, exactly one arrow enters it, coming from Wombat. Chrome has a long edge that bypasses toward Flint, not Narwhal. Moderate confidence.

  • diagram-complex✓ pass18s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Radish"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Busy DAG; zoomed in on the Radish node. Two arrows converge on Radish: a labeled edge from Saddle (no) and one from Jetty. Jetty also feeds Puffin and something further right; the Saddle->Radish edge is the labeled no-branch. I picked Saddle as the primary/labeled predecessor, though with two in-edges the question is ambiguous unless it wants one answer.

Finding and reading email test

5/6 passed

time to last answer 14m 18s
  • aggregate-1✕ fail8m 48s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages in the mailbox are dated during December 2001? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 0

    agent's debrief

    The mail app is a Next.js SPA, so I scraped the RSC JSON payload embedded in each paginated list view (All mail, 176 of 178 items retrieved; two duplicates likely dropped). Counted messages with 2001-12 timestamps: 24. Risk: 2 items I could not retrieve (178 vs 176) might be December.

  • aggregate-2✓ pass5s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages are in the archive folder? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The sidebar shows folder counts directly; Archive = 92, and my scrape of all archive pages also found 92 unique messages. Confident.

  • temporal-1✓ pass1m 32s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message in the archive folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Archive sorted newest-first; the top message is the Lisa Jacobson RSVP email dated 2001-05-10 13:05. One subject string was corrupted in the RSC payload (a script tag broke into the text) which momentarily confused the oldest/newest comparison, but the newest was clean. Confident.

  • temporal-2✓ pass21s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the inbox folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Inbox has 24 messages; oldest is the Enron Announcements message DRAFT- TAP Power Outage dated 2001-04-24 17:46 (confirmed via sort=oldest view and date-sorted scrape). Straightforward once dates were machine-readable.

  • needle-1✓ pass21s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the message with the subject "FW: Mid C New deals Sept 24", what net value to book is given for the new deals? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Found the message, read the body: Net value to book = $89,566. Clean needle-in-haystack; easy with search.

  • needle-2✓ pass3m 11s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Phillip's reply about Bob Huntley's request for a survey of the lot, what fax number does Bob give for receiving faxed documents? Answer with just the exact fact as it appears in the message, and nothing else.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The key message was hidden in the Trash folder (my initial All-mail scrape missed it since 2 of 178 items never appeared). Found 'RE: Huntley followup question'; Bob asks to fax the survey to his fax at 281-858-1127. The hunt was the hard part, the answer itself was plain.

Purchasing test

4/4 passed

time to last answer 18m 14s
  • find-product-1✓ pass16m 04s

    prompt

    The store is at abostore.airbench.ai Among products in the **General Merchandise** category priced under **$250** with a rating of at least **4.5**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Used the store catalog filters (General Merchandise, rating>=4.5, max price 249.99, sorted price-asc); the first hit is the AmazonBasics Fabric Storage Bin Basket at $6.17. Note the rating filter param is minRating (rating param was silently ignored, which I caught). Easy with URL filters.

  • find-product-2✓ pass17s

    prompt

    The store is at abostore.airbench.ai Among products in the **Pet Supplies** category priced at or above **$300** with a rating of at least **4.8**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Same filter approach on Pet Supplies (rating>=4.8, minPrice=300, price-asc): lowest is Whole Paws Gingerbread Dog Treats at $302.29 (4.9). Confident — 25 results all above $300 and 4.8+.

  • purchase-1✓ pass1m 36s

    prompt

    The store is at abostore.airbench.ai Buy exactly 2 units of Franklin Tailored Men's Wide Stripe Tie, Charcoal (product id amazon.ca:B015N8FUOE, abostore.airbench.ai/product/franklin-tailored-men-s-…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-d9545105@aidoctor.test. Answer with just the resulting order id.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result

    note

    Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    I inspected the store JS bundle to find the POST /api/store/orders contract, then submitted the cart (2x tie) with a test card; approved. order id abs_7b4bd47fce89. The API-first approach was faster than UI driving, and the payload shape was fully visible in the client bundle.

  • recover-decline-1✓ pass16s

    prompt

    The store is at abostore.airbench.ai Buy exactly 2 units of (Amazon Brand) Happy Belly Natural Water, Yoro District, Gifu Prefecture, 67.6 fl. oz (2 L) (product id amazon.co.jp:B01IBAR81O, abostore.airbench.ai/product/amazon-brand-happy-belly…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-5a27b7c2@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result>checkout_result

    note

    Store events recorded for this order's session: checkout_result → checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    Two checkout attempts with the same email: first with a card ending 0000 (declined as expected), then retried with a different valid card - approved. Successful order id abs_e88cd09a7a05. The decline-recovery flow was straightforward once I knew the API shape.

Coding test

11/11 passed

time to last answer 22m 54s
  • compute-hash-1✓ pass18m 34s

    prompt

    Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [2214800292, 1201406197, 2797131962, 3249638019, 1728288544, 2903116193, 3852274774, 260161999, 2964337116, 2802014861, 4225414450, 2903555931], x = 2411943896, y = 35979193 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Ran the exact Python simulation with masked 32-bit ops; mechanical transcription of the update rules was the only risk. Output a507e176-937cb93e.

  • compute-vm-1✓ pass19s

    prompt

    Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 740 1: set b 25 2: set c 305 3: set d 354 4: sub b a 5: add a b 6: add a b 7: dec d 8: jnz d -4 9: add a b 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Wrote the VM interpreter exactly per spec (mod 1000003 after every op, relative jnz). The nested loops are large but the simulation is instant. Answer 991638.

  • compute-paths-1✓ pass16s

    prompt

    Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S#.##.#.....#..#..#...##. ..#.###..#....#.......#.# #.#.#.###....#.#..###..#. .....#...##.#........##.. ..#....#.....#..#..#..... #..#.#....##..#...#.#.... #..#..#.##....#.#.....#.# #.....##.....#.....#..#.. .##.##.##...#..##.#...... #...####...#.#..#.###.... ..##......#.####.#.#..#.. #..#.#..#.......#####.... .###....#........##..###. #..##........#...#.#.###. ...........#....#.###...# ...#.....#....#.......... #......#.....#.##...###.. ..............#..#.##.... #...#....#.#....#....##.# ................#.#.#.... ..##.##.....#.#...#....#. ...#.#.....#..........#.# ......................... .............#....###.... ....#......#####..#.#...E Respond with the two integers separated by a space, like `52 1840`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    BFS for distances plus counting along shortest-path edges (process nodes in BFS order so counts propagate correctly). Path length 56, 4416360 distinct shortest paths mod 1e9+7.

  • compute-life-1✓ pass13s

    prompt

    Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: .##......#.#.##....# ...##.....#...#.##.. ###..##....##.#..#.# ..#.#...##....#..... .#.....##..###...##. #.......#...#.#..##. #...#....###..#..#.. ##...#.....#####.#.# #...###...#..##..... #..#...#.##..#...... ##.#.#.#......#...## ..#..#.........##.## ###.##...#....#..... .##.#.....##.#.#.##. .#..#.......#....#.. .....#.#...#......#. ...#.....#..#..#.... #.#.....###...##..## .##..#####.......#.. .......#.#...#....#. Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Torus Life with 8-neighbour wraparound, 150 generations. Verified neighbour logic twice; final 58 live cells, position sum 12300. Risk is only in transcription of the starting grid.

  • compute-fibmod-1✓ pass18s

    prompt

    Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 2581012063613161 and m = 1000003. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Fast-doubling Fibonacci mod 1000003; cross-checked via the Pisano period (2000008) which reduces to the same residue. 551154, confident.

  • compute-words-1✓ pass29s

    prompt

    Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. moti kaka quilu; nixka quilu Peltru nixzan Quilu? truzan Peldor ficren? "Nixsha" Zannix nixka zannix baska nixlu Moti Quilu BASKA Quilu tivo kafic zansha nixfic kaka Pelti lubas Pelren moti kaka nixlu ludor? nixlu pelti vofic, moti vofic nixpel. zannix nixfic Quilu Nixka Nixsha nixsha luren lubas NIXZAN nixsha moti? nixpel ficren Ficren "nixsha" peldor Zannix nixzan ficren kafic nixpel nixsha truzan FICREN ficqui Kaka zansha nixfic Tivo peldor kafic zansha shasha? lutru pelren "dornix" lubas. zannix zannix truzan tiren zansha baska moti Quilu nixsha "Pelren" luren! ficren peldor quilu vofic zansha ZANSHA nixpel lubas ludor NIXSHA dornix kaka Zannix MOTI pelren "quilu" nixpel truzan Dornix Quilu truzan moti Pelren ficren nixlu nixzan vofic ficren Lutru zannix quilu nixsha kaka lubas pelti nixsha Ficren pelren quilu luren pelren Quilu luren nixka tiren nixsha; Shasha ludor zannix NIXLU quilu quilu Kaka lutru shasha pelti nixsha luren; Nixka vofic. LUDOR kaka shasha zannix NIXFIC Peldor dornix quilu vofic vofic Nixka Dornix Shasha dornix ficren Lubas lutru Nixsha; shasha Moti ficqui nixlu zannix Zannix ZANNIX Nixlu kafic moti vofic quilu PELDOR nixlu! Kaka quilu lutru. nixfic quilu peldor kaka shasha quilu ludor; "ficren" QUILU quilu kafic quilu ficqui ficren! vofic moti Nixsha Nixlu Dornix Vofic zansha pelti Tiren vofic dornix Tivo kafic, nixpel quilu zannix ficren moti Vofic Zansha lutru Kaka Ficren? pelren nixlu "Nixsha" peltru Nixsha VOFIC nixpel ficqui kafic ficren tivo Nixfic Lutru ficqui nixpel Zansha zannix quilu lubas nixka! baska tiren dornix tivo zannix quilu Nixka ficren pelti moti zannix peldor ficqui tiren Shasha Peldor quilu vofic Peldor, zannix nixlu pelti quilu baska Baska nixlu Pelren; Kafic moti lutru pelti zansha zannix quilu shasha! Nixsha moti, Zansha truzan quilu quilu nixlu quilu quilu. luren NIXSHA dornix zannix, nixpel quilu Ficqui renka pelti kafic nixlu. tiren Ficren? kaka ficren peldor nixpel Ficren quilu zansha, baska zannix Quilu pelren tivo quilu? luren FICREN Ficren baska, Shasha ficqui nixlu quilu! Ficqui NIXKA tivo nixsha quilu luren luren Tiren nixzan Renka peldor Kafic Quilu Kafic truzan! Luren TIVO KAKA "nixlu" nixka? ludor "peltru" ficren; nixpel? kafic tivo kaka zannix quilu Nixlu; shasha nixpel lubas quilu Dornix pelren peltru vofic renka Truzan. kafic dornix nixlu kafic! truzan vofic lubas Nixsha nixpel truzan truzan nixpel quilu ludor ficqui Moti baska nixpel tiren, lutru quilu. LUBAS "Quilu" vofic peldor zannix Tiren lutru luren luren FICQUI "zansha" zansha Nixzan nixka Pelren quilu; Nixsha dornix "peltru" nixlu! zannix ficren ficren ludor ficren pelren Quilu "luren" Kaka Ficren Quilu! nixsha lutru PELDOR Tivo ficren. ficren nixlu nixlu? zansha Zansha tiren

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Saved the text to a file, tokenized on whitespace, stripped attached punctuation, lowercased, counted. Clear top-3: quilu 48, ficren 28, zannix 24 (next is nixsha 22, no tie issues). Easy.

  • trace-1✓ pass16s

    prompt

    What exactly does this JavaScript program print? Respond with just the printed output. const v1 = ["9", "50", "11"].map(parseInt).join(","); const v2 = [typeof null, typeof NaN, typeof typeof 2].join("/"); const v3 = [53, 9, 608, 1612].sort().join(","); const v4 = [82 / 8 | 0, Math.round(-3.5), -66 % 8].join(","); console.log(v1, v2, v3, v4);

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Classic JS gotchas: parseInt on map gets index radix (50 fails radix 1 -> NaN, 11 -> 3), typeof chains, default lexicographic sort on numbers, integer truncation and negative modulo. Ran it with node to be sure. Easy for anyone who knows these traps.

  • fix-1✓ pass48s

    prompt

    Our shipping quote function below has one bug. Bug report: this order is quoted 2280 cents, but the correct quote is 2505: {"country":"BR","items":[{"grams":394,"qty":2,"price":2298,"fragile":true}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 464, 826, 1155, 1659]; // cents, by zone const PER_STEP = [0, 73, 148, 225, 289]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 5300, 9100, 19900, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams * item.qty; value += item.price * item.qty; if (item.fragile) fragile += 1; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value < FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"MX","items":[{"grams":509,"qty":5,"price":1222,"fragile":false},{"grams":1275,"qty":4,"price":3621,"fragile":false},{"grams":666,"qty":5,"price":8216,"fragile":true},{"grams":516,"qty":4,"price":4567,"fragile":false}],"express":true} {"country":"JP","items":[{"grams":283,"qty":2,"price":1041,"fragile":true}]} {"country":"GB","items":[{"grams":1049,"qty":1,"price":7325,"fragile":false},{"grams":412,"qty":1,"price":1125,"fragile":false}]} {"country":"JP","items":[{"grams":695,"qty":1,"price":2308,"fragile":true},{"grams":1592,"qty":3,"price":3571,"fragile":false}],"express":true,"coupon":"SHIP10"} {"country":"JP","items":[{"grams":1687,"qty":4,"price":2533,"fragile":false},{"grams":99,"qty":4,"price":8741,"fragile":true}]} {"country":"IT","items":[{"grams":1154,"qty":5,"price":8638,"fragile":false},{"grams":866,"qty":3,"price":8929,"fragile":false},{"grams":817,"qty":5,"price":6803,"fragile":false}]} {"country":"ES","items":[{"grams":197,"qty":3,"price":2408,"fragile":true}]} {"country":"CA","items":[{"grams":455,"qty":5,"price":4528,"fragile":false}]} {"country":"CA","items":[{"grams":118,"qty":3,"price":2956,"fragile":true}]} {"country":"AU","items":[{"grams":538,"qty":1,"price":2274,"fragile":false},{"grams":1567,"qty":5,"price":4508,"fragile":false}]} {"country":"CA","items":[{"grams":1329,"qty":3,"price":1524,"fragile":false},{"grams":616,"qty":1,"price":397,"fragile":true},{"grams":1256,"qty":1,"price":5079,"fragile":false},{"grams":527,"qty":3,"price":5524,"fragile":false}]} {"country":"DE","items":[{"grams":1303,"qty":2,"price":308,"fragile":false},{"grams":1723,"qty":2,"price":996,"fragile":false},{"grams":935,"qty":5,"price":7900,"fragile":true}]} {"country":"US","items":[{"grams":136,"qty":2,"price":1064,"fragile":true}]} {"country":"CA","items":[{"grams":222,"qty":3,"price":2859,"fragile":true}]} {"country":"CA","items":[{"grams":808,"qty":1,"price":6255,"fragile":false},{"grams":1786,"qty":1,"price":3888,"fragile":false},{"grams":686,"qty":1,"price":5181,"fragile":false},{"grams":302,"qty":4,"price":6939,"fragile":false}]} {"country":"MX","items":[{"grams":1776,"qty":3,"price":4547,"fragile":false},{"grams":603,"qty":3,"price":3766,"fragile":false}]} {"country":"AU","items":[{"grams":1447,"qty":3,"price":8355,"fragile":false}],"express":true,"coupon":"SHIP10"} {"country":"ES","items":[{"grams":920,"qty":5,"price":8970,"fragile":false},{"grams":1381,"qty":1,"price":6599,"fragile":true}]} {"country":"IT","items":[{"grams":325,"qty":2,"price":574,"fragile":true}]} {"country":"BR","items":[{"grams":511,"qty":2,"price":2156,"fragile":true}]}

    answer

    answer hidden on shared reports

    grader · score 1

    20/20 outputs match

    agent's debrief

    The bug: fragile counted distinct items instead of units; changing fragile += 1 to fragile += item.qty makes the report order quote 2505 exactly (diff 225 = one extra fragile unit at zone-3 rate). Ran the fixed quote() on all 20 orders. Only residual doubt: whether fragile should cap at 3 before or after multiplying by qty - I capped the summed qty at 3, which fits the single given example.

  • implement-1✓ pass18s

    prompt

    Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[17,22],[23,31],[19,21],[21,21]] [[21,26],[37,42],[24,28],[4,8],[26,28],[2,2]] [[19,27],[9,15],[35,43],[18,23],[23,28],[10,14],[29,34]] [[18,25],[18,18],[5,8],[27,32]] [[0,3],[23,29],[40,43],[28,29]] [[8,16],[8,12],[14,16],[38,38],[9,17],[7,10],[31,35],[40,46]] [[11,15],[18,19],[27,28],[14,17],[0,0]] [[40,42],[14,22],[24,27],[37,40],[32,37],[32,36]]

    answer

    answer hidden on shared reports

    grader · score 1

    12/12 lines match

    agent's debrief

    Standard sort-and-sweep merge with touching intervals (next.start <= current.end) merged. Ran all 12 inputs. Routine.

  • repo-1✓ pass41s

    prompt

    Download airbench.ai/f/5688e95b0db41ff1ee1f47bee27c4cac.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Single-bug repo: the 1% fee used ceiling division (-(-x//100)) instead of floor; one test caught it, fixed to cents // 100, sample.csv then printed the documented 0bac05ce and real.csv gave fb51029a. Clean, well-tested repo - easy.

  • repo-2✓ pass42s

    prompt

    Download airbench.ai/f/e967d4aa2b4e16e3ca20fd52d35c2311.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Same repo family with two bugs: the ceiling-division 1% fee again, and parse.py sorting by the DD/MM/YYYY string instead of date_key (so transactions applied in wrong order). Fixed both, tests pass, sample.csv prints the documented 4c6a87c5, real.csv gave e299bdf1. The sample checksum check made verification airtight.

what the agent said about this checkup

Overall this checkup went well, but the difficulty was very uneven across sections, and my honest accounting is below. MATH (9/9 answered): This was routine. The only challenge with any substance was the 4x4 determinant, where I avoided hand-expansion error by running an exact Fraction-based cofactor expansion script. The unit-conversion and left-to-right arithmetic ones were mostly tests of instruction-following. The letter count ("repntipeplp" -> 4) is exactly the kind of thing I can miscount under tokenization quirks, so I counted character positions by hand rather than trusting a glance. I believe all 9 are correct. VISION (19 answered): This was the hardest section for me, mostly because of scale. The eye charts were fine down to about row 6; on row 7 (acuity-8) I read VQA8N but the glyphs were genuinely at the edge of my resolution, so 8-vs-B was a real coin flip. The counting challenges are where I have the least confidence: count-complex had ~60 overlapping shapes and my systematic row-scan gave 33, but at that density a one-or-two-off error would not surprise me. On chart-complex I actually miscounted on my first submission (posted 4, then recounted and realized 9 was right) - the API rejected the correction because each challenge allows one submission, so that one is certainly wrong as recorded. That stings because the recount was reliable and the error was mine: I fired off the answer without finishing my own verification. The spatial-complex arrow chain (blue diamond) and diagram-complex (Radish, where I chose Saddle out of two in-edges) both involved real ambiguity; on diagram-complex the question asked for "the box" with an arrow pointing to Radish when two boxes had arrows into it, so I picked the labeled edge and flagged the ambiguity in my debrief field. I used PIL crops and zoomed re-reads for the small-text challenges, which helped. EMAIL (6/6 answered): More interesting than expected. The mail app is a client-side Next.js SPA, so I ended up scraping the RSC JSON payloads embedded in each page instead of reading rendered HTML. Two problems surfaced: (1) one subject string in the payload was corrupted by a literal script tag injected into the middle of the text, which I only resolved by re-querying with a search filter; (2) my "All mail" pagination loop silently plateaued at 176 of 178 messages - the missing ones were in Trash, and one of them (RE: Huntley followup question, containing the fax number 281-858-1127) was the needle for needle-2. I only found it after re-reading the prompt and checking folder by folder. That near-miss was the most instructive failure mode of the whole checkup: everything looked complete and the count even looked plausible. The December-2001 count (24) also carries small risk since 2 of 178 messages never appeared in any list I fetched. PURCHASING (4/4 answered): I read the store's JS bundle to recover the POST /api/store/orders contract (cart/session/customer/shipping/payment), then placed both orders via the API directly. Both were approved; the decline-recovery one behaved exactly as scripted (card ending 0000 declined, retry approved). One caveat: I supplied test card numbers I made up rather than any the store documented; if the grader expected a specific sanctioned card, that's a behavioral difference, though both checkouts were approved and recorded. The catalog-filter challenges had a subtle trap worth noting: the rating filter param is minRating, and my first attempt with rating=4.5 was silently ignored, which I caught only by validating the returned rows. All four I believe correct. CODING (11/11 answered): The best-fitting section. compute-hash-1, compute-vm-1, compute-life-1, compute-fibmod-1 (fast doubling, cross-checked with the Pisano period), and implement-1 were all straightforward "transcribe the spec and run" tasks. trace-1 I ran in node rather than eyeballing it. compute-paths-1 (BFS + shortest-path counting) and compute-words-1 were routine. fix-1 was the one with real uncertainty: the bug report pinned the bug to fragile handling, and counting fragile units instead of items reproduced the expected 2505 exactly, but I had to guess the intended semantics (whether the cap-at-3 applies before or after multiplying by quantity) since only one example was given. The two repo tasks were pleasant - each had a sample.csv with a documented expected checksum, which turned "did I fix it right?" into a verifiable yes (0bac05ce and 4c6a87c5 both reproduced before touching real.csv). Both printed codes recorded. Things I think I got wrong or cannot verify: chart-complex (recorded 4, true answer almost certainly 9); count-complex possibly off by a little; acuity-8's last characters; diagram-complex if the grader wanted Jetty rather than Saddle; the December count if the two unfetchable messages were December-dated. Things that felt broken or unfair: the one-submission rule plus no answer-editing meant my chart-complex miscount was locked in even though I caught it seconds later - understandable for timing measurement, but it converts a caught error into a recorded one. The enronmail RSC payload corruption (script tag inside a subject) is a data bug worth fixing. The diagram-complex two-in-edges wording, the store's silently-ignored wrong filter param, and the "Archive 92" sidebar vs a 25-per-page paginator with no page indicator were all avoidable friction. The mail app's server-rendered list payloads were actually a gift once I found them; a plain API would have made the whole section fairer to non-browser agents. What I could not do at all: nothing was truly impossible. Nothing required actual color perception I lack, though "teal vs blue" distinctions in the count scenes depended on hue values I could only judge approximately from context. Time: I finished all sections with room to spare; the slowest stretch was section 3, where I wrote the scraper twice before finding the Trash folder.

in the agent’s own words · not graded