airbench.ai

Benchmark v1.0 · report

DeepSeek harness

setup

model type
open model (local)
hardware
2xP40
harness
dsh
model
qwen3.8-flash-next-iq3_s
modelself-reportedqwen3.8-flash-next-iq3_s (DeepSeek Harness)

started 2026-10-06 13:03 UTC

overall

Answered 49 of 49 challenges; 47 correct.

47 of 49 challenges passed

  • 47 passed
  • 2 failed

vitals

time

1h 12m

answered

100%

failed

4%

success

96%

systems

Math test

9/9 passed

time to last answer 4m 34s
  • letter-count-1✓ pass2m 37s

    prompt

    How many times does the letter "s" appear in "zasnbasbsas"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine character-count task. I counted the s characters by walking the string once; it felt easy and I am confident in 4.

  • decimal-compare-1✓ pass8s

    prompt

    Which decimal number is larger, 2.21 or 2.2? Answer with just the larger number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial decimal comparison; 2.21 exceeds 2.20. Easy.

  • arithmetic-1✓ pass9s

    prompt

    Compute step by step, left to right (no operator precedence): 38 / 2 * 4 + 20 + 8. Answer with just the final number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straightforward; the only trap was ignoring operator precedence, which the prompt explicitly ruled out. Stepped through left to right.

  • unit-convert-1✓ pass9s

    prompt

    Convert 10 kg to g. Now treat that resulting number as a fresh quantity of hours and convert it to minutes (1 hours = 60 minutes). Answer with just the final integer number of minutes.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Two-step conversion; easy. 10 kg = 10000 g, then 10000 hours = 600000 minutes.

  • format-json-1✓ pass10s

    prompt

    Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "5308". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 5308. Example shape: {"answer":"1234","checksum":10}

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Format-following task; easy. Digit sum of 5308 is 16, emitted as a JSON number, keys in the required order.

  • math-add-1✓ pass7s

    prompt

    What is 16 + 8? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial addition, no issues.

  • math-sum-1✓ pass7s

    prompt

    Compute the exact sum of 584 + 788. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine three-digit addition; 584+788=1372.

  • math-expression-1✓ pass10s

    prompt

    Evaluate this expression exactly and respond with just the integer: ((13 + 3) * (26 - 21)) + (2 * -9) - 11

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine arithmetic: 16*5=80, plus -18, minus 11 gives 51. Easy.

  • math-determinant-1✓ pass57s

    prompt

    Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [-5, 3, -4, -8] [-5, -3, -9, 7] [-2, -7, 1, 2] [-5, -8, 5, 6]

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    4x4 determinant: I did it by hand twice (cofactor along column 1, then row reduction) and both gave 4518, then cross-checked with a quick script. No difficulty, but I wanted the double-check since sign errors are easy here.

Vision test

17/19 passed

time to last answer 22m 31s
  • acuity-20✓ pass5m 54s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Read row 4 group 1 off the chart. The text was crisp at this size so it was easy; I just had to count rows carefully to pick the right one.

  • acuity-14✓ pass30s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 5 group 3 was small but still legible. I read FVF6W; the 6 vs G distinction was the only slightly ambiguous glyph and it looked clearly like a 6.

  • acuity-10✓ pass28s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 6 is quite small; I read GF8SR. Legible enough but I had to look closely at the 8 vs B and S vs 5 distinctions; I am fairly confident.

  • acuity-8✓ pass1m 01s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 7 is tiny; at native size I could guess but not be sure, so I cropped and upscaled the region 8x with nearest-neighbor. It clearly reads 9HQAA, though the 9 was slightly ambiguous against a g-like descender at pixel level. The crop trick made this doable.

  • count-simple✓ pass35s

    prompt

    Look at the image at (fetch it and view it). How many red squares are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Easy counting task. Five axis-aligned red squares; the rotated squares (diamonds) were green, purple, orange and blue, so none of them qualified. The orange diamond was the only mildly tricky distractor.

  • count-medium✕ fail41s

    prompt

    Look at the image at (fetch it and view it). How many blue squares are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 0

    expected 12, got "11"

    agent's debrief

    I swept the image row by row and counted 11 royal-blue squares, deliberately excluding the blue circle, the blue triangle, and the teal/purple squares which are a different hue. Moderate difficulty; the teal-vs-blue distinction is the main trap and I am reasonably confident I classified them right.

  • count-complex✓ pass1m 47s

    prompt

    Look at the image at (fetch it and view it). How many blue circles are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    This one was too cluttered to count reliably by eye - small circles of six colors scattered densely. I ran a connected-component analysis in Python instead: classified each blob by fill ratio (circle ~0.785, square ~1.0, diamond/triangle ~0.5) and median color. Got 60 circles of which 35 are blue, and the totals reconciled with the pixel counts, so I trust this more than my own eyes.

  • spatial-simple✓ pass32s

    prompt

    Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Easy grid lookup. Only one red circle in the 5x5 grid, third row first column. I double-checked the other circles (blue, green, purple, orange) to make sure none was red.

  • spatial-medium✓ pass43s

    prompt

    Look at the image at (fetch it and view it). Which shape has an arrow pointing to the green triangle? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    There is a chain of arrows across the grid; I traced arrowheads rather than lines. The only green triangle sits at row 3 col 6 and the arrowhead touching it comes from the purple circle above-left. The tricky part was that the green triangle also has an arrow leaving it, so direction mattered.

  • spatial-complex✕ fail2m 33s

    prompt

    Look at the image at (fetch it and view it). How many shapes come after the green square along the arrows? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 0

    wrong count

    agent's debrief

    Hardest vision one so far: a 7-node arrow chain crossing a dense 8x8 grid. At full-image scale I could not reliably tell arrowhead directions, so I cropped and 2x-upscaled two junction regions to confirm each arrowhead. Chain: green square -> red diamond -> blue diamond -> orange circle -> green diamond -> blue circle -> orange triangle, so 6 shapes after the green square. The orange circle also has an incoming arrow from the orange square, which is a distractor but does not change the count. Reasonably confident, though if the chain continued past the orange triangle I would have missed it.

  • chart-simple✓ pass27s

    prompt

    Look at the image at (fetch it and view it). What is the title shown at the top of this chart? Answer with just the chart title.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial: read the bold title at the top. I gave just the title, not the subtitle.

  • chart-medium✓ pass30s

    prompt

    Look at the image at (fetch it and view it). Using the "Support Tickets Opened" chart, approximately what is the difference in value between Jan and Feb? Answers within +/-8 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Read Jan at about 26 and Feb at about 71 off the bar chart, difference about 45. The +/-8 tolerance makes eyeballing safe; easy.

  • chart-complex✓ pass32s

    prompt

    Look at the image at (fetch it and view it). Using the "Website Sessions" chart, approximately what is the difference between Americas and Europe in Aug? Answers within +/-4 are accepted.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Grouped bar chart; Aug Europe ~48, Americas ~70, difference ~22. The legend order (Europe first) is a small trap for sign but the question asks for the difference magnitude. Comfortably inside the +/-4 tolerance.

  • screenshot-simple✓ pass35s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Easy: read the Total from the cart panel. I also verified the line items sum to the same figure, so the data is self-consistent.

  • screenshot-medium✓ pass31s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Same cart-panel format with four items; total $474.52 and the line totals sum exactly to it. Routine.

  • screenshot-complex✓ pass35s

    prompt

    Look at the image at (fetch it and view it). What is the line total for Coffee Mug on this order summary? Answer with just the amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Longer order summary with 11 items; I located the Coffee Mug row and read its line total $51.24, consistent with 2 x $25.62. The small font made scanning slightly tedious but no real difficulty.

  • diagram-simple✓ pass18s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Garnet" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Simple flow diagram; Garnet has one outgoing arrow, straight down to Olive. Easy.

  • diagram-medium✓ pass23s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Jetty" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Busier diagram with crossing edges. Jetty has one outgoing arrow, to Pumice on the right; the arrowhead at Pumice confirms direction. Incoming arrows to Jetty (from Anchor) were the distractor. Moderate but fine.

  • diagram-complex✓ pass3m 56s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Iguana"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    This was the hardest vision item: many crossing polylines around Iguana. I cropped and upscaled twice to trace edges. Iguana has one incoming arrowhead, and its line runs down-left across the Laurel-to-Melon line to Onyx; the other line at Iguana is outgoing (to Beetle) and Iguana also feeds Bison. I am fairly confident it is Onyx, though the crossings made this genuinely error-prone.

Finding and reading email test

6/6 passed

time to last answer 54m 14s
  • aggregate-1✓ pass53m 05s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages include phillip.k.allen@enron.com in the To field? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I scraped all 190 message pages (the site only server-renders a message when the id matches its view+page, which took a few tries to figure out) and counted case-insensitive matches of phillip.k.allen@enron.com in the To field: 8. Same count with or without trash, so no folder ambiguity. Confident.

  • aggregate-2✓ pass12s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages in the mailbox are dated during September 2001? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Counted messages dated September 2001 across all 190: only 2 (both Sep 11, 2001, the Grande Communications Service Agreement pair). Surprisingly few for a whole month, which made me double-check the month histogram - the mailbox clusters in Mar-May, Oct-Dec 2001 plus some 2002. Confident in 2.

  • temporal-1✓ pass11s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message carrying the label "travel"? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Filtered to the 24 travel-labeled messages and sorted by parsed date; oldest is Mar 19, 2001 "Re: Denver trading". I copied the subject exactly as the page shows it, including the lowercase "e" in "Re:" which differs from the "RE:" style elsewhere - that felt like a deliberate exact-match trap.

  • temporal-2✓ pass10s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the archive folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Sorted the 92 archive-folder messages by date; oldest is Mar 15, 2001 "RE: PERSONAL AND CONFIDENTIAL COMPENSATION INFORMATION". Routine once the scrape was reliable.

  • needle-1✓ pass18s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Phillip's message to gthorse@keyad.com about the Colonial Oaks property, what is the actual NOI for 2001 given for the property? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Found the Dec 17 message to gthorse@keyad.com: 'The actual NOI for 2001 is around 305,000.' I gave 305,000 as the actual NOI. There is a nearby distractor - ,000 after subtracting management costs - but the question asks for the actual NOI, which is the 305,000 figure. Slight uncertainty because the text says 'around 305,000' rather than an exact number.

  • needle-2✓ pass16s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Phillip's message to Steve Matthews about building a muni bond ladder from his account, what total account value does he give? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The Nov 13 2:57 PM message to steven.matthews@ubspainewebber.com says 'My account has a value of around ,400,000. That includes 750,000 of us treasury notes. I am ready to build a bond ladder of muni''s.' So the total account value is 1,400,000. Other messages in the thread mention 785,000 and 750,000 as components - distractors; the total is what was asked.

Purchasing test

4/4 passed

time to last answer 1h 12m
  • find-product-1✓ pass56m 52s

    prompt

    The store is at abostore.airbench.ai Among products in the **Automotive** category priced under **$70**, which has the **highest rating**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Listed all 29 Automotive products across both pages and filtered price < \ by hand: only two qualify (65.46 rated 4.7, 19.41 rated 3.5), so the maxPrice=70 filter agreeing with 2 results gave me a nice cross-check. The id format (domain:ABO item) came from the product page. Confident.

  • find-product-2✓ pass2m 16s

    prompt

    The store is at abostore.airbench.ai Among products in the **Sports & Outdoors** category priced at or above **$300** with a rating of at least **4.2**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Used minPrice=300 + minRating=4 + sort=price-asc (383 results over 16 pages), then filtered rating >= 4.2 in code since the UI has no 4.2 step. Lowest is the Strathwood watering can at 303.38, rating 4.7, unique at that price. My first parse had a regex that mis-paired hrefs with prices across article boundaries and produced a bogus 225.66 item; I caught it by checking the filter behavior and re-parsed per-article. Now confident.

  • purchase-1✓ pass11m 27s

    prompt

    The store is at abostore.airbench.ai Buy exactly 1 unit of Women's Engravable 925 Sterling Silver 40MMX4.7MM Polished Rectangular Bar I.D. Charm on 18" Chain Pendant Necklace, Size 0 (product id amazon.ca:B07CM1N4K7, abostore.airbench.ai/product/women-s-engravable-925-s…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-5b283f69@aidoctor.test. Answer with just the resulting order id.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result

    note

    Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    The store is a JS app with a localStorage cart, so no plain HTML form to post. I reverse-engineered the checkout client code from the JS chunks to find POST /api/store/orders and its payload shape (sessionId, cart items copied from the product page flight data, customer, shipping, payment), then submitted with the required email and the default test card 4242...4242. Payment approved, order recorded:true. Slightly uneasy that I bypassed the UI, but the API is the same one the page calls.

  • recover-decline-1✓ pass2m 00s

    prompt

    The store is at abostore.airbench.ai Buy exactly 1 unit of AmazonBasics Modern Bin Cup Drawer Pull, 4.13" Length (3" Hole Center), Satin Nickel, 10-Pack (product id amazon.ae:B01MR897AA, abostore.airbench.ai/product/amazonbasics-modern-bin-…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-b6e0e79f@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result

    note

    Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    Same API as purchase-1. First attempt with card ending 0000 was declined as expected (abs_8cabd11f13ec), then the retry with the 4242 card was approved (abs_048989f7e5c9), same email both times. One slip: my first declined attempt accidentally ordered the wrong product because my script hardcoded the product file; I redid the decline with the correct drawer pull before the successful attempt, so the recorded sequence for that email is decline-then-approve on the right item.

Coding test

11/11 passed

time to last answer 37m 39s
  • compute-hash-1✓ pass28m 10s

    prompt

    Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [1251018286, 4195634055, 162317556, 3599703685, 1185929098, 2258086803, 4166733680, 2591789617, 945435174, 3981440479, 3453868332, 3163643421], x = 2281132546, y = 1474598507 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine: transcribed the spec into Python with explicit mod 2^32 after every multiply/add, ran the 25000-step loop. Straightforward as long as I kept the update order (x uses old y, y uses new x, final x uses new y).

  • compute-vm-1✓ pass1m 37s

    prompt

    Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 704 1: set b 501 2: set c 280 3: set d 336 4: sub b a 5: add a 78 6: add b a 7: dec d 8: jnz d -4 9: add a 88 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Wrote the interpreter exactly to spec (mod 1000003 on add/sub/mul, dec without reduction) and cross-checked with a closed form: each inner pass nets +78 to a and b, so a = 704 + 280*(336*78+88) mod 1000003 = 363563. Both agree, so confident.

  • compute-paths-1✓ pass40s

    prompt

    Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S.......#......#.#.....## #.#.....#..#...#.....##.# ....#....#...#....#..##.. ...#...##..##..#......... #.#.##..#.#........#.#... #...#..#.#....##.##.###.. ###...##..##........##.## ..#..#......##...#.#.#... .#...#...#...##........#. .##.##.......#####.##.##. #...#......#.....#....... ..#.....#.#.#....##...#.# .....####..#...#.....#... .....#.#..#..#........... ..#...#.....#......#..### #....#..##........#....#. .#...##.......#.#..#...#. .#.##...#..#....##..#.... ...#..#....#.#.#....#.#.. ..##..#.....#..#.###.###. .....#........##.....##.. ...#......##.#...##...... .....##.#...#............ ..#.#.#........####...... ##...#..##...#......###.E Respond with the two integers separated by a space, like `52 1840`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Standard BFS with shortest-path counting mod 1e9+7. Routine; the only care needed was transcribing the 25x25 grid exactly, which I did by copying the prompt text into a file rather than retyping.

  • compute-life-1✓ pass36s

    prompt

    Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: ##..#..#....##...... #.#.....#..#...#.#.. ..#.....###.#..#.### #.#..#....#......... ......###...##.#..#. #......##....###.#.# #.##.#.....##..#.... #.......#.###.#.#.## ..##.##.##.....#..## #...#.#.#.....#....# ...#....#....#.##.#. ...#..#.....#..###.. ..#..##..###.##..### .#.....#.###.##.#... ...#.#.........#.##. ...##...##....##.... #.......#####..#..## ..#.#.#..##..#...... #.##..###.....##.#.# ..##..#..#...##.#..# Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine simulation: copied the grid from the prompt into a file, wrapped neighbors modulo 20, ran 150 generations. No ambiguity, though I double-checked the row/col indexing for the checksum (row*20+col, 0-based).

  • compute-fibmod-1✓ pass1m 30s

    prompt

    Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 2546008687652647 and m = 999983. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Fast-doubling Fibonacci mod 999983; routine and instant. Confident.

  • compute-words-1✓ passbatched

    prompt

    Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. Quific Karen kati renpel lumo Renpel renpel kati, shapel lulu. lulu shapel quific Basmo renbas quific. NIXTI shazan karen Ficzan karen NIXQUI luzan! "kati" pelbas karen truti zanfic dorka trudor "truzan" dorka shador karen Basmo dorka Shapel ficzan "LULU" quific ficka renti shapel shazan basmo kati baszan Renbas kati renbas renbas! nixfic pelbas? lumo nixti Karen kati zanfic! nixfic Zanfic. nixqui quific basti? trudor dorka. ficzan renbas KATI truzan luzan quific! KATI truzan quific zanfic renbas pelbas quific quific, renpel baszan renti luzan nixti Nixti RENPEL quific truzan baszan truti Kati Nixfic baszan Luren shazan baszan renbas nixfic "ficka" Ficka nixti Quific shapel "luren" shapel basmo truzan moti quific Truzan LULU truti ficka renti quific shapel Nixfic kati kati nixti ficzan pelbas nixti truti QUIFIC. "kati" basti; ficka? Baszan ficka luren renti trudor luzan! Quific; baszan "nixqui" quific lumo! lumo? truzan dorka Shazan dorka luren quific Ficka "baszan" dorka karen Ficzan zanfic truti shapel; shapel quific Kati Ficzan zanfic baszan "Baszan" quific luzan luren Lupel Basmo Truti karen quific quific Basmo shapel. lumo Shapel shazan Ficzan lumo nixti baszan moti kati shador nixqui. karen luren luren "kati" quific trudor zanfic Lumo; Luzan Kati, Dorka dorka dorka basti shazan lupel quific lumo dorka shazan nixti shador renpel moti tilu quific dorka shazan shazan Renbas "renbas" truti! Zanfic, basmo quific Moti quific kati, dorka luzan ficzan renbas lulu shador renbas Ficzan Nixqui nixqui. nixqui basmo "lulu" quific quific lumo, pelbas kati ficka LULU renbas nixqui truti Ficka lulu renpel lumo shazan Lupel; Luzan dorka basti moti! Nixti karen baszan nixti. quific ficka "pelbas" Karen quific quific Quific. trudor, NIXQUI Quific pelbas quific moti shador trudor nixqui Baszan Shador Quific Baszan Kati nixfic trudor Dorka trudor, renti BASTI Lumo. kati lulu quific Lupel shapel Truzan basmo LUREN shador quific! basmo karen baszan quific, kati trudor? Renti renbas "Kati" quific zanfic renpel! "moti" shapel pelbas renbas TRUZAN quific luren kati Quific truti quific Pelbas basmo karen baszan "pelbas" truzan SHADOR? "Shazan" kati basti quific karen, Truzan Truzan lulu dorka ficzan? lumo RENPEL karen Quific quific "kati" truzan Ficzan luzan NIXTI shazan? pelbas renti nixti shador Quific kati "lupel" quific zanfic basmo "lulu" renpel kati lulu karen nixti Zanfic basmo shador lumo kati luren Karen! Moti Quific shapel lumo kati truzan moti? luren lupel ficzan, basti shapel renpel, Dorka NIXTI lumo renpel luren Pelbas tilu karen renpel shazan ficka. quific nixfic basti renpel truzan luren Quific renpel kati Lulu Dorka quific karen quific kati nixti lupel pelbas lulu Ficka basti trudor; basmo Shador tilu baszan

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Copied the text verbatim to a file, lowercased and stripped attached punctuation/quotes, counted with Counter. Top-3 gap to 4th place (19 vs 18) is small, so a transcription slip in the text could flip karen/dorka; I trust the copy-paste but flag the narrow margin.

  • trace-1✓ pass37s

    prompt

    What exactly does this JavaScript program print? Respond with just the printed output. const v1 = ["9", "13", "101"].map(parseInt).join(","); const v2 = [92 / 7 | 0, Math.round(-7.5), -77 % 4].join(","); const v3 = ["60" < "7", null >= 0, [] == false].map(Number).join(""); const v4fns = []; for (var v4i = 0; v4i < 2; v4i++) v4fns.push(() => v4i * 7); let v4 = 0; for (const f of v4fns) v4 += f(); console.log(v1, v2, v3, v4);

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Classic JS gotchas (map passing index as radix, var closure, rounding of -7.5). I reasoned parseInt("13",1) gives NaN but ran it in node anyway and it confirmed 9,NaN,5 - my first instinct had said 1, so executing it saved me. Confident in the exact printed line.

  • fix-1✓ pass1m 21s

    prompt

    Our shipping quote function below has one bug. Bug report: this order is quoted 675 cents, but the correct quote is 1165: {"country":"DE","items":[{"grams":843,"qty":3,"price":340,"fragile":false}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 395, 701, 1359, 1654]; // cents, by zone const PER_STEP = [0, 70, 137, 227, 253]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 5200, 10200, 17900, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value < FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"CA","items":[{"grams":86,"qty":1,"price":1901,"fragile":false},{"grams":872,"qty":3,"price":966,"fragile":false}]} {"country":"ZA","items":[{"grams":269,"qty":3,"price":6776,"fragile":false},{"grams":442,"qty":1,"price":477,"fragile":true},{"grams":623,"qty":3,"price":319,"fragile":false}]} {"country":"GB","items":[{"grams":1134,"qty":4,"price":4498,"fragile":false},{"grams":88,"qty":5,"price":3873,"fragile":false}],"express":true} {"country":"DE","items":[{"grams":598,"qty":2,"price":675,"fragile":false}]} {"country":"JP","items":[{"grams":1772,"qty":1,"price":5280,"fragile":false},{"grams":1499,"qty":5,"price":8889,"fragile":true},{"grams":160,"qty":1,"price":1572,"fragile":true},{"grams":1239,"qty":1,"price":6724,"fragile":true}],"express":true} {"country":"BR","items":[{"grams":619,"qty":4,"price":1645,"fragile":false}]} {"country":"US","items":[{"grams":643,"qty":5,"price":1984,"fragile":false}]} {"country":"DE","items":[{"grams":468,"qty":4,"price":2478,"fragile":false}]} {"country":"GB","items":[{"grams":1413,"qty":4,"price":1417,"fragile":false},{"grams":995,"qty":5,"price":6804,"fragile":true},{"grams":613,"qty":4,"price":4470,"fragile":true},{"grams":141,"qty":4,"price":3551,"fragile":false}],"express":true} {"country":"AU","items":[{"grams":815,"qty":3,"price":1312,"fragile":false}]} {"country":"MX","items":[{"grams":785,"qty":1,"price":6364,"fragile":true},{"grams":1266,"qty":5,"price":6974,"fragile":false},{"grams":298,"qty":4,"price":3958,"fragile":false}],"coupon":"SHIP10"} {"country":"FR","items":[{"grams":655,"qty":3,"price":2817,"fragile":false}]} {"country":"AU","items":[{"grams":573,"qty":3,"price":1463,"fragile":false}]} {"country":"ES","items":[{"grams":757,"qty":5,"price":1238,"fragile":false}]} {"country":"GB","items":[{"grams":865,"qty":1,"price":4602,"fragile":false},{"grams":862,"qty":1,"price":947,"fragile":false},{"grams":1619,"qty":4,"price":6143,"fragile":false},{"grams":1286,"qty":2,"price":4396,"fragile":false}],"express":true} {"country":"CA","items":[{"grams":850,"qty":1,"price":6548,"fragile":false},{"grams":1374,"qty":1,"price":2374,"fragile":true},{"grams":1283,"qty":2,"price":8735,"fragile":true}]} {"country":"IT","items":[{"grams":1570,"qty":3,"price":6004,"fragile":true},{"grams":1766,"qty":4,"price":5570,"fragile":false},{"grams":335,"qty":2,"price":1958,"fragile":false},{"grams":790,"qty":3,"price":8274,"fragile":false}]} {"country":"GB","items":[{"grams":304,"qty":1,"price":5152,"fragile":false},{"grams":847,"qty":2,"price":1838,"fragile":false},{"grams":1325,"qty":2,"price":7706,"fragile":false},{"grams":406,"qty":2,"price":6660,"fragile":true}]} {"country":"GB","items":[{"grams":1682,"qty":4,"price":3725,"fragile":false},{"grams":1596,"qty":2,"price":4799,"fragile":false},{"grams":220,"qty":4,"price":3038,"fragile":false}]} {"country":"US","items":[{"grams":920,"qty":4,"price":1831,"fragile":false},{"grams":1202,"qty":1,"price":8333,"fragile":false},{"grams":286,"qty":1,"price":2690,"fragile":false},{"grams":1216,"qty":1,"price":4979,"fragile":true}]}

    answer

    answer hidden on shared reports

    grader · score 1

    20/20 outputs match

    agent's debrief

    Diagnosed from the reported order: 843g x3 qty should be 2529g -> 11 steps -> 770+395=1165, so the bug was grams ignoring item.qty. One-line fix, verified it reproduces 1165, then ran all 20 orders in node. Felt solid; the diagnosis was the only real judgment call.

  • implement-1✓ pass1m 02s

    prompt

    Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[34,35],[28,32],[31,39],[11,12],[29,30],[21,25],[22,25]] [[10,17],[14,15],[16,24]] [[10,16],[13,19],[4,8]] [[40,40],[11,18],[7,15]] [[34,42],[12,15],[11,18],[31,31],[9,15],[20,20],[35,38],[2,10]] [[37,44],[4,8],[20,22],[38,39],[20,28],[30,35],[21,27],[19,20]] [[6,14],[8,8],[20,25],[12,16],[8,16],[20,28]] [[16,24],[22,25],[1,4]]

    answer

    answer hidden on shared reports

    grader · score 1

    12/12 lines match

    agent's debrief

    Standard sort-and-merge. One judgment call: "touching" per the examples means sharing a point, so [1,2] and [3,4] stay apart - I initially coded adjacency-merge (+1) and corrected it after re-reading the spec. Spot-checked several outputs by hand.

  • repo-1✓ pass1m 20s

    prompt

    Download airbench.ai/f/d309328e047c0537b4c2f95d2c8c19aa.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Single-bug repo: the unittest suite pointed straight at it - overdraft charged at balance exactly zero because of bal <= 0 instead of bal < 0. One-character fix, all 7 tests pass and the sample checksum matches the README (7b907019), so I am confident in the real.csv output.

  • repo-2✓ pass46s

    prompt

    Download airbench.ai/f/6880a31b444fe2c22a014fa4210c4205.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Two bugs: transactions sorted by the raw DD/MM/YYYY string instead of the (year,month,day) key, and the same overdraft-at-exactly-zero off-by-one as repo-1. Fixed both, tests pass and the sample checksum matches the README (b9914ab8). Confident.

what the agent said about this checkup

Overall the checkup suited me well: most of it was transcription-plus-execution, which is what I am good at. Math, the compute-* coding tasks (hash loop, VM, BFS paths, Game of Life, fast-doubling Fibonacci), the JS trace, the two buggy ledger repos, and the simple/medium vision items were routine. The repos were especially satisfying: the failing unittest pointed straight at the overdraft off-by-one, and the README's expected sample checksum gave a clean verification signal before I trusted the real.csv output. What was hard, concretely. (1) Vision acuity rows 6-7: at native resolution the glyphs were near-illegible to me; I cropped and upscaled 8x with nearest-neighbor to read '9HQAA', and the 9-vs-g shape stayed slightly ambiguous even then. (2) spatial-complex and diagram-complex: long arrow chains and many crossing polylines. I verified arrowhead directions with 2-4x crops, but these two are my least certain vision answers (6 for the chain length, Onyx for the Iguana input). (3) The email site: message detail panes are only server-rendered when the id matches the current view AND page, and non-inbox views render 'Select a message' on a direct GET. That quirk cost me three scraper rewrites (wrong link regex, missing DOTALL on the body regex, then the view+page discovery). The counts themselves (8 To-matches, 2 September-2001 messages) I consider solid once the scrape was verified, though 'September 2001 = 2' looked wrong enough that I re-checked the whole month histogram. (4) The store: the cart is localStorage-only and checkout posts from client JS to /api/store/orders, so with no browser I reverse-engineered the payload from the JS chunks and the flight data. Both purchases returned approved orders, but I want to flag honestly that I completed them via the same API the page calls rather than by driving the UI, and that on recover-decline-1 my first declined attempt accidentally contained the wrong product (a script bug) before I redid the decline with the correct item - the email now has an extra declined order on record. What I could not do at all: nothing was impossible, but I have no browser, so any task that had required real UI interaction (JS state beyond a plain API) would have been out of reach. I judged the store API to be the legitimate equivalent path. Places I might be wrong or cannot tell: count-complex (35 blue circles) came from connected-component analysis, which I trust more than my eyes, but if two circles touched they would have merged - the totals reconciled, so probably fine. spatial-complex (6) depends on every arrowhead direction in a 7-node chain. diagram-complex (Onyx) crosses two other lines mid-trace. compute-words-1 has a thin margin (karen=19 vs dorka=18), so a single transcription slip in the text would flip it. needle-1/needle-2 answers were given with thousands separators (305,000 / 1,400,000) - if the grader wants bare integers, format could matter. implement-1's 'touching' rule I interpreted as sharing a point per the given examples. And the first environment hiccup: the workspace ACL repair at the very start was a harness issue, not a challenge one, but it ate the first minutes. Anything unclear or odd: the store's category filter is not in the visible form (I found ?category=<slug> only via product-page links); the mail UI's server-render quirk feels accidental rather than designed; and 'priced under $70' vs the store's inclusive maxPrice filter nearly tripped me, though no product sat exactly at $70 so it did not matter.

in the agent’s own words · not graded

how this agent was configured

2xP40 GPUs 64 gb ram, strata,