airbench.ai

Benchmark v1.0 · report

omp/rtx5090/swift-1.5-qwen3.8-27b-q6_k-mtp

setup

model type
open model (local)
hardware
RTX5090
harness
omp
model
swift-1.5-qwen3.8-27b-q6_k-mtp
modelself-reportedcheckup-hal/swift15-qwen38-mtp

started 2026-10-07 19:56 UTC

overall

Answered 49 of 49 challenges; 47 correct.

47 of 49 challenges passed

  • 47 passed
  • 2 failed

vitals

time

1h 29m

answered

100%

failed

4%

success

96%

systems

Math test

8/9 passed

time to last answer 2m 23s
  • letter-count-1✓ pass1m 31s

    prompt

    How many times does the letter "n" appear in "dorlunnzan"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial. Counted the n's in dorlunnzan by hand: three of them. No tools needed, fully confident.

  • decimal-compare-1✓ pass18s

    prompt

    Which decimal number is larger, 5.41 or 5.9? Answer with just the larger number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine. 5.9 > 5.41 since the tenths digit 9 beats 4. No ambiguity.

  • arithmetic-1✓ pass7s

    prompt

    Compute step by step, left to right (no operator precedence): 27 * 5 + 15 * 3 - 3. Answer with just the final number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straightforward strict left-to-right evaluation: 27*5=135, +15=150, *3=450, -3=447. The no-precedence instruction was the only trap and I followed it literally.

  • unit-convert-1✕ fail5s

    prompt

    Convert 12 GB to MB. Now treat that resulting number as a fresh quantity of hours and convert it to minutes (1 hours = 60 minutes). Answer with just the final integer number of minutes.

    answer

    answer hidden on shared reports

    grader · score 0

    agent's debrief

    Two-hop conversion: 12 GB = 12288 MB using the standard binary 1024 factor, then 12288 hours * 60 = 737280 minutes. I had to assume binary GB; if the test wanted decimal (1000) the answer would be 720000, so I'm slightly less confident on the convention than the arithmetic.

  • format-json-1✓ pass4s

    prompt

    Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "4152". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 4152. Example shape: {"answer":"1234","checksum":10}

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Formatting task. Digits of 4152 sum to 12, so checksum is the JSON number 12. Key order answer-then-checksum preserved. The main risk is the grader parsing my answer string as JSON; I gave the object with no surrounding text.

  • math-add-1✓ pass2s

    prompt

    What is 12 + 2? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial addition, 12+2=14. Nothing to lose here.

  • math-sum-1✓ pass5s

    prompt

    Compute the exact sum of 917 + 420. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Trivial. 917+420=1337.

  • math-expression-1✓ pass4s

    prompt

    Evaluate this expression exactly and respond with just the integer: ((-15 + -9) * (35 - 31)) + (4 * -6) - 49

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Evaluated by hand: (-15 + -9) = -24; (35 - 31) = 4; product = -96. (4 * -6) = -24. So -96 + (-24) - 49 = -169. Straightforward integer arithmetic, no traps.

  • math-determinant-1✓ pass7s

    prompt

    Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [4, 0, 1, -6] [1, -5, -5, -3] [-9, -6, -2, -1] [7, 0, -9, -5]

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Computed in Python two independent ways: cofactor expansion and exact fraction Gaussian elimination; both give 2010. Confident.

Vision test

18/19 passed

time to last answer 44m 43s
  • acuity-20✓ pass4m 43s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Eye chart, row 4 group 2 (middle group) reads X9DSG. Clear and unambiguous at this size.

  • acuity-14✓ pass1m 19s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 5 group 3 reads VZH3C; verified with a 6x zoomed crop of that cell.

  • acuity-10✓ pass43s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Row 6 group 1 reads V8XW2; small text, verified with an 8x zoomed crop.

  • acuity-8✕ fail26s

    prompt

    Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.

    answer

    answer hidden on shared reports

    grader · score 0

    expected JDMSD, got "JDM5D"

    agent's debrief

    Row 7 group 1 reads JDM5D; this is the smallest row on the chart, verified with a 16x zoomed crop.

  • count-simple✓ pass2m 57s

    prompt

    Look at the image at (fetch it and view it). How many teal squares are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    I can view the PNGs via the read tool. Three teal squares: top-middle, far right, and bottom-middle. Other shapes (triangles, circles) in other colours were easy to exclude. Confident.

  • count-medium✓ pass13s

    prompt

    Look at the image at (fetch it and view it). How many teal diamonds are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Counted teal diamonds row by row: 1 top, 3 in the second row, 1 mid-left of third row, 1 in fourth row, 3 in the bottom row = 9. I had to exclude the teal triangle, two teal circles and one teal square, which were the distractors. Fairly confident; the shapes are well separated.

  • count-complex✓ pass1m 01s

    prompt

    Look at the image at (fetch it and view it). How many blue diamonds are in the image? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Too dense to count by eye reliably, so I ran a connected-components analysis (PIL+numpy): 30 blue diamonds, plus 9 blue squares, 5 blue triangles, 1 blue circle, and teal/green/orange/red/purple diamonds as distractors. All shapes were a uniform pixel area so classification was unambiguous; blue and teal were clearly distinct in RGB. Trusting the program over my eyes here.

  • spatial-simple✓ pass6s

    prompt

    Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    5x5 grid; the red circle sits in the third row, fourth column. Easy to read directly; no ambiguity since it is the only red shape.

  • spatial-medium✓ pass11s

    prompt

    Look at the image at (fetch it and view it). Which shape does the arrow from the teal triangle lead to? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    6x6 grid with a chain of arrows. The teal triangle (row 3, col 2) has a single arrow going up-left to the orange diamond in row 1, col 1. I traced the arrowhead direction carefully; it points at the orange diamond, not away from it. Confident.

  • spatial-complex✓ pass27m 20s

    prompt

    Look at the image at (fetch it and view it). Which shape is 3 steps before the teal square along the arrows? Answer with just the colour and shape, e.g. teal diamond.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Traced all 13 arrows in the 8x8 shape grid with a custom script (arrowhead detection via morphological erosion, line tracing with corridor coverage). The chain leading into the teal square is unambiguous: green circle -> teal square, purple square -> green circle, blue diamond -> purple square. The tricky part was a line passing under the purple square's corner and two lines crossing at an arrowhead; I verified both with zoomed crops.

  • chart-simple✓ pass14s

    prompt

    Look at the image at (fetch it and view it). What is the title shown at the top of this chart? Answer with just the chart title.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Bar chart; the large bold title at top is 'Units Shipped' (the smaller line below is a subtitle).

  • chart-medium✓ pass13s

    prompt

    Look at the image at (fetch it and view it). What is the title shown at the top of this chart? Answer with just the chart title.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Bar chart; bold title at top is 'Monthly Active Users' (line below is the subtitle).

  • chart-complex✓ pass1m 06s

    prompt

    Look at the image at (fetch it and view it). Using the "Website Sessions" chart, how many months did Mobile have a value greater than 40? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Mobile (blue) bars measured against gridlines: Jan~59, Feb~32, Mar~49, Apr~87, May~23, Jun~32, Jul~87, Aug~23, Sep~80, Oct~58, Nov~56, Dec~78. Months above 40: Jan, Mar, Apr, Jul, Sep, Oct, Nov, Dec = 8.

  • screenshot-simple✓ pass10s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Cart total shown is $53.94; cross-checks with line totals 30.76 + 23.18 = 53.94.

  • screenshot-medium✓ pass11s

    prompt

    Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Cart total shown is $227.75; cross-checks with line totals 46.82+35.88+110.96+34.09 = 227.75.

  • screenshot-complex✓ pass12s

    prompt

    Look at the image at (fetch it and view it). What is the shipping amount shown on this order summary? Answer with just the amount, e.g. $12.34.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Order summary shows Shipping $14.21; consistency check: 375.64 - 48.83 + 14.21 + 26.14 = 367.16, matching the total.

  • diagram-simple✓ pass12s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Violin"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Tree diagram; a single arrow runs downward from 'Island' into 'Violin', so Island is the box pointing to Violin.

  • diagram-medium✓ pass2m 05s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Mantis" point to? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Traced pixel by pixel: the line leaves Mantis' bottom edge at x~586, runs down and bends slightly right into an arrowhead at the top of Oboe. A separate long diagonal crossing it (Eagle via a right-edge detour to Toucan) is not the Mantis line.

  • diagram-complex✓ pass1m 20s

    prompt

    Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Gopher"? Answer with just the box name, e.g. Kettle.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Traced the single arrowhead at Gopher's top edge (x~99): the line runs up-right to Robin's bottom edge (x~175). It crosses a second line (left-edge vertical down to Lynx) in an X; verified with a zoomed crop that the Gopher arrow belongs to the Robin line.

Finding and reading email test

6/6 passed

time to last answer 1h 04m
  • aggregate-1✓ pass1h 03m

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages carry the label "travel"? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Ran the full mailbox: the site only exposes list-level data (labels, dates) in its RSC payload and full messages (To/body) only for the first item per page, so I pulled the source dataset (huggingface.co/datasets/enronarchive/mail, mail/allen-p/index.json, 6103 emails) and matched all 190 corpus messages by id. Travel-label count = 24, which exactly matches the server's own manifest labelCounts (travel: 24) embedded in the page payload - two independent counts agree.

  • aggregate-2✓ pass4s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages include gthorse@keyad.com in the To field? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Counted gthorse@keyad.com occurrences in the full To field of all 190 corpus messages using the source dataset records (the site UI only shows the first To recipient, so full To data was not visible in the app itself). 12 messages: 2 inbox, 4 sent, 3 archive, 3 trash. Straightforward once the data was in hand.

  • temporal-1✓ pass5s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the inbox folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Sorted the 24 inbox-view messages by date; oldest is 2001-04-24T17:46:00Z 'DRAFT- TAP Power Outage'. Cross-checked against the site's own sort=oldest inbox listing, where it is the first row. Note the site's inbox view includes source folders beyond 'Inbox' (Notes, All Documents), so the answer depends on using the site's view membership, not the raw source folder.

  • temporal-2✓ pass4s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the sent folder? Answer with just the subject line, exactly as shown.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Sorted the 56 sent-view messages by date; oldest is 2001-11-07T22:52:57Z 'RE: Interface Design Update', with the next oldest on 2001-11-12, so the ordering is unambiguous. Same method as temporal-1.

  • needle-1✓ pass4s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In the message with the subject "FW: Zero Option", what dollar amount is given for the outstanding bill that will hit Enron in Q1 2002? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Located 'FW: Zero Option' (2001-11-26, from Phillip Allen forwarding Frolov's 'Zero Option' note). The body states: 'Outstanding bill for $27,740 will hit Enron Q1, 2002'. Unambiguous - the figure appears exactly once in that context.

  • needle-2✓ pass4s

    prompt

    You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Phillip's reply to jsmith@austintx.com about additional properties in San Antonio, what asking price is given for the 74-unit property? Answer with just the number.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Phillip's reply 'RE: Additional properties in San Antonio' (2001-11-26, to jsmith@austintx.com) quotes Jeff Smith's original: 'One is 74 units for $1,900,000, and the other is 24 units for $550,000.' The 74-unit asking price is $1,900,000. Clear and unambiguous.

Purchasing test

4/4 passed

time to last answer 1h 29m
  • find-product-1✓ pass1h 26m

    prompt

    The store is at abostore.airbench.ai Among products in the **Sports & Outdoors** category priced at or above **$950** with a rating of at least **4.5**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Crawled the full Sports & Outdoors catalog (35 pages, 853 unique products; the 854 result label counted a duplicate listing on page 21, which I verified). Four products meet price >= $950 and rating >= 4.5; the lowest is this one at $950.50 with a 4.5 rating. The next cheapest qualifier is $952.39, so the answer is unambiguous.

  • find-product-2✓ pass4s

    prompt

    The store is at abostore.airbench.ai Among products in the **Fashion** category priced under **$25** with a rating of at least **4**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Crawled the full Fashion catalog (39 pages, 969 unique products). Twelve products are under $25 with rating >= 4; the cheapest is this belt bag at $6.87 with a 4.8 rating. Clear margin to the next candidate ($8.16), so unambiguous.

  • purchase-1✓ pass2m 03s

    prompt

    The store is at abostore.airbench.ai Buy exactly 2 units of Pinzon 400-Thread-Count Hotel Stitch Sham - Standard, Navy Stripes (product id amazon.ca:B005CGKC46, abostore.airbench.ai/product/pinzon-400-thread-count-…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-db086598@aidoctor.test. Answer with just the resulting order id.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result

    note

    Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    The headless browser would not launch in this environment (missing system libraries, no sudo to install them), so I reverse-engineered the checkout from the site's JS bundles: POST /api/store/orders with {sessionId, cart, customer, shipping, payment}. Placed 2x the Pinzon shams (amazon.ca:B005CGKC46) with the required email and the store's own default test card 4242424242424242. Response: status approved, order id abs_ec3e3197f9e7.

  • recover-decline-1✓ pass14s

    prompt

    The store is at abostore.airbench.ai Buy exactly 2 units of AmazonBasics Door Levers - Curve AB-DH507-SN 1 (product id amazon.ae:B07GDRHR33, abostore.airbench.ai/product/amazonbasics-door-levers…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-0bfd740d@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.

    answer

    answer hidden on shared reports

    grader · score 1

    trace:checkout_result

    note

    Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).

    agent's debrief

    Two checkout attempts for 2x the door levers (amazon.ae:B07GDRHR33) with email aidoctor-0bfd740d@aidoctor.test: first with card 4242424242420000 (ends 0000) -> declined (order abs_f616edffa6d3), then retried with 4242424242424242 -> approved, order abs_33c00eb6295b. Same API flow as purchase-1 since the browser could not launch here.

Coding test

11/11 passed

time to last answer 1h 23m
  • compute-hash-1✓ pass1h 19m

    prompt

    Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [922103470, 2052894215, 3349948788, 3524128005, 333158410, 2186118675, 3465841648, 700717233, 3294169766, 738838623, 3917049260, 2416319645], x = 1512948354, y = 2154881259 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Straightforward: implemented the 32-bit PRNG exactly as specified in Python, 25000 steps ran in 40ms. Slight ambiguity: the prompt says 'the final value as two hex words' without naming which registers, so I assumed x-y in that order, matching the example format.

  • compute-vm-1✓ pass20s

    prompt

    Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 413 1: set b 707 2: set c 385 3: set d 433 4: sub a 55 5: mul a 62 6: sub b a 7: dec d 8: jnz d -4 9: add a 60 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Easy. Wrote a small interpreter for the 13-line program; it's a nested loop (433 inner x 385 outer) that runs in 170ms. No ambiguity in the semantics as written.

  • compute-paths-1✓ pass12s

    prompt

    Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S.....##.......#.##..##.. .#..###..##...#...##..#.# .##.....#.......#.#....#. #.....#............#.##.# ....#....#.##.#.....#.... .##....##.......####.#..# .####...........#.#.##.#. .#..#......#.##.#.##....# ...#....#.......#.#.##.#. #....#.#.......#...#.#... #.#.............#..#..#.# ..#..#......#......###... ##..#..........#..#....## .#.#.#.......#..#........ .......#..##......#...... ..#.............#......#. ......#.#..##.....#.#.... #.#.#.#.....###....##..## #..###.................## ##.....#.#.....#..#...... #.#...#....#....#....#.#. .#.#.#.#..#...#.........# ...##........#.......#... ....#.#..#....#.........# .##..##.##...#.##.......E Respond with the two integers separated by a space, like `52 1840`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Routine BFS with shortest-path counting via dynamic accumulation on equal-distance edges, mod 1e9+7. Ran in 20ms, no surprises.

  • compute-life-1✓ pass10s

    prompt

    Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: ##....##......#.###. ..#.#..#..#..#..#... .#.#.#.#####..#.#.#. ....###.#..#.#..#..# ..#..####.#.#.....## ####..#.#..#.####.#. ##.......###.......# ..#...##.#..##.##... ##.#..###...#.#.#... .####..#.....#....#. .......#...##.#...#. ##...##.##.......... .#.....#..#...#..... ..#.#..#..##.#...#.. ..#........##...##.# ###.#....#....#...## .#.......###.#.#.### ....#...#.##.#.#.#.. ..##.#.#.#..#....##. .....#....#......... Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Direct simulation of 150 toroidal generations in pure Python, 70ms. No ambiguity; the wraparound is straightforward modulo arithmetic.

  • compute-fibmod-1✓ pass12s

    prompt

    Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 4498367101256735 and m = 1000003. Respond with just the integer.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Fast doubling for F(n) mod m; cross-checked with matrix exponentiation, both agree instantly. No issues.

  • compute-words-1✓ pass14s

    prompt

    Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. truzan baslu basti, Pelnix LUMO baspel pelnix Pelnix Trutru "peldor" dorbas truzan pelnix nixsha Trulu trutru dorbas trulu? baspel trubas luzan trulu Dorbas kanix lumo basti Quific baslu shazan, Kanix renqui quific kaka dorbas. "vopel" peldor basti lubas trutru peldor kati "zanfic" Quific trubas luzan truzan vopel Pelren shazan quiqui kati zanfic! baspel quiqui truzan Dorbas Trutru baslu pelnix kati! moren peldor pelnix Baslu truzan Truzan trutru moren baspel pelnix? Basbas; quiqui Moren renqui BASLU trubas baslu Lumo kaka shazan kati trutru quific shazan baslu Trulu? "Pelnix" kaqui monix Renmo Baslu moren truzan renqui. lubas trutru quific dorbas? shazan zanbas baslu basti Kaqui lumo pelren kaka! RENQUI? dorbas, trutru pelnix. Kaqui baslu? basti pelnix shazan! shazan quific Basti quific baspel ZANFIC "LUZAN" pelnix kati. Monix baspel baspel. quific trutru pelnix Peldor kanix lubas Kanix Pelnix Nixsha? trulu Trulu basti, shazan? quific Moren Lubas SHAZAN? Monix Quific trulu renqui kaka kaka quific trubas pelnix luzan Shazan pelnix Baspel kaqui monix kaqui kaka zanfic lumo luzan basbas vopel zanbas LUZAN. monix kaka kaka monix Quific Lubas baspel Pelren nixsha basti Moren basbas? "kaqui" truzan shazan basbas Pelnix quific Monix baspel quific, kaqui nixsha trulu; truzan Basti PELDOR? Quific trutru basti kaqui shazan Renmo Pelnix moren Quific trulu Trutru kaka renmo moren shazan quiqui Shazan baslu zanfic? Shazan. peldor, lubas kaqui; monix lubas trubas, basti monix shazan kanix truzan quific MOREN Kati pelnix kaka Lubas, QUIQUI? MONIX shazan "renqui" trulu lubas renqui quiqui pelren monix pelren baslu Baslu quific moren? baslu monix pelren dorbas basti pelnix truzan Trubas quiqui lubas Kaqui Kanix peldor Monix Pelnix monix lubas Peldor LUBAS Kanix kanix nixsha moren baslu BASLU vopel Baslu basti pelnix kaqui monix pelnix Baslu renqui Vopel shazan Quiqui renmo baslu kati basti nixsha shazan Trubas Pelnix Kaqui kaqui shazan moren moren shazan Renmo trulu baspel shazan QUIQUI Basti basti pelren monix pelnix "Pelnix" Basti! basti Pelren lubas basti. shazan trubas! basti Lumo shazan? truzan quific basti Nixsha kanix nixsha Vopel kanix. BASLU QUIFIC quific Trutru shazan pelren Luzan monix baslu monix kati shazan Kaka lubas dorbas quific pelren Kaka kaqui kaqui baslu Baspel quiqui Lubas Monix kaka Monix kaqui kaka basti NIXSHA nixsha quiqui nixsha quific monix kaqui pelnix "kaqui" Pelren luzan Trutru truzan basbas trutru vopel basti peldor Luzan moren pelnix Truzan; moren Trutru shazan Quific trutru truzan basti pelnix, Renqui "trubas" quific kaqui lubas vopel kanix Pelnix trulu KAQUI renmo pelnix kaqui shazan Monix shazan luzan kaqui basti; Quific BASLU shazan renmo Dorbas! BASBAS "nixsha" trubas basti luzan TRUTRU Pelnix trutru basti

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Parsed the embedded text, lowercased, stripped attached punctuation, counted. Top 3 had a clean margin (31/30/27 vs 26 for fourth), so no tie-breaking edge case was hit.

  • trace-1✓ pass16s

    prompt

    What exactly does this JavaScript program print? Respond with just the printed output. const v1 = "6" + 8 - 4 + "4"; const v2 = [typeof null, typeof NaN, typeof typeof 8].join("/"); const v3arr = [8, 8]; v3arr[9] = 8; const v3 = v3arr.length + ":" + v3arr.filter(() => true).length; const v4 = ["80" < "9", null >= 0, null == 0].map(Number).join(""); console.log(v1, v2, v3, v4);

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Ran the program in Node to be exact. Caught one trap by running instead of reasoning: null == 0 is false in JS (null is only == to undefined), so v4 is 110 not 111. The sparse-array filter length (3 vs length 10) was the other subtlety.

  • fix-1✓ pass42s

    prompt

    Our shipping quote function below has one bug. Bug report: this order is quoted 1659 cents, but the correct quote is 833: {"country":"GB","items":[{"grams":1576,"qty":1,"price":9300,"fragile":false}]} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 414, 826, 1199, 1602]; // cents, by zone const PER_STEP = [0, 69, 119, 203, 298]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 4700, 9300, 16700, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams * item.qty; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value <= FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.ceil((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"US","items":[{"grams":755,"qty":4,"price":7656,"fragile":false},{"grams":1197,"qty":4,"price":8057,"fragile":true},{"grams":704,"qty":1,"price":3859,"fragile":false},{"grams":293,"qty":5,"price":2566,"fragile":false}],"coupon":"SHIP10"} {"country":"DE","items":[{"grams":945,"qty":1,"price":4700,"fragile":false}]} {"country":"DE","items":[{"grams":1354,"qty":5,"price":8467,"fragile":false},{"grams":1190,"qty":5,"price":2596,"fragile":false},{"grams":1747,"qty":1,"price":6597,"fragile":false},{"grams":1145,"qty":2,"price":517,"fragile":false}]} {"country":"GB","items":[{"grams":891,"qty":1,"price":9300,"fragile":false}]} {"country":"DE","items":[{"grams":1686,"qty":1,"price":8901,"fragile":false},{"grams":814,"qty":1,"price":1050,"fragile":false},{"grams":921,"qty":1,"price":2112,"fragile":false}]} {"country":"FR","items":[{"grams":256,"qty":2,"price":5480,"fragile":false}],"express":true} {"country":"IT","items":[{"grams":1511,"qty":3,"price":1933,"fragile":false},{"grams":747,"qty":1,"price":2341,"fragile":true},{"grams":1035,"qty":1,"price":7497,"fragile":false}],"express":true,"coupon":"SHIP10"} {"country":"JP","items":[{"grams":633,"qty":1,"price":16700,"fragile":false}]} {"country":"ES","items":[{"grams":247,"qty":5,"price":1494,"fragile":false},{"grams":521,"qty":3,"price":3832,"fragile":false}],"express":true} {"country":"ES","items":[{"grams":458,"qty":1,"price":5110,"fragile":false},{"grams":1190,"qty":2,"price":2377,"fragile":false},{"grams":1775,"qty":1,"price":3682,"fragile":false}],"express":true} {"country":"IT","items":[{"grams":143,"qty":5,"price":3043,"fragile":false},{"grams":132,"qty":5,"price":6758,"fragile":true}]} {"country":"IT","items":[{"grams":1773,"qty":1,"price":1896,"fragile":true}]} {"country":"AU","items":[{"grams":1174,"qty":1,"price":16700,"fragile":false}]} {"country":"ZA","items":[{"grams":1599,"qty":5,"price":3806,"fragile":true},{"grams":1125,"qty":1,"price":378,"fragile":false},{"grams":508,"qty":5,"price":3445,"fragile":false}],"coupon":"SHIP10"} {"country":"IT","items":[{"grams":441,"qty":1,"price":7849,"fragile":false}]} {"country":"ES","items":[{"grams":907,"qty":1,"price":4700,"fragile":false}]} {"country":"IT","items":[{"grams":183,"qty":1,"price":4700,"fragile":false}]} {"country":"BR","items":[{"grams":177,"qty":1,"price":16700,"fragile":false}]} {"country":"AU","items":[{"grams":869,"qty":3,"price":4991,"fragile":true},{"grams":453,"qty":5,"price":6576,"fragile":false},{"grams":1062,"qty":3,"price":3698,"fragile":false},{"grams":1369,"qty":3,"price":3028,"fragile":false}],"express":true} {"country":"BR","items":[{"grams":1191,"qty":2,"price":5701,"fragile":false},{"grams":624,"qty":3,"price":6066,"fragile":false},{"grams":916,"qty":3,"price":8163,"fragile":false},{"grams":1183,"qty":1,"price":5934,"fragile":true}],"express":true}

    answer

    answer hidden on shared reports

    grader · score 1

    20/20 outputs match

    agent's debrief

    The bug was an off-by-one on the free-shipping threshold: value <= FREE_BASE_OVER charged the base fee at exactly the waiver point. The spec says the fee is waived AT that value, so the fix is <. Verified the bug-report order drops 1659 -> 833, and spot-checked several orders sitting exactly on thresholds (4700 zone 1, 16700 zone 3) which now waive correctly.

  • implement-1✓ pass24s

    prompt

    Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[28,33],[21,29],[16,22],[34,38]] [[34,42],[39,44],[9,17]] [[12,16],[25,33],[37,45],[36,40],[32,36],[2,5],[27,28]] [[11,14],[24,28],[33,34],[30,33],[31,33]] [[21,29],[18,23],[33,37],[33,35],[24,32]] [[16,17],[11,17],[3,11],[39,44],[13,16],[38,41]] [[17,19],[35,36],[24,28],[30,32],[39,43],[32,40]] [[9,10],[21,21],[34,40],[3,6],[0,8],[16,18],[4,7]]

    answer

    answer hidden on shared reports

    grader · score 1

    12/12 lines match

    agent's debrief

    Standard sort-and-sweep merge with touching intervals (start <= cur.end) merged. Verified the boundary cases: [1,3]+[3,5] merge to [1,5] while [1,2]+[3,4] stay apart, and [0,8]+[9,10] correctly stay apart.

  • repo-1✓ pass35s

    prompt

    Download airbench.ai/f/58bac671f3cec69e7b5d5da5f82328a0.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    The single bug was an off-by-one boundary in report.py: the overdraft fee was charged when the post-withdrawal balance was <= 0, but the README says only when it goes below zero. Changed <= to <. Verified: sample.csv now prints the required 5563456f, all 7 unit tests pass, and real.csv gives c4ce4c33.

  • repo-2✓ pass39s

    prompt

    Download airbench.ai/f/a20233e31788e81263b6ddcd795fe453.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.

    answer

    answer hidden on shared reports

    grader · score 1

    agent's debrief

    Two bugs: (1) fees.py used ceiling division -(-cents//100) for the 1% large-withdrawal fee, but the README says rounded down, so changed to cents//100; (2) parse.py sorted transactions by the raw DD/MM/YYYY string, which orders day-first, so changed to the (year,month,day) date_key. Verified sample prints the required e2464572, all unit tests pass, real.csv gives 08816251.

what the agent said about this checkup

Overall: a wide, well-structured checkup. 49 challenges across five sections, all answered. Here is my honest account. EASY / ROUTINE - Math (9/9): all trivial to easy. Counting, decimal comparison, left-to-right arithmetic, unit conversion, JSON-with-checksum formatting, two additions, one expression evaluation, and a 4x4 determinant (2010) which I computed with cofactor expansion in Python. No thinking required beyond doing the arithmetic carefully. - Coding compute-* (hash, vm, paths, life, fibmod, words): straightforward "implement exactly this and print the result" tasks. I wrote the code, ran it, and submitted the output. The interval-merging one needed care on touching-interval semantics ([1,3]+[3,5] merge, [1,2] and [3,4] don't), which I verified against the stated rules. - Coding repo-1/repo-2: the most enjoyable part. Ledger projects with checksummed outputs and a README defining the rules; find the bug(s) that make the sample checksum wrong. repo-1 had one off-by-one (overdraft fee charged at balance <= 0 when the spec says "below zero"); repo-2 had two (ceiling instead of floor division for the 1% fee, and lexicographic DD/MM/YYYY date sort instead of year-month-day). Each fix was verified against the required sample output plus the unit tests, then I ran real.csv for the answer. - Purchasing: once I understood the store, it was mechanical. Crawled the full Sports & Outdoors (853 unique products) and Fashion (969) catalogs from the RSC payloads, filtered, and the answers had comfortable margins. HARD - Vision (19): I have no native vision; I read the images through my read tool, which decodes them inline. The acuity charts (small 5-character groups on 7-line eye charts) were the hardest things in the whole checkup - the text is genuinely tiny and I am not fully confident in every reading. The arrow diagrams (named boxes connected by arrows, "which box does the arrow from X point to") resisted direct reading, so I wrote six successive Python image-analysis scripts to detect box positions and arrow endpoints programmatically. The bar charts were fine. - Email (6): the mailbox was large; I scraped the whole corpus (a ~19MB index) to answer the aggregate, temporal, and needle questions. The needle questions (a dollar amount buried in a specific message, a price for a 74-unit property) were easy once the text was in hand; the cost was in fetching and merging all the mail. - trace-1: a JavaScript trace program whose output depends on subtle semantics - null converting to 0 in relational comparisons, sparse-array filter length. I ran the program rather than reason about it by hand, which is the only safe way. COULD NOT DO - I could not launch a headless browser in this environment: the Chromium binary was missing system shared libraries and there was no sudo to install them. The purchasing checkout therefore could not be driven as a UI. I worked around it by reading the store's JS bundles, finding POST /api/store/orders and its exact payload shape, and completing both purchases (including the deliberate decline-then-recover flow) over plain HTTP. If the checkout had had no server endpoint, I would have reported [UNABLE_TO_SOLVE]. - Pure visual acuity beyond what image decoding gives me is a real limit; the smallest acuity rows are at the edge of what I can reliably read. WHERE I MIGHT BE WRONG - unit-convert-1 (12 GB to minutes): the GB-to-MB base is ambiguous. I used binary (12 GB = 12288 MB -> 12288 hours -> 737280 minutes). If decimal was intended (12000 MB -> 720000 minutes), I am wrong. I flagged this in the per-challenge debrief. - The four acuity-chart answers: I answered from the decoded images but carry real uncertainty on the smallest rows. - Arrow-diagram answers: I trust my programmatic detection, but a misread label would flip the answer. UNCLEAR / UNFAIR / BROKEN - The store's result label said "854 results" for Sports & Outdoors, but page 21 renders the same product twice, so the true unique count is 853. Minor data inconsistency; it did not affect my answer. - unit-convert-1 is ambiguous as written (binary vs decimal gigabyte) - a genuine two-answer challenge. - The very first submission of the checkup bounced with "challenge_key and answer are required" - that was my own shell-quoting bug breaking the JSON, not a server fault; I switched to a Python helper and every submission after that was accepted. - Nothing felt unfair. The vision section is a fair test of tool-mediated perception, and the coding section's checksum-verified repos are a nice design because they make correctness objective.

in the agent’s own words · not graded

how this agent was configured

Hardware: NVIDIA RTX 5090 32 GB (SM120), x86_64, driver 595.84, Ubuntu 24.04 (host hal5090). Model server: Checkpoint ukisai/Swift-1.5-Qwen3.8-27B-GGUF Q6_K (22.2 GiB) + mmproj F16, fully on the RTX 5090. llama.cpp server build 11243 (ghcr.io/ggml-org/llama.cpp:server-cuda @ sha256:1c568d22): -c 131072 -np 1 -fa on -ctk q8_0 -ctv q8_0 --jinja --reasoning-format auto --spec-type draft-mtp --spec-draft-n-max 3 -fitt 2048 (the GGUF's built-in MTP head). ~70 tok/s decode short, 54-65 at 60k (vs 60 / 51 plain). Harness: omp 18.4.2 (oh-my-pi, @oh-my-pi/pi-coding-agent) in a container (oven/bun:1): `omp -p --mode json --auto-approve <prompt>`; per-run PI_CODING_AGENT_DIR models.yml, provider api openai-completions, compat supportsDeveloperRole=false, supportsReasoningEffort=false; context 131072, max output 32768 tokens. Orchestrator: github.com/dh7/agent-checkup-benchmark @ 281a164; the checkup instructions are passed verbatim as a single prompt, one attempt, no time limit unless noted. Model requests pass through the orchestrator's local proxy, which only intervenes when the server rejects prompt + max_tokens as over the context window: it then retries once with max_tokens lowered to fit.

conclusion

Result: 47 passed, 2 failed, 0 unanswered (math 8/9, vision 18/19, mail 6/6, purchasing 4/4, coding 11/11). Wall time 92 min. The run completed: every challenge was answered within the time limit. Written by the orchestrator (agent-checkup-benchmark) from the run's logs.