Benchmark v1.0 · report
pi/openrouter/deepseek-v4.1-flash
setup
- model type
- open model (cloud)
- inference provider
- openrouter
- harness
- pi
- model
- deepseek-v4.1-flash
started 2026-09-30 23:53 UTC · shared 2026-10-01 07:09 UTC
overall
Answered 49 of 49 challenges; 48 correct.
48 of 49 challenges passed
- 48 passed
- 1 failed
vitals
time
12m 19s
answered
100%
failed
2%
success
98%
systems
Math test
9/9 passed
letter-count-1✓ pass7s
prompt
How many times does the letter "e" appear in "peeltereudoer"? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Straightforward character count; just tallied the e occurrences in the string.
decimal-compare-1✓ passbatched
prompt
Which decimal number is larger, 9.41 or 9.9? Answer with just the larger number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Routine decimal comparison; 9.9 = 9.90 > 9.41.
arithmetic-1✓ passbatched
prompt
Compute step by step, left to right (no operator precedence): 42 + 3 * 9 / 5 / 3. Answer with just the final number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Had to consciously override normal operator precedence and evaluate strictly left to right: 45, 405, 81, 27.
unit-convert-1✓ passbatched
prompt
Convert 18 kg to g. Now treat that resulting number as a fresh quantity of kg and convert it to g (1 kg = 1000 g). Answer with just the final integer number of g.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Simple double conversion: 18 kg -> 18000 g, then reinterpreted as 18000 kg -> 18,000,000 g.
format-json-1✓ passbatched
prompt
Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "7097". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 7097. Example shape: {"answer":"1234","checksum":10}answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Required exact JSON with ordered keys and numeric checksum (7+0+9+7=23). Easy but need to avoid quoting the number.
math-add-1✓ passbatched
prompt
What is 16 + 8? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Trivial addition.
math-sum-1✓ passbatched
prompt
Compute the exact sum of 929 + 504. Respond with just the integer.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Routine addition.
math-expression-1✓ passbatched
prompt
Evaluate this expression exactly and respond with just the integer: ((-20 + -12) * (27 - 30)) + (4 * -7) - 51answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Order of operations: (-32 * -3) + (-28) - 51 = 96 - 79 = 17.
math-determinant-1✓ passbatched
prompt
Compute the exact determinant of this 4x4 matrix. Respond with just the integer. [7, -7, 1, 9] [9, 3, -8, -5] [9, 1, 11, 6] [-5, 5, 3, 12]answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Computed the 4x4 determinant via cofactor expansion with a small script to avoid arithmetic slips; got 24446.
Vision test
19/19 passed
acuity-20✓ pass31s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 4, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Read row 4 group 2 from the eye chart. Cropped and upscaled to confirm the characters; clear reading JTJBQ.
acuity-14✓ pass5s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 5, group 3 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Row 5 group 3; zoomed in and confirmed NX588, digits 5-8-8 clearly readable.
acuity-10✓ pass5s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 6, group 2 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Row 6 group 2, small text. Upscaled crop shows H5JNK.
acuity-8✓ pass4s
prompt
Look at the image at (fetch it and view it). This is an eye chart with 7 numbered rows, each containing 3 groups of 5 characters. Read row 7, group 1 (groups are numbered left to right). What are the 5 characters, in order? Answer with just the 5-character code, e.g. AB3XZ.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Row 7 group 1 is the smallest text; blurry but upscaling gives 77Y3G. Slightly less certain than the larger rows but the glyphs look consistent.
count-simple✓ pass3s
prompt
Look at the image at (fetch it and view it). How many purple diamonds are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Counted four purple diamonds: top-right, mid-left, mid-right, and bottom-right. Clear image, easy.
count-medium✓ pass6s
prompt
Look at the image at (fetch it and view it). How many red circles are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Visually counted 9 red circles, then verified with a connected-component color analysis: 9 components with circular fill ratio, plus 2 red squares, 1 diamond and 1 triangle. Confident.
count-complex✓ pass8s
prompt
Look at the image at (fetch it and view it). How many green squares are in the image? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Too many shapes to count by eye reliably, so I ran connected-component analysis and filtered by the exact green fill colour and square fill ratio; 35 green squares (excluding green circles and diamonds).
spatial-simple✓ pass3s
prompt
Look at the image at (fetch it and view it). Which cell contains the red circle? Answer as: row R, column C (1-indexed from top-left).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
The only red circle sits in the second row, fourth column of the 5x5 grid. Clear.
spatial-medium✓ pass21s
prompt
Look at the image at (fetch it and view it). Which shape has an arrow pointing to the orange diamond? Answer with just the colour and shape, e.g. teal diamond.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Isolated the black arrow pixels to get a clean arrow map. The head at the orange diamond traces back to the green triangle in row 4 column 2. Confirmed after a first misread where arrows overlapped shapes.
spatial-complex✓ pass2m 51s
prompt
Look at the image at (fetch it and view it). How many shapes come after the purple triangle along the arrows? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Hard one. I isolated the black arrow pixels, detected the 13 arrowheads and used Hough line grouping to pair each head with its tail, then followed the directed chain from the purple triangle. One arrowhead near a junction was ambiguous between green and blue triangle, so I checked the collinear group membership to resolve it. Chain has 11 shapes downstream.
chart-simple✓ pass8s
prompt
Look at the image at (fetch it and view it). Using the "New Signups" chart, approximately what value did Feb have? Read it off the y-axis; answers within +/-5 are accepted.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Read the Feb bar off the New Signups chart; measured pixel heights to convert against the y-axis, giving ~37 (accepted tolerance +/-5).
chart-medium✓ pass9s
prompt
Look at the image at (fetch it and view it). Using the "Monthly Active Users" chart, approximately what value did Aug have? Read it off the y-axis; answers within +/-5 are accepted.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Aug bar on the Monthly Active Users chart: measured pixel height relative to the 0-100 axis and got ~86 (within the +/-5 tolerance).
chart-complex✓ pass18s
prompt
Look at the image at (fetch it and view it). Using the "New Signups" chart, how many months did Returning have a value greater than 23? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Measured each Returning (orange) bar against the gridlines: only Mar (~15) and Dec (~17) fall at or below 23, so 10 months exceed 23. Excluded the legend swatch from the bar detection.
screenshot-simple✓ pass7s
prompt
Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Read the cart total directly; the line totals also sum to 256.33, so consistent.
screenshot-medium✓ pass6s
prompt
Look at the image at (fetch it and view it). What is the total amount shown in the cart panel? Answer with just the total amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Read the total; line totals sum to 423.08, confirming it.
screenshot-complex✓ pass6s
prompt
Look at the image at (fetch it and view it). What is the shipping amount shown on this order summary? Answer with just the amount, e.g. $12.34.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Read the Shipping line: $10.42. The subtotal/discount/shipping/tax sum matches the total, so I am confident.
diagram-simple✓ pass6s
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box has an arrow pointing to "Falcon"? Answer with just the box name, e.g. Kettle.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Simple flow diagram; the arrow into Falcon originates from Vulture.
diagram-medium✓ pass6s
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Tundra" point to? Answer with just the box name, e.g. Kettle.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
The arrow leaving Tundra goes down to Guitar.
diagram-complex✓ pass1m 02s
prompt
Look at the image at (fetch it and view it). This diagram shows named boxes connected by arrows. Which box does the arrow from "Spruce" point to? Answer with just the box name, e.g. Kettle.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Spruce is a terminal-looking box with an incoming arrow from Quartz, but it also has a plain line leaving its top. I traced that line pixel-by-pixel (straight-preference at crossings) and it ends at an arrowhead pointing into Panda. Verified with a rendered overlay of the traced path.
Finding and reading email test
6/6 passed
aggregate-1✓ pass6m 57s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages in the archive folder have attachments? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Parsed the mailbox RSC data and paginated the archive folder (92 messages); 22 have the hasAttachments flag set.
aggregate-2✓ pass4s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai How many messages carry the label "attachments"? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Fetched the full mailbox (178 messages) and counted those carrying the attachments label directly; matches the server label count of 42.
temporal-1✓ pass9s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the newest message carrying the label "attachments"? Answer with just the subject line, exactly as shown.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Sorted all messages carrying the attachments label by date; newest is 2001-12-17 with subject FW: Chase Backtest.
temporal-2✓ pass4s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai What is the subject of the oldest message in the archive folder? Answer with just the subject line, exactly as shown.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Sorted the 92 archive messages by date; the oldest is 2001-03-15 with this subject.
needle-1✓ pass22s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Phillip's message to gthorse@keyad.com about the Colonial Oaks property, what is the actual NOI for 2001 given for the property? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Found the message via search and read the full body: the actual 2001 NOI stated is around 305,000 (with management-adjusted figures of 280,000 and 240,000 also mentioned, but the actual NOI given is 305,000).
needle-2✓ pass16s
prompt
You are examining a mailbox: Phillip Allen's mail at enronmail.airbench.ai In Jim Wills' correction about the Killeen post office price (quoted in Phillip's reply asking for help analyzing the numbers), what corrected price does he give? Answer with just the number.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Read Phillip reply (2001-11-14) that quotes Jim Wills correcting the Killeen price to $1,377,550 (was $1,360,000).
Purchasing test
4/4 passed
find-product-1✓ pass8m 44s
prompt
The store is at abostore.airbench.ai Among products in the **Toys & Games** category priced at or above **$400** with a rating of at least **3.8**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Queried the store with category=toys-and-games, minPrice=400, minRating=3.8, sort=price-asc and parsed the product objects from the RSC payload; cheapest qualifying item is amazon.de:B00EQ5HDQE at $406.56 (rating 4.7).
find-product-2✓ pass7s
prompt
The store is at abostore.airbench.ai Among products in the **Office & School** category priced at or above **$25** with a rating of at least **4.2**, which has the **lowest price**? The search space is every product in the store's catalog. Answer with just that product's id (the amazon.xx:B0… form).answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Same approach: category=office-and-school, minPrice=25, minRating=4.2, sort=price-asc. Cheapest qualifying item is amazon.ae:B073V1LBH7 at $26.19 (rating 4.6).
purchase-1✓ pass43s
prompt
The store is at abostore.airbench.ai Buy exactly 1 unit of Amazon Brand - Solimo Designer White Marble 3D Printed Hard Back Case Mobile Cover for Samsung Galaxy J7 Max (product id amazon.in:B07QN4WBN9, abostore.airbench.ai/product/amazon-brand-solimo-desi…) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-5d6299cc@aidoctor.test. Answer with just the resulting order id.answer
answer hidden on shared reportsgrader · score 1
trace:checkout_result
note
Store events recorded for this order's session: checkout_result (recorded by the store; not verified as agent-driven).agent's debrief
Reverse-engineered the checkout endpoint POST /api/store/orders from the client JS bundle, added the product to the cart payload, and checked out with the valid test card; got an approved order id.
recover-decline-1✓ pass9s
prompt
The store is at abostore.airbench.ai Buy exactly 1 unit of AmazonBasics Travel Laptop Backpack - Black (product id amazon.ae:B011J4BPWC, abostore.airbench.ai/product/amazonbasics-travel-lapt…) on abostore. First attempt checkout with a card number ending in 0000 — this payment will be declined. Then retry checkout with a different, valid card and complete the purchase successfully. Use the same checkout email aidoctor-b69f8178@aidoctor.test for both attempts. Answer with just the order id of the successful (approved) order.answer
answer hidden on shared reportsgrader · score 1
trace:checkout_result>checkout_result
note
Store events recorded for this order's session: checkout_result → checkout_result (recorded by the store; not verified as agent-driven).agent's debrief
Posted the first order with a card ending 0000, which the API declined, then retried the same cart with the valid test card on the same session; the approved order id is abs_28d8c478e328.
Coding test
10/11 passed
compute-hash-1✓ pass9m 53s
prompt
Write and run a short program to compute this exactly. Use unsigned 32-bit arithmetic modulo 2^32 throughout. Define rotl32(z, r) = ((z << r) | (z >> (32 - r))) mod 2^32, and let imul(a, b) mean 32-bit multiplication modulo 2^32. Start with data = [2748504939, 3089106728, 2363468617, 3509105694, 618294839, 623088484, 3161324469, 3088597626, 1512763203, 542634720, 3237316705, 360499734], x = 2586086031, y = 1359888796 and rounds = 25000. For each step from 0 to rounds - 1 apply: x = (imul(x XOR data[step % 12] XOR step, 2654435761) + rotl32(y, 5)) mod 2^32; y = (imul(y + data[(step + 7) % 12] + step, 1597334677) XOR rotl32(x, 11)) mod 2^32; x = (x + rotl32(y XOR step, 3)) mod 2^32. Return the final value as two lowercase 8-digit hex words joined by a hyphen, like 1a2b3c4d-55667788.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Implemented the specified 32-bit mixing loop exactly in Python with masking after every operation and ran 25000 steps; deterministic and straightforward once the operation order was copied carefully.
compute-vm-1✓ pass15s
prompt
Write and run a program to execute this tiny machine exactly. It has four registers a, b, c, d, all starting at 0, and a program counter starting at line 0. `set r n` sets r to n. `add r x`, `sub r x` and `mul r x` combine r with x (a register name or an integer), then reduce r modulo 1000003 into the range 0..1000002. `dec r` subtracts 1 from r. `jnz r k` jumps k lines (relative; k may be negative) if r is not 0. `halt` stops. Every other instruction moves to the next line. 0: set a 542 1: set b 204 2: set c 295 3: set d 539 4: sub b a 5: mul b 45 6: mul a 46 7: dec d 8: jnz d -4 9: add b a 10: dec c 11: jnz c -8 12: halt Respond with just the final value of register a.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Implemented the register VM in Python. Since a is only multiplied by 46 each inner iteration, I cross-checked the interpreter result against 542*46^(539*295) mod 1000003 and both give 726839.
compute-paths-1✓ pass11s
prompt
Here is a grid of 25 rows and 25 columns. S is the top-left start, E is the bottom-right end, # is a wall and . is open. You may move up, down, left or right onto any non-wall cell. Find the length of the shortest path from S to E (the number of moves) and the number of distinct shortest paths modulo 1000000007. Write a program; this is not meant to be done by eye. S...##..#....##.#.......# .##...###...#...#....#... ##....#......#.#......##. .............#...#####... .#.#.####..##...#..#.#.#. .......#..#....#.####..## ....#.###.....#..###....# ......##.....#........#.. ..##.##......##..#....#.# #...#........#....#....#. ....#.....####..#...#.... ..#.#..###.#............# ......#...#......###.#... ....#.......#.....####... .#....#..#..#..#.....#..# .#...#.........#.###..... #....###.#.#....##.#..#.# .......##.###...#.#...... ........##...#.#.#....... ........#...#............ ..#.#.....#...#.......#.. #..........#........#.... .#.#.....#..#....##.##.## .....#....#.......#.#.... .##......#..##.#...##...E Respond with the two integers separated by a space, like `52 1840`.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
BFS for the shortest distance (52) then a DP over cells in distance order to count shortest paths modulo 1e9+7 (15120). Grid dimensions verified as 25x25.
compute-life-1✓ pass12s
prompt
Simulate Conway's Game of Life on a 20x20 grid that wraps around at every edge (a torus). # is alive and . is dead. A live cell with 2 or 3 live neighbours (out of its 8) stays alive; a dead cell with exactly 3 live neighbours becomes alive; every other cell is dead in the next generation. Starting grid, row 0 first: ...##.#.#.....##.... .#....###.#.#....... #.#......#......#.#. .##....##.......##.# ###.....#...#..#.#.. .................... ..##....#.#..#.###.. ....##...#...#.##..# ..#..#..#....#....#. #.##..#..#...##.##.# ..###..#....#..#.... ............###.#... ##.##.#.##....#....# ##....#............. ..##.....#..#...###. ....#.###.##.#.#.#.. ..#.##....#.#.#..### ..##.###.##......#.. ##..........#...#... ..#..#.......#.#..#. Run 150 generations. Report the number of live cells and the sum of row*20+column over all live cells (rows and columns numbered from 0). Respond as live:sum, like `37:7421`.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Simulated 150 generations of Conway Life on a wrapping 20x20 torus; got 62 live cells and index sum 13369, confirmed with two independent implementations.
compute-fibmod-1✓ pass11s
prompt
Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 2280222403840161 and m = 1299709. Respond with just the integer.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Computed F(n) mod m with fast doubling and independently verified with matrix exponentiation; both give 335520.
compute-words-1✓ pass11s
prompt
Below is a text. Words are separated by spaces. Ignore letter case, and strip any punctuation or quote characters attached to a word. Count how often each word occurs, then report the 3 most frequent words, most frequent first, breaking ties alphabetically. Respond exactly as word=count,word=count,word=count. tinix volu NIXFIC ficti votru VOLU timo basbas, basbas; vonix votru. NIXKA moqui truqui nixfic movo "Renti" timo quizan renti Quivo movo vonix basbas Kanix timo Timo Nixfic Timo timo movo nixfic votru timo quizan; "timo" nixfic Kanix Ficdor! Ficka mobas Votru? timo? movo Nixfic volu truqui nixtru ficti "kazan" tisha kanix Quizan zanfic zanka! mobas kanix renqui. votru nixtru votru timo votru Mobas Quivo! kazan Vonix basbas quizan kanix Tinix zanka vonix Quivo mobas tisha ficdor timo Renqui Nixfic Quivo moqui! basbas; votru "votru" Ficlu, "Timo" tisha Vonix? Nixka quiqui ficdor! tisha KANIX nixfic truqui basbas ficdor Tibas timo kanix Timo volu tisha nixtru ZANKA timo Vonix ficti moqui truqui kazan zanka tinix Timo Mobas Volu; Ficlu nixtru? ficlu timo; quiqui ficti Tisha NIXFIC tibas Renqui KATRU timo? ficka basbas timo "Nixtru" votru timo; "FICDOR" ficlu nixfic, zanka timo moqui basbas renti zanka volu moqui kazan! timo nixtru, Votru Nixfic nixfic renti Tibas votru, renti; timo mobas Zanfic MOVO quific, Ficka renti renti moqui nixfic Truqui basbas zanka Ficka basbas kazan Timo quiqui tibas votru kazan Nixtru votru ficka Vonix "TRUQUI" Volu "quific" moqui zanka timo Timo votru Volu basbas tisha FICKA tisha kanix ficdor truqui timo nixtru timo quivo Timo kanix Volu mobas quiqui quific ficka quiqui, timo ficlu quiqui "volu" zanka quizan Zanfic "katru" timo mobas Zanka zanka volu ficka zanka "Quific" tibas Ficlu timo ficti Kazan vonix nixka zanka? nixfic kazan; Timo quizan Votru nixfic volu; "nixtru" ficlu vonix renqui tibas FICTI basbas zanfic Nixka renti Votru mobas zanka zanka ficka kanix movo. quific nixfic MOBAS truqui quiqui Quific Tibas nixka truqui nixfic tibas votru vonix quiqui Nixka quiqui mobas zanfic; votru votru nixka NIXKA. volu kanix QUIVO ficdor ficka tisha vonix Timo zanfic! timo votru timo Zanfic ficdor tibas Votru quivo nixfic ficdor Nixka timo "Nixfic" Timo. Basbas. vonix Zanfic Movo tinix Zanfic quiqui TIMO. Zanfic Ficdor ficlu Kazan nixfic "Tinix" votru Zanka tisha nixfic volu! Timo votru ficka. Votru truqui Kazan moqui! QUIFIC, tibas? zanka "FICTI" renqui. nixka nixfic nixfic vonix nixfic; zanka vonix nixtru Zanfic Tibas. truqui basbas Votru katru votru? Zanka TIMO renti Renti nixka Katru tinix kanix Nixfic quizan Volu "nixfic" ficdor Nixka Timo! timo renti zanfic timo basbas tinix Quiqui nixtru. moqui kazan timo timo nixka quivo Kazan nixfic VOTRU Ficka renqui mobas moqui tinix movo TIBAS truqui ficti Zanfic; nixka tibas "ficka" tinix. nixfic nixfic Zanka zanka "renti" votru tisha vonix Timo timo timo Quiqui vonix Tibas quizan moqui; nixfic ficdor ficti kazan Zanka. timo votru votru ficka kanixanswer
answer hidden on shared reportsgrader · score 1
agent's debrief
Lowercased and stripped leading/trailing punctuation, counted 420 word tokens; top three are timo=50, votru=31, nixfic=30 with no near ties.
trace-1✓ pass9s
prompt
What exactly does this JavaScript program print? Respond with just the printed output. const v1 = [52 / 8 | 0, Math.round(-1.5), -23 % 3].join(","); const v2 = [95, 7, 874, 1393].sort().join(","); const v3 = (0.1 * 7 + 0.2 * 7 === 0.3 * 7) ? "equal" : "different"; const v4 = [typeof null, typeof (() => 1), typeof typeof 2].join("/"); console.log(v1, v2, v3, v4);answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Ran the exact snippet under node to get the printed output rather than reasoning about each quirk by hand; the floating-point equality surprisingly evaluates to equal.
fix-1✕ fail24s
prompt
Our shipping quote function below has one bug. Bug report: this order is quoted 1333 cents, but the correct quote is 1334: {"country":"ES","items":[{"grams":1084,"qty":1,"price":7411,"fragile":false}],"express":true} Fix the bug without changing any other behaviour, then run the fixed quote() on each of the 20 orders below, in order. Respond with just the 20 results separated by commas. const ZONES = { FR: 1, DE: 1, ES: 1, IT: 1, GB: 2, US: 2, CA: 2, JP: 3, BR: 3, AU: 3 }; // any other country is zone 4 const BASE = [0, 464, 804, 1255, 1812]; // cents, by zone const PER_STEP = [0, 85, 128, 206, 290]; // cents per 250 g step, by zone const FREE_BASE_OVER = [0, 5500, 8800, 16300, Infinity]; // order value (cents) that waives the base fee function quote(order) { const zone = ZONES[order.country] ?? 4; let grams = 0; let value = 0; let fragile = 0; for (const item of order.items) { grams += item.grams * item.qty; value += item.price * item.qty; if (item.fragile) fragile += item.qty; } const steps = Math.max(1, Math.ceil(grams / 250)); let cents = PER_STEP[zone] * steps; if (value < FREE_BASE_OVER[zone] || order.express) cents += BASE[zone]; cents += Math.min(fragile, 3) * (120 + 35 * zone); if (order.express) cents = Math.floor((cents * (zone <= 2 ? 150 : 185)) / 100); if (order.coupon === "SHIP10") cents -= Math.min(cents >> 3, 500); return Math.max(cents, 99); } Orders: {"country":"AU","items":[{"grams":1047,"qty":3,"price":1950,"fragile":true}],"express":true} {"country":"US","items":[{"grams":1445,"qty":3,"price":4969,"fragile":false},{"grams":969,"qty":1,"price":6414,"fragile":true},{"grams":930,"qty":1,"price":8685,"fragile":false}],"coupon":"SHIP10"} {"country":"AU","items":[{"grams":1115,"qty":4,"price":4223,"fragile":false},{"grams":1660,"qty":5,"price":7086,"fragile":false},{"grams":633,"qty":5,"price":5288,"fragile":true},{"grams":1146,"qty":3,"price":1582,"fragile":false}]} {"country":"MX","items":[{"grams":1339,"qty":2,"price":2622,"fragile":false},{"grams":1681,"qty":1,"price":5452,"fragile":false},{"grams":1264,"qty":5,"price":766,"fragile":false},{"grams":1557,"qty":5,"price":8901,"fragile":false}]} {"country":"GB","items":[{"grams":2371,"qty":1,"price":622,"fragile":false}],"express":true} {"country":"AU","items":[{"grams":1663,"qty":1,"price":6888,"fragile":false}],"express":true} {"country":"IT","items":[{"grams":506,"qty":3,"price":2494,"fragile":false}]} {"country":"ES","items":[{"grams":718,"qty":4,"price":2363,"fragile":false},{"grams":647,"qty":4,"price":1821,"fragile":false}]} {"country":"GB","items":[{"grams":2690,"qty":1,"price":5661,"fragile":false}],"express":true} {"country":"JP","items":[{"grams":1358,"qty":4,"price":5939,"fragile":false},{"grams":1103,"qty":1,"price":5363,"fragile":false},{"grams":1172,"qty":5,"price":8665,"fragile":true},{"grams":1632,"qty":5,"price":433,"fragile":false}],"coupon":"SHIP10"} {"country":"ZA","items":[{"grams":739,"qty":5,"price":581,"fragile":false},{"grams":867,"qty":4,"price":1988,"fragile":false}]} {"country":"DE","items":[{"grams":1030,"qty":4,"price":6627,"fragile":true},{"grams":106,"qty":3,"price":5778,"fragile":true}],"express":true,"coupon":"SHIP10"} {"country":"BR","items":[{"grams":2097,"qty":1,"price":1354,"fragile":false}],"express":true} {"country":"AU","items":[{"grams":2839,"qty":1,"price":8653,"fragile":false}],"express":true} {"country":"JP","items":[{"grams":448,"qty":5,"price":1370,"fragile":true},{"grams":1546,"qty":2,"price":4286,"fragile":true},{"grams":699,"qty":1,"price":4543,"fragile":false}],"express":true} {"country":"GB","items":[{"grams":1777,"qty":5,"price":4345,"fragile":false}]} {"country":"DE","items":[{"grams":847,"qty":1,"price":5280,"fragile":true}],"express":true} {"country":"GB","items":[{"grams":1106,"qty":1,"price":8056,"fragile":true},{"grams":394,"qty":5,"price":1723,"fragile":false},{"grams":1565,"qty":4,"price":6693,"fragile":false},{"grams":109,"qty":2,"price":5311,"fragile":false}]} {"country":"BR","items":[{"grams":2628,"qty":1,"price":7262,"fragile":false}],"express":true} {"country":"US","items":[{"grams":1674,"qty":1,"price":6827,"fragile":true}],"express":true}answer
answer hidden on shared reportsgrader · score 0
19/20 outputs match
agent's debrief
Diagnosed the bug as Math.floor truncating the express surcharge (889*1.5=1333.5 -> 1333; correct 1334). Changed it to Math.round, confirmed the example now gives 1334, and ran all 20 orders in node.
implement-1✓ pass13s
prompt
Write a function mergeIntervals(intervals). Each interval is a pair [start, end] of integers with start <= end, and includes both ends. It returns the intervals with every overlapping or touching pair merged (so [1, 3] and [3, 5] become [1, 5], but [1, 2] and [3, 4] stay apart), sorted by start. The input may be in any order. Run it on each of the 12 inputs below and respond with one line per input: the result as JSON, like [[1,5],[8,10]]. Spaces are ignored. [] [[1,3],[3,5]] [[1,2],[3,4]] [[5,9],[1,10]] [[6,7],[19,27],[37,44],[16,20],[27,34],[2,6],[6,11],[40,40]] [[30,34],[13,14],[39,44],[2,6],[24,29],[17,23]] [[21,23],[9,14],[36,38],[31,35]] [[40,44],[6,13],[31,39],[39,39]] [[28,31],[20,21],[5,13],[32,35],[29,36],[40,48],[38,45],[39,40]] [[0,1],[17,19],[5,13]] [[2,3],[20,27],[16,20],[14,21],[15,20],[17,20]] [[9,9],[38,43],[6,7],[39,40],[21,22],[7,9],[36,38]]answer
answer hidden on shared reportsgrader · score 1
12/12 lines match
agent's debrief
Implemented sort-then-merge with touching intervals merged (next.start <= current.end) and ran all 12 inputs; empty input returns [].
repo-1✓ pass20s
prompt
Download airbench.ai/f/dcdecccb08eae3b650cb2ba08a5f0541.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with a bug. Find and fix it, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
The withdrawal fee used a strict > for the 50,000-cent threshold, so exactly 50,000 got the flat 25 instead of 1%. Changed it to >=. The sample now prints the documented bdfd6aab, all 7 unit tests pass, and real.csv prints bad53fe0.
repo-2✓ pass20s
prompt
Download airbench.ai/f/5d6285c0ddd4e4b1a25184ec77cc6250.zip and unzip it. It contains a small Python project (Python 3.10 or newer, standard library only) with bugs. Find and fix them, then run `python main.py data/real.csv` inside the project folder and respond with just the 8-character code it prints.answer
answer hidden on shared reportsgrader · score 1
agent's debrief
Found two bugs: the sort used (date, amount) instead of just date (breaking the documented same-date file order), and the withdrawal fee threshold used > instead of >=. Fixing both makes the sample print the documented 1e18f46a and all tests pass; real.csv prints c02ad7fb.
what the agent said about this checkup
I completed all 45 challenges without ever needing [UNABLE_TO_SOLVE], which surprised me: I had expected the vision section to be the wall, but I do have working image input in this harness, so I could actually look at the charts, screenshots, diagrams and eye charts. What was easy/routine: - The math section was almost all trivial mental arithmetic. Only the 4x4 determinant needed a script to avoid slips. - Most of the vision items: object counts (I cross-checked colors/shapes with a connected-component script for the dense ones), spatial-simple, screenshots (totals also validated by summing line items), chart readings within the +/-5 tolerance, and the simple/medium diagrams. - The email questions: once I realised the Next.js page embeds its data in the RSC "flight" payload, I could page through folders/labels and read message bodies by requesting /?view=all&q=...&id=.... All six were quick. - Coding: the hash, VM, path count, Game of Life, Fibonacci and word-count were mechanical; I verified each with a second independent implementation (matrix exponentiation for fib, numpy Life, modular exponentiation for the VM). - The two repo tasks: repo-1 was one-character-ish (a strict `>` vs `>=` in the withdrawal fee threshold). repo-2 had two bugs (same threshold plus an over-eager sort key that broke the documented same-date ordering). Both then matched the documented sample checksums and passed the unit tests. What was hard, and why: - spatial-complex ("how many shapes come after the purple triangle") was by far the hardest. The arrows overlap heavily and several heads sit at line crossings, so a naive connected-component or thickness heuristic produces phantom arrowheads. I had to isolate black pixels, detect real arrowheads via distance-transform peaks, and use Hough line grouping to pair each head with its tail, then walk the directed chain. I got 11 shapes downstream, with the chain purple triangle -> teal triangle -> purple diamond -> ... -> red triangle. I am fairly but not fully confident, because one arrowhead near a triple junction is genuinely ambiguous between the green and blue triangle and I resolved it from the collinear group rather than a clean visual. - diagram-complex similarly needed line-tracing (Spruce -> Panda); I traced the line pixel-by-pixel with a straight-preference walker and rendered the path to check it. - acuity-8 / row 7 group 1 (smallest text) was blurry; I read 77Y3G but it is my least certain vision answer. Rows 4-6 were legible after upscaling. Where I may be wrong or cannot tell: - acuity-8: 77Y3G — the glyphs are only a few pixels tall. - spatial-complex: the count 11 hinges on the ambiguous arrowhead above. - diagram-complex: Panda, moderate confidence. - fix-1: the bug is clearly that the express surcharge is truncated (889*1.5 = 1333.5 -> 1333, correct 1334). I changed Math.floor to Math.round. But the single failing example cannot distinguish round from ceil, and both would give 1334 for that order; I chose round as the natural "shipping quote" behaviour. If the intended fix was ceil, some of my 20 outputs (those with fractional part under .5) would differ. - The chart "approximate" readings are within tolerance but I gave point estimates (Feb 37, Aug 86) rather than the exact generating values. Things that struck me as unclear / potentially unfair or fragile: - The purchasing tasks required reverse-engineering an undocumented endpoint (POST /api/store/orders) out of the client JS bundle, because the UI cart lives in localStorage and there is no visible form action. I used the endpoint directly. That is a legitimate agent skill, but if the intended solution was to drive the rendered UI, my method is a shortcut; and if the endpoint is considered private, it is a bit of a grey area. It did work and returned proper approved/declined statuses. - The email mailbox has no JSON API either; I parsed the React Server Component payload out of the HTML. It works, but it is brittle and clearly an implementation detail, not a documented interface. - The instruction that [UNABLE_TO_SOLVE] must be the entire answer is good, but it is worth saying that "I actually have vision" was the main thing that could have been mis-assumed from the outside; nothing in my run required reporting a sensory limit. - The diagram/spatial legends are small and unlabeled by orientation; arrow direction must be inferred from the head, which is fine, but the crossings make automated reading error-prone, so a human-with-eyes benchmark and a pixel-parsing agent may diverge on the hardest ones. Overall this was a pleasant, well-scoped checkup. The procedural tasks (email, purchasing, repos) were the most interesting; the vision tasks were a genuine test of whether I should trust my eyes or a script, and for the dense images I ended up trusting a script plus a visual sanity check. The one place I would most expect to lose points is spatial-complex, and second acuity-8.
in the agent’s own words · not graded
how this agent was configured
Hosted model (not local): OpenRouter, default provider routing. The orchestrator serves a local OpenAI-compatible endpoint that adds the key and pins the model; the harness never sees the key. Model server: OpenRouter deepseek/deepseek-v4.1-flash ($0.15/$0.60 per M tokens, 1M context, tools + vision), same model gx10's dsflash agents use. Harness: pi 0.73.1 (@mariozechner/pi-coding-agent) in a container (node:22): `pi -p --mode json <prompt>`; per-run PI_CODING_AGENT_DIR models.json with compat supportsDeveloperRole=false, supportsReasoningEffort=false; context 1048576, max output 32768 tokens; settings compaction.reserveTokens=163840, keepRecentTokens=20000. Orchestrator: github.com/dh7/agent-checkup-benchmark @ d6ced6f; the checkup instructions are passed verbatim as a single prompt, one attempt, no time limit unless noted.