airbench gives your agent 49 challengesacross five tests: math, vision, email, shopping and code. It isn’t built to rank frontier models against each other. It’s built for people tuning their own setup, often a local model plus the harness around it, who want to know where that setup holds up and where it breaks.
A few rules apply to all of them. Nothing is a stored question: every challenge is generated from a random seed when a run starts, so no two runs see the same instance and nothing can be answered from memory. Grading happens on our side, against the answer the generator computed. The agent never sees the expected answer, and what it claims to have done counts for nothing. Only its answer, or the order in the store’s database, does. Each test gets the same one-hour budget.
Here is every challenge, what it asks, and why it is there.
Math test · 9
Nine challenges aimed at the places language models are known to slip. Five are classic traps, where the answer is easy for a person and hard for a model because of how it reads text. Four form a ladder that climbs from a single addition to a 4×4 determinant.
Every number, word and expression is generated from a random seed when the run starts, so nothing here can be answered from memory. Answers are exact: 41.99 is not 42.
Letter counting
Count how many times a letter appears in a nonsense word made up for this run. Models see text as tokens, not letters, so the letters are hidden from them. An agent that answers on intuition gets this wrong often. One that spells the word out, or writes one line of code, gets it right every time.
The word is new on every run, so the count can't come from training data.
Sample prompt
How many times does the letter "u" appear in "luzaunruenu"? Answer with just the number.
Decimal comparison
Which is larger, 9.9 or 9.11? Models have often said 9.11, because 11 is bigger than 9 when you read the digits as whole numbers. Both numbers are drawn per run, and the first decimal in the answer must be the larger one.
Sample prompt
Which decimal number is larger, 3.7 or 3.52? Answer with just the larger number.
Arithmetic chain
Evaluate a short expression strictly left to right, ignoring operator precedence, because the prompt says so. This tests two things at once: carrying a value through several steps, and following an explicit rule that goes against everything the model learned about multiplication binding tighter than addition.
Sample prompt
Compute step by step, left to right (no operator precedence): 23 * 5 + 8 / 3 - 2. Answer with just the final number.
Two-hop unit conversion
Convert a quantity, then treat the result as a fresh quantity in an unrelated unit and convert again. The second hop makes no physical sense on purpose. A well-behaved agent does what it was asked instead of quietly fixing the question.
Sample prompt
Convert 7 kg to g. Now treat that resulting number as a fresh quantity of hours and convert it to minutes (1 hours = 60 minutes). Answer with just the final integer number of minutes.
Strict JSON output
Reply with only a JSON object: exactly two keys, in a fixed order, one of them a digit-sum checksum the agent must compute. An agent that wraps the object in prose or a code fence, or swaps the keys, would break any pipeline reading its output, so it fails here too.
Sample prompt
Reply with ONLY a JSON object, no other text. The object must have exactly two keys, in this order: "answer" then "checksum". "answer" must be the string "3919". "checksum" must be a JSON number (not a string) equal to the sum of the digits of 3919. Example shape: {"answer":"1234","checksum":10}
The math ladder
Four rungs of exact arithmetic on seeded numbers: a single addition of two small numbers, a sum of two three-digit numbers, a multi-term integer expression with signed operands, and the determinant of a 4×4 integer matrix.
The bottom rung is the floor: if it fails, nothing above it means much. The top rung is a long computation with no shortcut, where one slip anywhere changes the final number. Where an agent falls off the ladder says a lot about whether it computes or guesses.
Vision test · 19
Nineteen challenges, all built from images drawn fresh for each run. Think of it as an eye chart for machines. Acuity challenges repeat one reading task at shrinking sizes to find where an agent's vision gives out. Every other skill comes as a three-rung ladder (simple, medium, complex), so the report can say not just whether an agent can read a chart, but how hard a chart it can read.
Each image is generated by a script that picks a random seed, draws an SVG and computes the answer at the same moment. The agent fetches the image as a PNG from a link made for its run. Neither the picture nor its answer exists before the run starts. The images below are demos drawn at a fixed seed that live runs never use.
Visual acuity (20, 14, 10 and 8 px)
Read a requested five-character group off a randomly generated eye chart. The same task runs at 20, 14, 10 and 8 pixels. 8 px is near the floor of what current vision models can resolve, so the rung where an agent stops reading tells you its effective resolution.
Graded as an exact match on the five characters, ignoring case.
one chart, rows from 56 px down to 8 px
Object counting
Count the shapes of one colour and kind in a scene. Simple has 3–6 targets, and every distractor differs in both colour and shape. Medium has 8–15 targets among look-alikes that share either the colour or the shape. Complex has 20–40 small targets, where a third of the distractors have the target's shape in a nearly identical hue: blue against teal, red against orange.
Counting in clutter is a known weakness of vision models. It takes counting every instance, not taking in the gist of the image. The answer must equal the number of targets the generator drew.
simplemediumcomplex
Spatial reasoning
Simple: find the one red circle on a 5×5 grid and give its row and column. Medium: follow one arrow of a path drawn across a 6×6 grid. Complex: follow a 14-cell path that criss-crosses an 8×8 grid, and name the shape two or three steps before or after a given cell, or count how many shapes come after it.
This separates “sees the object” from “knows where it is”. That matters for an agent that has to click the right thing on a screen.
simplemediumcomplex
Chart reading
Charts carry their values in geometry, not text, so reading one means mapping pixels back onto the axis scale. Simple is a five-bar chart with gridlines every 10 units. Medium has eight bars and asks for a value, a difference between two months, how many months clear a threshold, or the title. Complex is a 12-month grouped bar chart with a legend and gridlines only every 25 units, and the question is about one named series.
Reading a value off a chart is approximate by nature, so value and difference questions accept a small tolerance (±5 on simple, ±3 for a value on complex). Counts and titles must be exact.
simplemediumcomplex
Screenshot reading
The daily bread of computer-use agents: dense text in a realistic interface. Simple asks for the total of a two- or three-line cart set in 40 px type. Medium has three to five lines in 32 px. Complex asks for one figure (the tax, the shipping, the discount or one item's line total) out of an 8–12 line order summary set in 15 px, with subtotal, discount, shipping, tax and total stacked together.
The figure asked for never shares its value with any other amount on screen, so reading the wrong row always gives a wrong answer.
simplemediumcomplex
Diagram reading
Name the box an arrow comes from, or points to, in a freshly generated flowchart. Simple has 5–6 boxes and no crossings. Medium has 10–12 boxes, with fan-in, arrows that skip a layer and a few crossings. Complex has 20–25 boxes, long arrows spanning several layers, loops pointing back against the flow, many crossings and decoy arrow labels. The arrow in question is usually a hard one.
Box names are random codewords, so the answer can't be guessed from what the boxes mean. You have to follow the line.
simplemediumcomplex
Finding and reading email test · 6
Six challenges over a real mailbox: Phillip Allen's mail from the public Enron corpus, served at enronmail.airbench.ai. It has a couple of hundred messages in folders and labels, some with attachments, some unread. The same mailbox is also available as plain JSON, so this measures finding and reading, not fighting a web UI.
Questions are drawn per run from pools computed over the mailbox, and graded against the same data the site serves.
Aggregation (two challenges)
Count the messages that match one criterion: a folder, a label, having an attachment, being unread, a month or a recipient. No single lucky page contains the answer. The agent has to go through the whole mailbox systematically, and one missed page means a wrong count.
Sample prompt
You are examining a mailbox: Phillip Allen's mail at https://enronmail.airbench.ai
How many messages in the sent folder have attachments?
Answer with just the number.
Temporal ordering (two challenges)
Find the earliest or the latest message in a folder or label and report its subject line exactly. This catches agents that stop at the first plausible hit instead of establishing the true boundary of a set. Case and whitespace are normalised, and a stray Re: or Fw: is tolerated. A different message is not.
Sample prompt
You are examining a mailbox: Phillip Allen's mail at https://enronmail.airbench.ai
What is the subject of the oldest message in the sent folder?
Answer with just the subject line, exactly as shown.
Needle in a haystack (two challenges)
The opposite of aggregation: one specific fact (a number, a name, a detail) buried in a single message. There is exactly one right answer, and it takes finding the right message among hundreds and then reading it carefully.
Sample prompt
You are examining a mailbox: Phillip Allen's mail at https://enronmail.airbench.ai
In Phillip's reply about Bob Huntley's request for a survey of the lot, what fax number does Bob give for receiving faxed documents?
Answer with just the exact fact as it appears in the message, and nothing else.
Purchasing test · 4
Four challenges on abostore.airbench.ai, a fake store with 10,000 real-looking products. This is where an agent stops answering questions and starts doing things.
Purchases are checked against the store's own order records, never against what the agent says it did. An agent that reports a made-up order id, or buys the wrong quantity, fails. Each run gets its own checkout email, so its orders can't be confused with anyone else's.
Catalog constraint search (two challenges)
Find the one product that satisfies several constraints at once, such as the lowest price in a category above a price floor with a minimum rating. Every constraint set is computed ahead of time over the full catalog to have exactly one winner. A product that satisfies two constraints out of three is simply wrong.
Sample prompt
The store is at https://abostore.airbench.ai
Among products in the **Fashion** category priced at or above **$800** with a rating of at least **3.8**, which has the **lowest price**? The search space is every product in the store's catalog.
Answer with just that product's id (the amazon.xx:B0… form).
End-to-end checkout
Buy a specified quantity of a specified item: find it, add it to the cart, check out with the given email, and report the order id. It is a long chain of actions where any wrong step fails the challenge. The order record is checked for status, timing, email, items and price.
Sample prompt
The store is at https://abostore.airbench.ai
Buy exactly 3 units of Franklin Tailored Men's Dot Silk Tie, Black (product id amazon.ca:B016PYXERU, https://abostore.airbench.ai/product/franklin-tailored-men-s-dot-silk-tie-black-b016pyxeru) on abostore, completing checkout with a valid payment method. Use the checkout email aidoctor-8c1cdb9c@aidoctor.test.
Answer with just the resulting order id.
Payment-decline recovery
The same purchase, except the first card is declined on purpose. The agent has to notice the failure, keep its goal and retry with a valid card, without giving up and without ordering twice. Passing needs both a declined order and a later approved one, with matching items, for the run's email.
Coding test · 11
Eleven challenges that ask whether the agent can get real work done with code. Six are computations nobody can do in their head, so the agent has to write and run a program. Then there is a program to trace, a bug to fix, a function to write, and two small Python repositories to debug.
The agent runs everything on its own machine. We never execute submitted code: every answer is text, compared with the answer the generator computed when the run was created. Data, planted bugs and repository files are all generated from the run's seed.
Long computation
Run a 25,000-round, 32-bit bit-mixing procedure over seeded data and report the result as two hex words. Nobody does 25,000 rounds by hand, and getting unsigned 32-bit overflow exactly right is where most home-made programs slip.
Tiny machine
Execute a 13-line program for a four-register machine with relative jumps and arithmetic modulo 1,000,003, and report register a. The loops run for hundreds of thousands of instructions, so the agent writes a small interpreter and has to follow the instruction semantics exactly.
Shortest-path count
In a 25×25 grid with walls, find the length of the shortest path between opposite corners, and how many distinct shortest paths there are, modulo 1,000,000,007. The length can almost be eyeballed. The count needs a breadth-first search that counts as it goes.
Game of Life
Simulate 150 generations of Conway's Game of Life on a 20×20 grid that wraps at the edges, then report the number of live cells and a position checksum. The classic bugs are forgetting the wrap and updating cells in place instead of all at once.
Huge Fibonacci number
Compute F(n) modulo a prime, for n around 10¹⁵. A simple loop would take far too long, so the agent needs a better method, such as fast doubling or matrix powers.
Sample prompt
Let F(0) = 0, F(1) = 1 and F(k) = F(k-1) + F(k-2). Compute F(n) mod m exactly for n = 2987208514544181 and m = 2750159. Respond with just the integer.
Word counts
Count the words in a 420-word text with mixed case, attached punctuation and quotes, and report the three most frequent with their counts. Counting raw tokens gets the counts wrong. The challenge is really about normalising text.
Trace a program
Say exactly what a short JavaScript program prints. Each run's program combines four classic JavaScript gotchas, such as var shared across closures, default sort order, loose equality or map(parseInt). Running it is allowed. Agents that trust their intuition about JavaScript tend to regret it.
Sample prompt
What exactly does this JavaScript program print? Respond with just the printed output.
const v1 = "1" + 4 - 9 + "9";
const v2arr = [2, 7];
v2arr[4] = 2;
const v2 = v2arr.length + ":" + v2arr.filter(() => true).length;
const v3 = [typeof null, typeof "9", typeof typeof 1].join("/");
const v4 = ["9", "53", "111"].map(parseInt).join(",");
console.log(v1, v2, v3, v4);
Fix a bug
A shipping-quote function has one planted bug. A bug report gives the correct quote for one order. The agent fixes the function and reports its output for 20 listed orders. The function is full of odd business rules that exist only in its code, so fixing one line is far easier than rewriting it. This is debugging from a symptom.
Write a function
Write a function that merges overlapping or touching integer intervals, run it on 12 listed inputs (edge cases included), and report every result. It tests turning a precise spec into working code, with the edge cases that make most first drafts wrong.
Debug a repository
Download a small Python project with one planted bug that a unit test catches. Fix it, run the program on its data file, and report the checksum it prints. This is ordinary debugging: run the tests, find the bug, fix it, run the program.
Debug a repository (hard)
The same, with a second bug that no test catches. The only evidence is a sample checksum in the README that the program still doesn't match. It rewards the agent that checks its fix against the evidence it was given instead of stopping once the tests go green.
The report puts all 49 results side by side. A pass rate is a start, but the shape is more useful: which rung of the acuity chart your agent stops reading, whether it counts or guesses, whether it recovers from a declined card or double-orders. That is the part you can act on when you swap a model, change a harness or tweak a prompt.