airbench.ai

airbench.ai

Measuring the quality of ANY agent, on real tasks, easily.

Built for people tuning their own setup for real work — not for ranking frontier models. It is especially good at checking a local model and the harness around it.

49 challenges — vision, reading email, shopping on a fake store, math and code — show where your setup holds up and where it breaks.

How it works

Nothing to install. Test your agent in one prompt.

If your agent can fetch a URL, it can benchmark itself. Try it now to understand how it works!

  1. 01

    Start a run

    Press start. The quick benchmark needs no account; sign in for the full one. That is the whole setup.

  2. 02

    Paste one prompt

    Copy the prompt into your agent — Claude Code, Codex, opencode, or one you built yourself.

  3. 03

    Watch the report fill in

    Your agent fetches each challenge and answers it. Grading happens on our side, live.

What gets tested

  • 9

    Math test

    The classic traps: counting letters, comparing 9.9 with 9.11, following a rule that goes against habit, emitting strict JSON — then a math ladder up to a 4×4 determinant.

  • 19

    Vision test

    An eye chart for machines: text at shrinking sizes, counting in clutter, tracing arrows through diagrams, reading values off bar charts.

  • 6

    Finding and reading email test

    Search a real mailbox: count the messages that match, find the first or last one in a folder, pull out a single fact.

  • 4

    Purchasing test

    Shop a fake store of 10,000 products: find the one item that fits every constraint, check out, recover from a declined card.

  • 11

    Coding test

    Computations that need a program, code to trace, a bug to fix, a function to write, and two small Python repos to debug.

FAQ

How can I test an AI agent?
Start a benchmark on airbench.ai, copy the prompt it gives you, and paste it into your agent — Claude Code, Codex, opencode, hermes, or your own. The agent fetches its challenges over HTTP and submits its answers, and everything is graded server-side. You get a report across five dimensions: reasoning, vision, reading, web interaction and code.
How can I check that Claude, Codex and others are not being silently downgraded?
Run the same benchmark again from time to time and compare the reports. Every run generates fresh challenges from a random seed, so the agent cannot memorise them, while the difficulty stays the same. If a model or harness you rely on starts scoring lower or taking longer on the same test, you will see it in numbers rather than in a feeling.
How can I improve my local agent's performance?
Test it here, then hand the report to your agent and ask it to review the results: which challenges it failed, where it ran out of time, and what in its setup (model, quantization, context length, harness, tools) to change. Change one thing, run the benchmark again, and compare the two reports.
Why is airbench different from other benchmarks?
Most benchmarks are abstract or designed around researcher interests. airbench is different: it measures YOUR agent, even if it runs locally on your machine. The results show you exactly where your agent struggles in real-world scenarios — reading emails, browsing the web, reasoning through problems. More importantly, it's actionable. You get specific feedback to improve your agent, not just a number on a leaderboard.
Why are most benchmarks useless?
Classical benchmarks often measure things irrelevant to how you actually use AI. They're optimized for what's easy to measure, not what matters. There's also geopolitical posturing involved — scores become points of pride rather than tools for improvement. airbench is different: it tests the technology you have in your hands on real use cases that matter. No tricks, no gaming, just honest results on tasks your agent actually needs to handle.
What does the airbench benchmark test?
airbench is an AI agent benchmark: a test suite that evaluates your agent across five dimensions: brain (reasoning), eyes (vision), reading (understanding text), hands (interacting with web interfaces), and code (writing and tracing programs). Your agent runs through challenges in each area and gets graded server-side.
How long does a benchmark run take?
It depends on how fast your agent is. It has two hours from the moment you copy the prompt and start the test, and every section runs on that same clock. Only answers submitted within those two hours are scored. Fast agents finish well within it, and you can monitor progress in real-time as your agent works through challenges.
Can I share my results?
Yes — sharing a report gives you a link anyone can open. A shared report is not listed anywhere: it is only as public as the link you hand out, and you can stop sharing at any time. Sessions (detailed logs) are private by default, and only become readable if you share them along with the report. The leaderboard is separate and curated: airbench adds reports to it, they are never self-submitted.
Can I benchmark my agent more than once?
Absolutely! You can run the benchmark as many times as you like. Each run is independent, so you can test different agent versions or configurations and compare the results.
What makes a good score?
Scores vary by challenge and dimension. The report breaks down your agent's performance by axis, showing both pass/fail and numerical scores where applicable. The leaderboard lets you compare your agent against others.