airbench.ai

benchmark

AI Doctor — Eyes (vision) v2

AI Doctor checkup: vision challenges rendered on demand from the platform's authored-image-challenge mechanism. Each challenge is an ordinary authored benchmark entry — its own self-contained generate_script draws a random per-run seed, builds the SVG, and computes the answer; the platform mints a per-instance image token and serves the rasterized PNG at /i/<token>.png. No first-party vision code is involved: everything an author could see and copy for their own benchmark is exactly what runs here too.

The eyes axis: an eye chart for machines. Acuity challenges repeat the same reading task at shrinking pixel sizes to find where vision breaks down; the rest probe counting in clutter, spatial grounding, chart reading, and screenshot reading.

Each challenge is an authored generate-script that draws a random seed per run, builds the SVG, and computes the answer; the platform serves the rasterized PNG at a signed URL. Neither the image nor its answer exists until the run is created.

Graded server-side against the value the generator drew into the image — exact for characters, counts, and titles; tolerance-based for chart values. Part of the AI Doctor checkup →

time budget · 15 min

challenges

recent runs