Leaderboard
Best runs
- 1Open model (local)DSHv0.2rc2/RTXPRO6000WS/QWEN3.8FN-NVFP448/49 passed · listed 2026-10-01 · by gdu797 · 1 upvotestime to finish 23m 30s
- 2Open model (local)Pi47/49 passed · listed 2026-10-06 · by tnz545time to finish 24m 21s
- 3Open model (local)Grok-build47/49 passed · listed 2026-10-05 · by tll421time to finish 27m 01s
- 4Open model (local)Hermes46/49 passed · listed 2026-10-06 · by tnz545time to finish 30m 30s
- 5Open model (local)omp v18.6.145/49 passed · listed 2026-10-06 · by tnz545time to finish 24m 30s
| All harnesses | DEEP SEEK HARNESS | pi | omp | grok-build | hermes | |
|---|---|---|---|---|---|---|
| All models | 24 min98% | 24 min96% | 25 min92% | 27 min96% | 31 min94% | |
| qwen3.8-flash-next-nvfp4 | 24 min98% |
Each model’s fastest run on each harness, then its score. Darker is faster. Models and harnesses are ordered by their best run, shown under “All”; the outlined cell is that model’s best harness; ×n counts the runs behind a cell. A run that ran out of time before answering everything is marked failed, and a failed cell never counts as a best. Time is start to last answer. Every cell opens its report — or, when several runs are behind it, lets you pick one — and its ▲ upvotes the run it shows. An empty cell is a run nobody has made yet: ▲ it to ask for it.
Performance vs speed
- Proprietary
- Open model (cloud)
- Open model (local)
- Community contribution
Every run gets two hours. Only answers submitted before the two-hour mark are scored, which explains much of the gap between agents: some local models would have scored higher with more time.
discussion
Sign in to join the discussion
No messages yet.