pass@1 is the mean per-task success rate across all runs in the column. Repeat-swept models aggregate three or more attempts per task; single-pass columns (n≤23) carry wider uncertainty. Cost is total spend at list API prices ÷ runs in the column; time is sandbox wall-clock per task.
| # | Model | pass@1 | ±1 se | $/task run | min/task | Runs |
|---|---|---|---|---|---|---|
| 1 | Grok 4.5 highLeader | 89.9% | 6.1 | $0.47 | 8.4 | 50 |
| 2 | Claude Fable 5 low · 19/23 tasks* | 89.5% | 7.2 | $0.51 | 3.5 | 19 |
| 3 | DeepSeek V4-Flash maxCheapest | 88.4% | 6.2 | $0.06 | 11.2 | 68 |
| 4 | DeepSeek V4 Pro high | 87.0% | 6.2 | $0.11 | 7.4 | 69 |
| 5 | GPT-5.6 Terra mediumFastest | 87.0% | 7.2 | $0.19 | 2.3 | 69 |
| 6 | Claude Opus 5 low · Report 10† | 87.0% | 7.2 | $0.61 | 5.2 | 23 |
| 7 | Grok 4.6 medium | 87.0% | 7.2 | $0.69 | 14.2 | 23 |
| 8 | GPT-5.6 Sol high | 87.0% | 7.2 | $0.69 | 4.2 | 23 |
| 9 | GPT-5.6 Luna high | 85.5% | 6.6 | $0.17 | 3.8 | 69 |
| 10 | Qwen3.8-Max low | 81.2% | 7.8 | $0.61 | 19.6 | 69 |
| 11 | Claude Haiku 4.5 default · 21/23 tasks* | 76.2% | 9.5 | $0.42 | 5.3 | 21 |
| 12 | Kimi K3 extra-high · 19/23 tasks* | 73.7% | 10.4 | $0.69 | 17.2 | 19 |
* Partial task coverage: Claude Fable 5 excludes tasks refused by its safety filters (19/23 at low); Kimi K3 19/23; Claude Haiku 4.5 21/23. † Claude Opus 5 rows come from Report 10 (single runs, 2026-07-26). Effort names are each provider’s own enum: DeepSeek’s top level is “max”, Qwen’s and Grok 4.6’s is “xhigh”. Claude Haiku 4.5 and Kimi K3 have a single tested setting.
The rows above show one effort level per model. This card shows every level that was tested, on one shared y-scale: more reasoning is not reliably better on this suite, and for several models it is worse.
VulcanBench Eval Suite 3: 23 tasks from real merged post-cutoff PRs (Python 9, TypeScript 4, Rust 4, Go 3, JavaScript 3), graded by deterministic hidden tests in Docker with model judges disabled. Each row aggregates every fresh run of that model at that effort whose task hashes match the frozen suite; pass@1 is the mean per-task success rate and its standard error is the spread of per-task means. The board is regenerated from the public run data with the harness’s scripts/rankings-chart pipeline; nothing here is edited by hand. Individual reports on the Benchmarks page carry the per-task detail, failure modes, and caveats behind each row.