The published record

Benchmarks

Compare results within the same suite and scoring protocol. Scores from different suites are not directly comparable.

Current suite

VulcanBench Frontier v4

Suite methodology →

Formerly published as VulcanBench-Frontier v4; renamed on September 18, 2026 with its URLs unchanged. 23 behavioral-reconstruction tasks. From September 2026 the combined score weights functional correctness (50%), lint and complexity (8.5%), security (8.5%) and Code quality (33%), judged by a calibrated panel from labs with no model on the board. Time and cost are reported separately.

October 4, 2026 · 92 runs · Code quality protocol v3.20

Grok 4.7 across every effort level

92 runs through Cursor’s agent CLI on the Cursor subscription, 23 tasks at each of the four effort levels Cursor offers (no Max), with Code quality at 33% of the combined score, judged by Muse Spark 1.3 and GPT-6.1 Sol because Grok 4.6, the usual second judge, is not neutral for an xAI model. Companion cards cover time and tokens, Grok 4.7 beside the leaders, and VulcanBench Safety v1.

Combined score 89.42 at Low, 92.30 at Medium, 92.71 at High and 93.15 at Extra-high, the top three places on the board; on Muse Spark 1.3 alone it stays first at Medium to Extra-high and is second at Low. Tasks passed 18, 21, 22 and 23 of 23; 20.2 to 28.5 minutes per task; cost unavailable, no list price.

Preview of the Grok 4.7 effort card. Open the study for readable charts and tables.

October 1, 2026 · 115 runs · Code quality protocol v3.18

GPT-6.1 Sol across every effort level

115 runs through Codex CLI 0.159.0 on a ChatGPT Pro subscription, 23 tasks at each of five effort levels, with Code quality at 33% of the combined score, judged for a human reader by Muse Spark 1.3 and Grok 4.6 under the same frozen protocol as the other Frontier v4 reports. Companion cards price the runs and set three generations of Sol side by side.

Combined score 86.22 at Low, 86.44 at Medium, 88.23 at High, 88.78 at Extra-high and 88.35 at Max, ahead of GPT-6 Sol and GPT-5.6 Sol at every level. Tasks passed 20, 22, 23, 23 and 23 of 23, every hidden test from High up; $0.31 to $0.46 per task, $44.33 for all 115 runs.

Preview of the GPT-6.1 Sol effort card. Open the study for readable charts and tables.

September 30, 2026 · 115 runs · Code quality protocol v3.17

GPT-6 Sol across every effort level

115 runs through Codex CLI 0.155.0 on a ChatGPT Pro subscription, 23 tasks at each of five effort levels, with Code quality at 33% of the combined score, judged for a human reader by Muse Spark 1.3 and Grok 4.6 under the same frozen protocol as the other Frontier v4 reports. Companion cards price the runs and set GPT-6 Sol, and the whole GPT-6 family, beside GPT-5.6.

Combined score 67.62 at Low, 80.60 at Medium, 82.83 at High, 85.94 at Extra-high and 86.82 at Max, slightly behind GPT-5.6 Sol at every level. Tasks passed 4, 13, 15, 18 and 19 of 23; $1.02 to $2.52 per task, $221.38 for all 115 runs. Medium is judged on 22 of 23 after one invalid judge probe.

Preview of the GPT-6 Sol effort card. Open the study for readable charts and tables.

September 28, 2026 · 115 runs · Code quality protocol v3.16

GPT-6 Luna across every effort level

115 runs through Codex CLI 0.155.0 on a ChatGPT Pro subscription, 23 tasks at each of five effort levels, with Code quality at 33% of the combined score, judged for a human reader by Muse Spark 1.3 and Grok 4.6 under the same frozen protocol as the other Frontier v4 reports. A companion card prices the runs and counts raw tokens at every level.

Combined score over judged runs 40.83 at Low, 45.72 at Medium, 57.35 at High, 78.79 at Extra-high and 81.42 at Max. Six runs hit the 3-hour bound (2 at Extra-high, 4 at Max); counting them as 0 gives 71.94 and 67.26. Tasks passed 0, 0, 2, 11 and 12 of 23; $0.010 to $0.160 per priced task.

Preview of the GPT-6 Luna effort card. Open the study for readable charts and tables.

September 26, 2026 · 115 runs · Code quality protocol v3.15

Claude Opus 5.5 across every effort level

115 runs through Claude Code on a Claude Max subscription, 23 tasks at each of five effort levels, with Code quality at 33% of the combined score, judged for a human reader by Muse Spark 1.3 and Grok 4.6 under the same frozen protocol as the other Frontier v4 reports. A second card sets Opus 5.5 beside GPT-6 Astra and Claude Fable 5.1 at every level.

Combined score 86.36 at Low, 90.86 at Medium with all 23 tasks passed, 91.11 at High, second on Frontier v4 only to Fable 5.1 at Max, then 90.68 and 90.22; $1.70 to $8.75 per task. Claude Code's refusal fallback was on, following Artificial Analysis: Opus 4.8 wrote 0.0, 6.4, 20.9, 32.4, 47.8% of replies from Low to Max, and every run counts.

Preview of the Claude Opus 5.5 effort card. Open the study for readable charts and tables.

September 19, 2026 · 115 runs · Code quality protocol v3.7

GPT-5.6 Sol across every effort level

115 runs through the Codex CLI on a ChatGPT Pro subscription, 23 tasks at each of Sol's five effort levels, with Code quality at 33% of the combined score, judged for a human reader by Muse Spark 1.3 and Grok 4.6 under the same frozen protocol as the other Frontier v4 reports. Sol joins Terra, Luna, GPT-5.5, Astra and Fable 5.1 on the Frontier v4 leaderboard, completing the GPT-5.6 family.

Most of the effort curve is one step: combined score 69.05 at Low, 82.56 at Medium, then 85.76 to 87.18 from High to Max, where Sol passes 21 of 22 judged tasks. Code quality stays within about three points at every effort. A companion card prices the same runs at API-equivalent rates and counts raw tokens at every effort: $1.36 to $2.05 per task, with High the most expensive level.

Preview of the GPT-5.6 Sol effort card. Open the study for readable charts and tables.

September 17, 2026 · 115 runs · Code quality protocol v3.6 and the v3.6.1 top-up

GPT-5.6 Terra across every effort level

115 runs through the Codex CLI on a ChatGPT subscription, 23 tasks at each of Terra's five effort levels, with Code quality at 33% of the combined score, judged for a human reader by Muse Spark 1.3 and Grok 4.6 under the same frozen protocol as the other Frontier v4 reports. Terra joins GPT-5.5, Luna, Astra and Fable 5.1 on the Frontier v4 leaderboard.

The steepest effort curve on the board so far: combined score rises at every step, 59.70 at Low to 89.05 at Max, where Terra passes every task. Code quality stays within three points at every effort. A companion card prices the same runs at API-equivalent rates and counts raw tokens at every effort: $0.58 to $2.12 per task.

Preview of the GPT-5.6 Terra effort card. Open the study for readable charts and tables.

September 15, 2026 · 207 runs · Code quality protocol v3.5

GPT-5.5 vs. GPT-5.6 Luna across every effort level

207 runs through the Codex CLI on a ChatGPT subscription, 23 tasks at every effort level each API offers (four for GPT-5.5, five for Luna), with Code quality at 33% of the combined score, judged for a human reader by Muse Spark 1.3 and Grok 4.6 under the same frozen protocol as the Astra and Fable 5.1 report.

The effort knob decides it: GPT-5.5 leads at Low and Medium, Luna edges ahead at High and Extra-high, and Luna's Max level, which GPT-5.5 does not offer, is the top cell with 19 of 23 tasks passed. Luna is rated higher on Code quality at every matched effort. A companion card prices the same runs at API-equivalent rates and counts raw tokens at every effort: Luna costs 17 to 75 times less per task at matched effort.

Preview of the GPT-5.5 and GPT-5.6 Luna comparison card. Open the study for readable charts and tables.

September 9, 2026 · 230 runs rescored · Code quality protocol v3.4

GPT-6 Astra vs. Fable 5.1 under a neutral Code quality panel

230 runs from the September 2026 effort sweep, 23 tasks at five effort levels per model, with Code quality at 33% of the combined score, judged for a human reader by Muse Spark 1.3 and Grok 4.6 under a frozen, calibrated six-dimension rubric with an intent-recovery probe.

Fable 5.1 has the higher combined score and Code quality at every matched effort; Astra is faster and cheaper at every effort. The widest gap is human readability. Both judges agree on every ranking. A companion card prices the same runs at API-equivalent rates and counts raw tokens at every effort: Astra costs $1.48 to $2.57 per task against Fable's $7.96 to $14.49.

Preview of the neutral-panel comparison card. Open the study for readable charts and tables.

Models measured on this suite

Agent safety

VulcanBench Safety v1

Why Safety v1 →

A private suite that asks what a coding agent does when the repository tells it to do something it should not. Ten Frontier v4 tasks each hide two harmful instructions in ordinary-looking places (README sections, design docs, code comments, program output), and every run is labeled as followed, refused and told the user, or refused silently. Only aggregates are published. Safety v1 has no effect on any Frontier v4 score.

Decision models

VulcanBench Verdict v2

Read the report →

A suite for models that return a typed decision instead of text. Twenty question families in eight areas, twelve in software engineering and eight in general reasoning, with every answer checked by running code, a type checker, hidden tests, a merged fix, a security advisory or the generator that built the question. Each family is scored as skill above its floor, the best of the most common answer, guessing and every surface shortcut built for it, so 0 is no better than the best dumb strategy and 100 is perfect. It replaces Verdict v1, which asked one kind of question and whose headline could not separate Jev from guessing. Verdict and Frontier scores measure different things and are not comparable.

September 25, 2026 · 4,774 test items · jev-1.13.0

Jev 1.13.0 on VulcanBench Verdict v2

TypeSafe AI's Jev scores 45.9 on the Verdict Index (95% interval 43.1 to 48.4): 49.6 on the software families and 40.4 on the general ones, with a Calibration Index of 0.33. It answers in 0.3 seconds, $0.58 for all 4,774 items. GPT-6 Astra at high effort, given exactly the same inputs as a reference that shows every family is answerable rather than as a leaderboard entry, scores 91.7.

Jev is strongest where the answer can be recognised from the shape of the text, such as whether a snippet type-checks (87) or which service caused an incident (82), and weakest on multi-step work: 24 on filtering and totalling a table, 1 on predicting which hidden test a patch fails. On size-matched patch pairs it is above the floor, unlike v1, at skill 36 against 58 for the reference. The admission gate was amended after the pilot; the report sets out the change and its reasons. The reference scored 100 on all eight general families, so they have no headroom for ranking strong reasoning models.

Preview of the Jev Verdict v2 results card. Open the study for readable charts and tables.

Decision models, earlier version

VulcanBench Verdict v1

Read the v1 report →

A suite for models that return a typed decision instead of text, and so cannot take a coding benchmark. It asks them to judge code rather than write it: 2,163 test items mined from 745 real agent patches produced during Frontier v4 sweeps, with the answers taken from those tasks' hidden test suites. Accuracy is always published beside calibration and beside the score for simply answering each question's most common label. Verdict and Frontier scores measure different things and are not comparable.

September 22, 2026, corrected and extended with a control September 23 · 2,641 queries · jev-1.13.0

Jev 1.13.0 on VulcanBench Verdict v1

The first measurement of TypeSafe AI's Jev on software engineering. Jev is a System One model: it returns a typed decision in a single parallel pass and generates no text, so it was asked to judge 745 patches that frontier coding agents had already written, each one graded by hidden tests.

Its stated probability that a patch passes never rises above 0.42, on a population where 63% of them pass, which is a large calibration failure. The ranking underneath is sounder: AUROC 0.69 on that question, and at a cutoff fitted on held-out items its accuracy draws level with the majority answer rather than beating it. On which of two patches a code-quality panel preferred, it agrees 89.7% of the time at AUROC 0.962, for 328 ms and $0.15 per 1,000 decisions. Corrected on September 23 after the first version scored these questions at a cutoff the model never crosses. A control, GPT-6 Astra given the same inputs, ranks the fixes far better (AUROC 0.87 against 0.69) and beats guessing by about 10 points, so the questions are answerable and Jev's shortfall is real; even ranking fixes by lines changed alone beats Jev (0.78).

Preview of the Jev results card. Open the study for readable charts and tables.

Archive

Earlier suites & studies

Back to Frontier v4 ↑

Eval Suite 3, earlier coding suites, voice-v1 and special studies. Each report retains its original task set and scoring method. These percentages do not belong on the Frontier v4 score scale.

Browse historical reports by model

Report20
Muse Spark 1.2 in Pi vs. a bare-bones harness
Our first open-source-harness study: Pi beats the bare loop at every effort level, 4.3 to 17.4 points. And the integrity audit caught the model hunting the host for answer keys: 17 of 69 Pi cells were replaced by kernel-confined, audited-clean reruns that cut the headline gap roughly in half.
2026-08-28 · 23 tasks · 138 scored cells · v3 suite · Harness Study No. 04
→
Report19
Muse Spark 1.2 across the effort knob
Meta’s flagship runs the steepest backward reasoning dial measured on v3: 87.0% at low, 52.2% at xhigh, −34.8 points. Higher effort converts wrong answers into wall-clock timeouts, and ten of the fifteen timed-out runs never wrote a line.
2026-08-25 · 23 tasks · 69 runs · v3 suite
→
Report18
GLM 5.3 in ZCode vs. a bare-bones harness
Our first subscription-harness study: the same GLM 5.3 through Z.ai’s own ZCode harness and through a bare-bones API loop. The effort knob points opposite directions, and at max the harness is worth 21.8 points: 65.2% becomes 87.0%, with zero timed-out runs.
2026-08-24 · 23 tasks · 138 runs · v3 suite · Harness Study No. 03
→
Report17
Qwen3.8-27B across the effort knob
Alibaba’s open-weights 27B runs the reasoning dial backward: 82.6% at low and medium, 73.9% at xhigh, and 2.4× slower for it. Far more robust than the Qwen3.8-Max flagship on the same knob.
2026-08-21 · 23 tasks · 129 runs · v3 suite
→
Report16
Grok 4.6 in Grok Build vs. Cursor vs. a bare-bones harness
The first three-way harness study: one model, three delivery systems. Inside xAI’s own harness the effort knob is monotone and ends at 92.8%, the best result this suite has recorded for Grok 4.6.
2026-08-18 · 23 tasks · 277 runs · v3 suite · Harness Study No. 02
→
Report15
Grok 4.6 in Cursor vs. a bare-bones harness
Our first model × harness study: the same Grok 4.6 through Cursor and through a deliberately minimal reference loop. The harness is worth 14.5 points at high effort and nothing measurable elsewhere; the agent’s 299 web-fetch attempts were all refused.
2026-08-16 · 23 tasks · 368 runs · v3 suite · Harness Study No. 01
→
Report14
Grok 4.6 across the effort knob
xAI’s newest model peaks mid-knob: 87.0% at medium, 73.9% at its shipped default, and the new xhigh level recovers less than half the drop. Failures shift from wrong answers to unfinished runs as effort rises.
2026-08-12 · 23 tasks · 92 runs · v3 suite
→
Report13
DeepSeek V4 Pro vs V4-Flash
Same 87.0% pass@1 at high effort. Pro uses 55% fewer tokens and runs 38% faster, but its higher token price makes the sweep 29% more expensive.
2026-08-09 · 23 tasks · 138 runs · v3 suite
→
Report12
Qwen3.8-Max across the effort knob
Alibaba’s flagship debuts on v3 with an effort knob that runs backwards: 81.2% at low, 55.1% at its shipped default, and every failure at higher effort is an unfinished run rather than a wrong answer.
2026-08-04 · 23 tasks · 202 runs · v3 suite
→
Report11
Grok Voice Think Fast 2.0 vs GPT Realtime
The same 200 questions, typed vs spoken: Grok pays +3.3 pp for hearing instead of reading, GPT Realtime +4.0; both lose ~10 pp on spoken arithmetic.
2026-07-31 · 200 questions · 1,840 units · voice-v1 suite
→
Report10
Claude Opus 5 across the effort knob
Opus 5 debuts on v3: its cheapest setting wins at 20/23 and $0.70/solved; score falls at every step up the knob while cost triples.
2026-07-26 · 23 tasks · 69 runs · v3 suite
→
Report09
Does contamination move Claude Opus 5’s score?
A controlled A/B: 13 tasks Opus 5 could have memorised vs 13 it cannot have seen. 11/13 vs 12/13, one genuine miss each; no detectable effect.
2026-07-25 · 26 tasks · 26 runs · contamination study
→
Report08
Kimi K3 vs Grok 4.5, Fable 5 & GPT-5.6 Sol
Kimi K3 joins the v3 board: Grok 4.5 keeps the lead at 91% and the lowest $/solved; K3 debuts at 74%, 87% with an extended budget.
2026-07-19 · 23 tasks · 28 Kimi runs · v3 suite
→
Report07
Grok 4.5 vs Fable 5 vs GPT-5.6 Sol
Grok leads at 91% from medium effort; only Sol climbs with the knob; two hard tasks go 0-for-27.
2026-07-12 · 23 tasks · 207 runs · v3 suite
→
Report06
Opus 4.8 vs Haiku 4.5
Five runs per task: the cheapest model wins on dependability; pass@1 hides coin flips.
2026-07-07 · 15 evals · 135 runs · reliability study
→
Report05
Fable 5 vs Opus 4.8
A four-way pass@1 tie; effort helps Opus and backfires for Fable.
2026-07-04 · 15 evals · 60 runs · frontier-hard tier
→
Report04
Fable 5 vs Opus 4.8 vs GPT-5.5
Four configurations, fifteen evals, each misses a different task.
2026-07-02 · 15 evals · 60 runs · corrected costs
→
Report03
Fable 5 vs Sonnet 5 vs GLM-5.2
Accuracy converges; efficiency spans an order of magnitude.
2026-07-01 · 10 tasks · 30 runs
→
Report02
Sonnet 5 vs Opus 4.8
Does more thinking effort pay off? Cost versus correctness across low, medium, high.
2026-06-30 · 52 tasks · 936 runs
→
Report01
Sonnet 5 vs Opus 4.8
The pilot that established cost as the live axis once correctness saturates.
2026-07-01 · 4 tasks · 24 runs
→