All VulcanBench Frontier v4 reports

September 26, 2026 · 115 runs · Code quality protocol v3.15

Claude Opus 5.5 across every effort level

Across every effort level

Claude Opus 5.5 ran every Frontier v4 task once at each of five effort levels, 23 tasks per cell, through Claude Code 2.1.280 on a Claude Max subscription on September 22 to 24, 2026. Combined score weights are 50% functional correctness, 8.5% lint and complexity, 8.5% security and 33% Code quality, judged by Muse Spark 1.3 and Grok 4.6 with a ground-truth intent-recovery probe. Judging ran on September 24 and 25.

Effort pays once, then flattens. Low scores 86.36 and passes 17 of 23 tasks. Medium, Anthropic's shipped default, passes all 23 at 90.86 for $2.84 a task. High, extra-high and max land between 90.22 and 91.11, with standard errors that overlap medium's, while cost per task climbs to $8.75 at max.

Claude Opus 5.5 (with fallback) at each effort level, 23 runs per cell (judged: 22 at high, 23 elsewhere)
MetricLowMediumHighExtra-highMax
Combined score86.3690.8691.1190.6890.22
Standard error1.300.640.571.101.02
Tasks passed17/2323/2322/2322/2321/23
Code quality72.5377.3277.9979.9879.77
Human readability67.174.675.578.677.5
Intent recovery78.377.879.379.781.0
Runs with Opus 4.8 turns037812
Opus 4.8 share of replies0.0%6.4%20.9%32.4%47.8%
Minutes per task13.718.621.420.737.5
Raw tokens per task2.77M5.01M4.11M3.90M8.72M
Cost per task$1.70$2.84$3.27$4.11$8.75

Tasks passed counts every run; the high cell's combined score and Code quality are over its 22 judged runs. Runtime and cost cover every run, so high's runtime here (21.4 minutes) differs slightly from the card, which averages the judged runs.

Claude Opus 5.5 across five effort levels under the v3.15 Code quality protocol. Combined score is 86.36 at low, 90.86 at medium, 91.11 at high, 90.68 at extra-high and 90.22 at max. A table gives Code quality, tasks passed, runs with Opus 4.8 turns, Opus 4.8 share of replies and cost per task at every level. Exact values are in the adjacent table.
Combined score uses a focused scale; runtime starts at zero. Whiskers show ±1 task standard error. Open full-size card.

Beside GPT-6 Astra and Claude Fable 5.1

On this suite Opus 5.5's best level, high at 91.11, is second only to Fable 5.1 at max (91.84) and ahead of GPT-6 Astra's best (89.30 at max). Opus 5.5 leads Fable 5.1 at medium and high and trails it at low, extra-high and max, costing 21 to 34% of Fable 5.1's price per task below max and about the same at max. Astra is the cheapest of the three at every level except low, where Opus 5.5 is two cents cheaper. Astra and Fable 5.1 are the published v3.4 rows; nothing was re-judged.

Frontier v4 combined score and cost per task at every effort level for Claude Opus 5.5, GPT-6 Astra and Claude Fable 5.1, with a table of combined score, tasks passed, cost and Opus 4.8 reply share for each model and level.
Both Claude columns ran with Claude Code's refusal fallback on and carry their Opus 4.8 reply share. Opus 5.5 and Fable 5.1 ran on different Claude Code versions (2.1.280 and 2.1.259 to 2.1.261), so small gaps between them are harness confounded. Open full-size card.

Refusal fallback, and what it means for these numbers

When a safeguard classifier refuses a turn, Claude Code can finish the session on another model, here claude-opus-4-8. We follow Artificial Analysis, which publishes Opus 5.5 as "Default Fallback": the fallback stays on and every run counts. We go one step further and publish how much of each cell the fallback model wrote. Opus 4.8 wrote some replies in 0, 3, 7, 8 and 12 runs from low to max, and 0.0, 6.4, 20.9, 32.4 and 47.8% of all replies. The higher effort levels therefore partly measure Opus 4.8, and max is close to an even blend. The published Fable 5.1 column carries the same effect: 11 of its runs fell back.

Cost is Claude Code's own list-price total for every model that served a run. The harness's per-run estimate omitted the fallback model's usage (one extra-high run showed $0.15 against $19.69), so it is not used here.

What was judged, and by whom

114 of 115 runs are judged. On high depotcore a safeguard classifier stop cut off the turn that would have written the module, so the patch is empty and there is nothing to review; the run scored 0 and stays in tasks passed, runtime and cost. Both judges, Muse Spark 1.3 and Grok 4.6, passed the full calibration exam under v3.15 with no allowance used, and no review fell back to another judge model. They are neutral for an Anthropic submission.

Evidence

Public files contain all 115 run measurements, both judges' sub-scores and intent recovery per judged run, raw tokens by serving model, replies by serving model, cost per run, per-cell aggregates, every calibration gate for both judges and the exact judge protocol text. The combined score is 100 × (0.50F + 0.085Q + 0.085S + 0.33C / 100). Summed solver time is 42.94 hours and API-equivalent solver cost $475.33 at list prices; judging is excluded.

No human rated anything; the scores are model judgment for a human reader. Runs were not repeated. Withheld: raw prompts and responses, submitted patches and reconstructed sources, the quirk answer keys (they describe hidden-test behaviour), and reviewer session identifiers.