Model results · Anthropic

Claude Sonnet 5.5

Suitev4
Claude Sonnet 5.5 across every effort level
In Claude Code across five effort levels: combined score 81.63 at Low, 84.10 at Medium, 86.79 at High, 90.20 at Extra-high and 92.02 at Max. Tasks passed 15, 17, 20, 22 and 23 of 23; $2.39 to $5.52 a task. Judge versions and the pre-tagged-worktree sweep disclosed.
October 8, 2026 · 23 tasks · 5 effort levels · 115 runs of Claude Sonnet 5.5 · VulcanBench Frontier v4 · Code quality protocol v3.23
→
Suitev4
Claude Sonnet 5.5 vs Claude Opus 5.5 vs GPT-6.1 Sol
The three models side by side at every effort level on the same 23 tasks and judges: combined score with ±1 SE, Code quality, tasks solved, cost and minutes per task. Sonnet 5.5 solves 15 to 23 tasks at $2.39 to $5.52 a task.
October 8, 2026 · 23 tasks · 5 effort levels · 345 runs · VulcanBench Frontier v4
→
Suitev4
Cost and tokens across effort levels
Claude Code’s own list-price total per run, because Anthropic’s page gives two cache-read prices for Sonnet 5.5. $2.88, $3.00, $2.39, $3.38 and $5.52 per task from Low to Max; the whole sweep priced at $394.77 for 115 runs.
October 8, 2026 · API-equivalent estimates, solver inference only · VulcanBench Frontier v4
→
Suitev4
Beside Claude Opus 5.5 and Claude Fable 5.1
Combined score, tasks passed, Code quality, cost and minutes at every level for the three Claude columns on Frontier v4, each judged in its own protocol run with the same rubric and judges.
October 8, 2026 · VulcanBench Frontier v4
→

How would Claude Sonnet 5.5 do on your stack?

Public results answer "which model is stronger on this suite." Whether Claude Sonnet 5.5 belongs in your routing table depends on your languages, your codebases, and your task mix, that gets measured, not guessed. Grading rules are in the methodology.