Model results · Anthropic
Claude Sonnet 5.5
Every published VulcanBench measurement of Claude Sonnet 5.5, newest first: combined score, Code quality, runtime, tokens and Claude Code’s list-price cost at every effort level on Frontier v4, graded by deterministic hidden tests on real tasks. Measured through Claude Code on a Claude Max subscription, with Claude Code’s refusal fallback on (no run used it). For the earlier Claude Sonnet 5, see its own page; Claude Opus 5.5, the other Claude 5.5 model on the board, has its own page too.