Model results · Anthropic

Claude Sonnet 5

Report03
Claude Fable 5 vs Sonnet 5 vs GLM-5.2 benchmark
Accuracy converges across three models while efficiency spans an order of magnitude. Fable 5 at low effort matches Sonnet 5 at high effort for a quarter of the cost.
2026-07-01
Report01
Claude Sonnet 5 vs Opus 4.8 benchmark
The pilot run that established cost as the live axis: both frontier models solve every real, decontaminated bug, and the gap is entirely economic.
2026-07-01
Report02
Claude Sonnet 5 vs Opus 4.8 reasoning-effort benchmark
Does more reasoning effort pay off? A 936-run sweep of Sonnet 5 and Opus 4.8 across low, medium, and high effort.
2026-06-30

How would Claude Sonnet 5 do on your stack?

Public results answer "which model is stronger on this suite." Whether Claude Sonnet 5 belongs in your routing table depends on your languages, your codebases, and your task mix — that gets measured, not guessed. Grading rules are in the methodology.