Model results · Anthropic

Claude Opus 5

Report09
Claude Opus 5 benchmark: does contamination move the score?
A controlled A/B on training-data contamination: 13 tasks Claude Opus 5 could have memorised vs 13 it cannot have seen. Result: 11/13 vs 12/13, one genuine miss each — no detectable effect.
2026-07-29
Report10
Claude Opus 5 reasoning-effort benchmark
First measurement of Claude Opus 5 on the full VulcanBench v3 suite: its cheapest setting wins at 20/23 and $0.70 per solved task; score falls at every step up the effort knob while cost triples.
2026-07-28

How would Claude Opus 5 do on your stack?

Public results answer "which model is stronger on this suite." Whether Claude Opus 5 belongs in your routing table depends on your languages, your codebases, and your task mix — that gets measured, not guessed. Grading rules are in the methodology.