Every published VulcanBench measurement of Claude Opus 5, newest first: accuracy, cost, tokens, and time, graded by deterministic hidden tests on real tasks.
Public results answer "which model is stronger on this suite." Whether Claude Opus 5 belongs in
your routing table depends on your languages, your codebases, and your task mix — that gets measured, not
guessed. Grading rules are in the methodology.