Model results · Anthropic

Claude Haiku 4.5

Report06
Claude Opus 4.8 vs Haiku 4.5 reliability benchmark
The first reliability study on the suite: five runs per task. On pass^5 the cheapest model wins, the frontier model is the flaky one, and the frontier-hard tier holds at 0 solves in 30 attempts.
2026-07-08

How would Claude Haiku 4.5 do on your stack?

Public results answer "which model is stronger on this suite." Whether Claude Haiku 4.5 belongs in your routing table depends on your languages, your codebases, and your task mix — that gets measured, not guessed. Grading rules are in the methodology.