Model results · Anthropic

Claude Opus 4.8

Report06
Claude Opus 4.8 vs Haiku 4.5 reliability benchmark
The first reliability study on the suite: five runs per task. On pass^5 the cheapest model wins, the frontier model is the flaky one, and the frontier-hard tier holds at 0 solves in 30 attempts.
2026-07-08
Report05
Claude Fable 5 vs Opus 4.8 frontier-hard benchmark
Two models, two effort levels, on a suite with five frontier-hard tasks. Binary pass@1 is a four-way tie; effort helps Opus and backfires for Fable.
2026-07-04
Report04
Claude Fable 5 vs Opus 4.8 vs GPT-5.5 benchmark
Four frontier configurations on 15 real engineering evals. All score 14/15 but each misses a different task; cost spans 4x.
2026-07-02
Report01
Claude Sonnet 5 vs Opus 4.8 benchmark
The pilot run that established cost as the live axis: both frontier models solve every real, decontaminated bug, and the gap is entirely economic.
2026-07-01
Report02
Claude Sonnet 5 vs Opus 4.8 reasoning-effort benchmark
Does more reasoning effort pay off? A 936-run sweep of Sonnet 5 and Opus 4.8 across low, medium, and high effort.
2026-06-30

How would Claude Opus 4.8 do on your stack?

Public results answer "which model is stronger on this suite." Whether Claude Opus 4.8 belongs in your routing table depends on your languages, your codebases, and your task mix — that gets measured, not guessed. Grading rules are in the methodology.