Model results · Anthropic

Claude Fable 5

Report08
Kimi K3 vs Grok 4.5, Claude Fable 5 & GPT-5.6 Sol benchmark
Kimi K3 joins the VulcanBench v3 field: Grok 4.5 leads all eleven configurations at 91.3% with the lowest cost per solved task; K3 debuts at 74%, rising to 87% under an extended-budget ablation.
2026-07-19
Report07
Grok 4.5 vs Claude Fable 5 vs GPT-5.6 Sol benchmark
Grok 4.5 leads at 91% from medium effort onward on the v3 suite; only GPT-5.6 Sol rewards the effort knob; two frontier-hard tasks go 0-for-27.
2026-07-12
Report05
Claude Fable 5 vs Opus 4.8 frontier-hard benchmark
Two models, two effort levels, on a suite with five frontier-hard tasks. Binary pass@1 is a four-way tie; effort helps Opus and backfires for Fable.
2026-07-04
Report04
Claude Fable 5 vs Opus 4.8 vs GPT-5.5 benchmark
Four frontier configurations on 15 real engineering evals. All score 14/15 but each misses a different task; cost spans 4x.
2026-07-02
Report03
Claude Fable 5 vs Sonnet 5 vs GLM-5.2 benchmark
Accuracy converges across three models while efficiency spans an order of magnitude. Fable 5 at low effort matches Sonnet 5 at high effort for a quarter of the cost.
2026-07-01

How would Claude Fable 5 do on your stack?

Public results answer "which model is stronger on this suite." Whether Claude Fable 5 belongs in your routing table depends on your languages, your codebases, and your task mix — that gets measured, not guessed. Grading rules are in the methodology.