Model results · Anthropic

Fable 5.1

Suitev4
GPT-6 Astra vs. Fable 5.1 under a neutral Code quality panel
In Claude Code across five effort levels: combined score 89.12 to 91.84 from Low to Max and Code quality 78.36 to 82.43, the higher score on both at every effort, with the widest lead on human readability. Slower than Astra at every effort, at 26.9 to 38.8 minutes per task.
September 9, 2026 · 23 tasks · 5 effort levels · 115 runs of Fable 5.1 · VulcanBench-SWE v4
Suitev4
Cost and tokens across effort levels
$7.96 to 14.49 per task at list API rates and 3.82M to 11.44M raw tokens per task across the five effort levels, 3.5 to 6.3 times Astra's cost; cached prefixes make the token count a poor cost proxy.
September 9, 2026 · API-equivalent estimates at list rates, solver inference only · VulcanBench-SWE v4

How would Fable 5.1 do on your stack?

Public results answer which model is stronger on this suite. Whether Fable 5.1 in Claude Code belongs in your routing table depends on your languages, your codebases and your task mix, and that gets measured, not guessed. Grading rules are in the methodology.