Model results · OpenAI

GPT-6 Astra

Suitev4
GPT-6 Astra vs. Fable 5.1 under a neutral Code quality panel
In Codex across five effort levels: combined score 87.43 to 89.30 from Low to Max and Code quality 70.31 to 76.26, second to Fable 5.1 at every effort on both. The faster model at every effort, at 3.8 to 10.3 minutes per task.
September 9, 2026 · 23 tasks · 5 effort levels · 115 runs of GPT-6 Astra · VulcanBench-SWE v4
Suitev4
Cost and tokens across effort levels
$1.48 to 2.57 per task at list API rates and 0.68M to 0.95M raw tokens per task across the five effort levels; the cheaper model at every effort, with a long-context upper bound that keeps it cheaper still.
September 9, 2026 · API-equivalent estimates at list rates, solver inference only · VulcanBench-SWE v4

How would GPT-6 Astra do on your stack?

Public results answer which model is stronger on this suite. Whether GPT-6 Astra in Codex belongs in your routing table depends on your languages, your codebases and your task mix, and that gets measured, not guessed. Grading rules are in the methodology.