Model results · OpenAI

GPT-5.5

Suitev4
GPT-5.5 vs. GPT-5.6 Luna across every effort level
In Codex across four effort levels: combined score 49.80 to 78.46 from Low to Extra-high and Code quality 61.75 to 67.67. Ahead of Luna at Low and Medium, behind by less than one standard error at High and Extra-high; its API has no Max level. 1 to 11 of 23 tasks passed, at 8.2 to 20.8 minutes per task.
September 15, 2026 · 23 tasks · 4 effort levels · 92 runs of GPT-5.5 · VulcanBench Frontier v4
→
Suitev4
Cost and tokens across effort levels
$1.79 to $4.56 per task at list API rates and 1.72M to 4.72M raw tokens per task across the four effort levels; 17 to 75 times Luna's cost at matched effort.
September 15, 2026 · API-equivalent estimates at list rates, solver inference only · VulcanBench Frontier v4
→
Report04
Claude Fable 5 vs Opus 4.8 vs GPT-5.5 benchmark
Four frontier configurations on 15 real engineering evals. All score 14/15 but each misses a different task; cost spans 4x.
2026-07-02
→

How would GPT-5.5 do on your stack?

Public results answer "which model is stronger on this suite." Whether GPT-5.5 belongs in your routing table depends on your languages, your codebases, and your task mix, that gets measured, not guessed. Grading rules are in the methodology.