Model results · OpenAI

GPT-5.6 Luna

Suitev4
GPT-5.5 vs. GPT-5.6 Luna across every effort level
In Codex across five effort levels: combined score 41.29 to 84.15 from Low to Max and Code quality 65.35 to 71.73. Behind GPT-5.5 at Low and Medium, ahead by less than one standard error at High and Extra-high, and the top cell on the board at Max with 19 of 23 tasks passed. Rated higher on Code quality at every matched effort. 2.7 to 44.1 minutes per task.
September 15, 2026 · 23 tasks · 5 effort levels · 115 runs of GPT-5.6 Luna · VulcanBench-SWE v4
Suitev4
Cost and tokens across effort levels
$0.02 to $0.45 per task at list API rates and 0.42M to 11.88M raw tokens per task across the five effort levels; 17 to 75 times cheaper than GPT-5.5 at matched effort.
September 15, 2026 · API-equivalent estimates at list rates, solver inference only · VulcanBench-SWE v4

How would GPT-5.6 Luna do on your stack?

Public results answer "which model is stronger on this suite." Whether GPT-5.6 Luna belongs in your routing table depends on your languages, your codebases, and your task mix, that gets measured, not guessed. Grading rules are in the methodology.