Model results · OpenAI

GPT-6 Luna

Suitev4
GPT-6 Luna across every effort level
In Codex across five effort levels: combined score 40.83 at Low, 45.72 at Medium, 57.35 at High, 78.79 at Extra-high and 81.42 at Max over judged runs. Six runs hit the 3-hour bound (2 at Extra-high, 4 at Max); counting them as 0 gives 71.94 at Extra-high and 67.26 at Max. Tasks passed 0, 0, 2, 11 and 12 of 23; Code quality 65.84 to 71.51; 3.6 to 60.5 minutes per task.
September 28, 2026 · 23 tasks · 5 effort levels · 115 runs of GPT-6 Luna · VulcanBench Frontier v4 · Code quality protocol v3.16
→
Suitev4
Cost and tokens across effort levels
$0.010 to $0.160 per priced task at list API rates and 0.40M to 8.47M raw tokens per task across the five effort levels; the whole sweep prices at $7.26 for 109 finished runs. The six timeouts are unpriced, not $0.
September 28, 2026 · API-equivalent estimates at list rates, solver inference only · VulcanBench Frontier v4
→

How would GPT-6 Luna do on your stack?

Public results answer "which model is stronger on this suite." Whether GPT-6 Luna belongs in your routing table depends on your languages, your codebases, and your task mix, that gets measured, not guessed. Grading rules are in the methodology.