Model results · OpenAI

GPT-6 Sol

Suitev4
GPT-6 Sol across every effort level
In Codex across five effort levels: combined score 67.62 at Low, 80.60 at Medium, 82.83 at High, 85.94 at Extra-high and 86.82 at Max. Tasks passed 4, 13, 15, 18 and 19 of 23; Code quality 62.53 to 70.78; 13.0 to 24.5 minutes per task. Medium is judged on 22 of 23 after one invalid judge probe. Slightly behind GPT-5.6 Sol at every level.
September 30, 2026 · 23 tasks · 5 effort levels · 115 runs of GPT-6 Sol · VulcanBench Frontier v4 · Code quality protocol v3.17
→
Suitev4
Cost and tokens across effort levels
$1.02 to $2.52 per task at list API rates and 3.26M to 8.32M raw tokens per task across the five effort levels; the whole sweep prices at $221.38 for 115 runs, $1.93 per task.
September 30, 2026 · API-equivalent estimates at list rates, solver inference only · VulcanBench Frontier v4
→
Suitev4
Beside GPT-5.6 Sol, and the GPT-6 family
GPT-6 Sol scores slightly below GPT-5.6 Sol at every effort level, within one standard error at Extra-high and Max. At half the list price per token it uses 1.8 to 3.9 times the tokens, so it costs more per task from Medium up, and it is slower at every level. A second card sets GPT-6 Sol and GPT-6 Luna beside GPT-5.6 Sol and GPT-5.6 Luna.
September 30, 2026 · comparison cards · VulcanBench Frontier v4
→

How would GPT-6 Sol do on your stack?

Public results answer "which model is stronger on this suite." Whether GPT-6 Sol belongs in your routing table depends on your languages, your codebases, and your task mix, that gets measured, not guessed. Grading rules are in the methodology.