Model results · OpenAI

GPT-5.6 Sol

Suitev4
GPT-5.6 Sol across every effort level
In Codex across five effort levels: combined score 69.05 to 87.18 from Low to Max, with 13.5 of the 18 points gained between Low and Medium and the top three levels within 1.5 points, and Code quality 64.98 to 68.19, flat across the ladder. Passes 21 of 22 judged tasks at Max. 9.8 to 12.0 minutes per task.
September 19, 2026 · 23 tasks · 5 effort levels · 115 runs of GPT-5.6 Sol · VulcanBench Frontier v4 · Code quality protocol v3.7
→
Suitev4
Cost and tokens across effort levels
$1.36 to $2.05 per task at list API rates and 1.79M to 2.83M raw tokens per task across the five effort levels, with High the most expensive level; the whole sweep priced at $195.55.
September 19, 2026 · API-equivalent estimates at list rates, solver inference only · VulcanBench Frontier v4
→
Report08
Kimi K3 vs Grok 4.5, Claude Fable 5 & GPT-5.6 Sol benchmark
Kimi K3 joins the VulcanBench v3 field: Grok 4.5 leads all eleven configurations at 91.3% with the lowest cost per solved task; K3 debuts at 74%, rising to 87% under an extended-budget ablation.
2026-07-19
→
Report07
Grok 4.5 vs Claude Fable 5 vs GPT-5.6 Sol benchmark
Grok 4.5 leads at 91% from medium effort onward on the v3 suite; only GPT-5.6 Sol rewards the effort knob; two frontier-hard tasks go 0-for-27.
2026-07-12
→

How would GPT-5.6 Sol do on your stack?

Public results answer "which model is stronger on this suite." Whether GPT-5.6 Sol belongs in your routing table depends on your languages, your codebases, and your task mix, that gets measured, not guessed. Grading rules are in the methodology.