Model results · OpenAI

GPT-5.6 Sol

Report08
Kimi K3 vs Grok 4.5, Claude Fable 5 & GPT-5.6 Sol benchmark
Kimi K3 joins the VulcanBench v3 field: Grok 4.5 leads all eleven configurations at 91.3% with the lowest cost per solved task; K3 debuts at 74%, rising to 87% under an extended-budget ablation.
2026-07-19
Report07
Grok 4.5 vs Claude Fable 5 vs GPT-5.6 Sol benchmark
Grok 4.5 leads at 91% from medium effort onward on the v3 suite; only GPT-5.6 Sol rewards the effort knob; two frontier-hard tasks go 0-for-27.
2026-07-12

How would GPT-5.6 Sol do on your stack?

Public results answer "which model is stronger on this suite." Whether GPT-5.6 Sol belongs in your routing table depends on your languages, your codebases, and your task mix — that gets measured, not guessed. Grading rules are in the methodology.