Model results · OpenAI

GPT-6.1 Sol

Suitev4
GPT-6.1 Sol across every effort level
In Codex across five effort levels: combined score 86.22 at Low, 86.44 at Medium, 88.23 at High, 88.78 at Extra-high and 88.35 at Max. Tasks passed 20, 22, 23, 23 and 23 of 23, with every hidden test passed from High up; Code quality 70.25 to 74.99; 6.9 to 15.4 minutes per task. Every cell is judged on 23 runs.
October 1, 2026 · 23 tasks · 5 effort levels · 115 runs of GPT-6.1 Sol · VulcanBench Frontier v4 · Code quality protocol v3.18
→
Suitev4
Cost and tokens across effort levels
$0.31 to $0.46 per task at list API rates and 0.91M to 1.62M raw tokens per task across the five effort levels; the whole sweep prices at $44.33 for 115 runs, $0.39 per task.
October 1, 2026 · API-equivalent estimates at list rates, solver inference only · VulcanBench Frontier v4
→
Suitev4
Three generations of Sol
GPT-6.1 Sol leads GPT-6 Sol and GPT-5.6 Sol at every effort level on combined score, tasks passed and cost per task, at 14% to 30% of GPT-6 Sol's cost per task. On the Frontier v4 board its best level ranks 13th of 49 columns, below Fable 5.1, Opus 5.5 and GPT-6 Astra, mainly on Code quality.
October 1, 2026 · comparison card · VulcanBench Frontier v4
→

How would GPT-6.1 Sol do on your stack?

Public results answer "which model is stronger on this suite." Whether GPT-6.1 Sol belongs in your routing table depends on your languages, your codebases, and your task mix, that gets measured, not guessed. Grading rules are in the methodology.