Model results · OpenAI

GPT-5.6 Terra

Suitev4
GPT-5.6 Terra across every effort level
In Codex across five effort levels: combined score 59.70 to 89.46 from Low to Max, rising at every step, and Code quality 70.93 to 73.90, flat across the ladder. Passes every judged task at Max. 6.7 to 19.2 minutes per task.
September 16, 2026 · 23 tasks · 5 effort levels · 114 runs of GPT-5.6 Terra · VulcanBench-SWE v4
Suitev4
Cost and tokens across effort levels
$0.58 to $2.12 per task at list API rates and 1.21M to 5.86M raw tokens per task across the five effort levels; the whole sweep priced at $142.73.
September 16, 2026 · API-equivalent estimates at list rates, solver inference only · VulcanBench-SWE v4

How would GPT-5.6 Terra do on your stack?

Public results answer "which model is stronger on this suite." Whether GPT-5.6 Terra belongs in your routing table depends on your languages, your codebases, and your task mix, that gets measured, not guessed. Grading rules are in the methodology.