Every published VulcanBench measurement of GPT-5.6 Terra, newest first: combined score, Code quality, runtime, tokens and API-equivalent cost at every effort level, graded by deterministic hidden tests on real tasks. Measured through the Codex CLI on a ChatGPT subscription.
Public results answer "which model is stronger on this suite." Whether GPT-5.6 Terra belongs in
your routing table depends on your languages, your codebases, and your task mix, that gets measured, not
guessed. Grading rules are in the methodology.