Every published VulcanBench measurement of GPT-6.1 Sol, newest first: combined score, Code quality, runtime, tokens and API-equivalent cost at every effort level, graded by deterministic hidden tests on real tasks. Measured through the Codex CLI on a ChatGPT subscription. GPT-6.1 Sol is a newer model than GPT-6 Sol and GPT-5.6 Sol, which have their own pages.
Public results answer "which model is stronger on this suite." Whether GPT-6.1 Sol belongs in
your routing table depends on your languages, your codebases, and your task mix, that gets measured, not
guessed. Grading rules are in the methodology.