Every published VulcanBench measurement of GPT-5.5, newest first: accuracy, cost, tokens, and time, graded by deterministic hidden tests on real tasks.
Public results answer "which model is stronger on this suite." Whether GPT-5.5 belongs in
your routing table depends on your languages, your codebases, and your task mix — that gets measured, not
guessed. Grading rules are in the methodology.