Every published VulcanBench measurement of Grok 4.6, newest first: accuracy, cost, tokens, and time, graded by deterministic hidden tests on real tasks.
Public results answer "which model is stronger on this suite." Whether Grok 4.6 belongs in
your routing table depends on your languages, your codebases, and your task mix — that gets measured, not
guessed. Grading rules are in the methodology.