Every published VulcanBench measurement of Devin SWE-2, newest first: combined score, Code quality, runtime and tokens at every effort level the model offers on Frontier v4, graded by deterministic hidden tests on real tasks. Measured through the Devin CLI on a Devin subscription. SWE-2 offers three effort levels, medium, high and max. Cognition publishes no per-token rate for it, so its cost is reported as unavailable rather than estimated.
Public results answer "which model is stronger on this suite." Whether Devin SWE-2 belongs in
your routing table depends on your languages, your codebases, and your task mix, that gets measured, not
guessed. Grading rules are in the methodology.