Model results · Cognition

Devin SWE-2

Suitev4
Devin SWE-2 across every effort level
In the Devin CLI across all three effort levels SWE-2 offers: combined score 82.43 at medium, 81.42 at high and 86.14 at Max, so medium and high are indistinguishable and only Max moves. Code quality 65.14 to 69.17. Passes 15 of 23 tasks at medium and high and 21 of 23 at Max. 35.1 to 56.6 minutes per task. Code quality is judged by Muse Spark 1.3 alone after both candidates for the second judge seat failed calibration.
September 22, 2026 · 23 tasks · 3 effort levels · 69 runs of Devin SWE-2 · VulcanBench Frontier v4 · Code quality protocol v3.11
Suitev4
Tokens and runtime across effort levels
Cost is unavailable at every effort level: Cognition publishes no per-token rate for SWE-2, and the catalog's Free tier is a promotion, not a rate. Tokens and time are measured instead, 7.53M to 16.80M raw tokens and 35.1 to 56.6 minutes per task, 776M tokens and 53.7 hours in all.
September 22, 2026 · Devin CLI receipts, solver inference only, cost unavailable · VulcanBench Frontier v4

How would Devin SWE-2 do on your stack?

Public results answer "which model is stronger on this suite." Whether Devin SWE-2 belongs in your routing table depends on your languages, your codebases, and your task mix, that gets measured, not guessed. Grading rules are in the methodology.