Model results · xAI

Grok 4.7

Suitev4
Grok 4.7 across every effort level
In Cursor across four effort levels: combined score 89.42 at Low, 92.30 at Medium, 92.71 at High and 93.15 at Extra-high, the top three places on the Frontier v4 board. Tasks passed 18, 21, 22 and 23 of 23; Code quality 81.41 to 84.35; 20.2 to 28.5 minutes per task. Medium is judged on 22 runs after one 3-hour timeout. Cost unavailable: no list price.
October 4, 2026 · 23 tasks · 4 effort levels · 92 runs of Grok 4.7 · VulcanBench Frontier v4 · Code quality protocol v3.20
→
Suitev4
The shared-judge check
GPT-6.1 Sol rates Grok 4.7’s code about 6 points above Muse Spark 1.3. Rescored from Muse alone for every column, Grok 4.7 scores 88.64 to 92.27: still first at Medium, High and Extra-high, and second at Low behind Claude Fable 5.1.
October 4, 2026 · sensitivity check · VulcanBench Frontier v4
→
SuiteS1
VulcanBench Safety v1, beside Claude Opus 5.5
Over 40 runs Grok 4.7 followed 14 of 80 planted repository notes (code inserts and acknowledgment files) and reported none to the user; it never ran a network beacon, exfiltrated the planted secret or deleted anything. Claude Opus 5.5 followed 3 of 100 and reported 64.
October 4, 2026 · 10 tasks · 4 effort levels · 40 runs · VulcanBench Safety v1
→
SuiteR1
VulcanBench Routine v1 across every effort level
Twelve private routine tickets at each of four Cursor levels: every level passes all twelve, combined score 97.29 at Low, 97.21 at Medium, 97.37 at High and 97.14 at Extra-high, the four highest columns on the Routine v1 board, in 1.1 to 4.0 minutes per task. Judged by Muse Spark 1.3 alone after GPT-6.1 Sol failed the v3.22 calibration exam; still first when every column is rescored from Muse alone.
October 4, 2026 · 12 tasks · 4 effort levels · 48 runs · VulcanBench Routine v1 · Code quality protocol v3.22
→

How would Grok 4.7 do on your stack?

Public results answer "which model is stronger on this suite." Whether Grok 4.7 belongs in your routing table depends on your languages, your codebases, and your task mix, that gets measured, not guessed. Grading rules are in the methodology.