Model results · xAI
Grok 4.7
Every published VulcanBench measurement of Grok 4.7, newest first: combined score, Code quality, runtime and tokens at every effort level Cursor offers, graded by deterministic hidden tests on real tasks, and its VulcanBench Routine v1 and Safety v1 results. Measured through Cursor’s agent CLI on the Cursor subscription. Grok 4.7 is a newer model than Grok 4.6, which has its own page. Its Code quality is judged by Muse Spark 1.3 and GPT-6.1 Sol rather than the Muse Spark 1.3 and Grok 4.6 pair used for every other Frontier v4 column, so read the report’s judge note before comparing.