Model results · xAI

Grok 4.6

Report14
Grok 4.6 reasoning-effort benchmark
xAI’s newest model peaks mid-knob: 87.0% at medium, 73.9% at its shipped default, and the new xhigh level recovers less than half the drop. Failures shift from wrong answers to unfinished runs as effort rises.
2026-08-13

How would Grok 4.6 do on your stack?

Public results answer "which model is stronger on this suite." Whether Grok 4.6 belongs in your routing table depends on your languages, your codebases, and your task mix — that gets measured, not guessed. Grading rules are in the methodology.