Model results · xAI

Grok 4.6

Report16
Grok 4.6 in Grok Build vs. Cursor vs. a bare-bones harness
The first three-way harness study. Inside xAI’s own agent CLI the effort knob rises at every step, from 84.1% to 92.8%, the best pass@1 this suite has recorded for the model.
2026-08-18
→
Report15
Grok 4.6 in Cursor vs. a bare-bones harness
Holding the model fixed and varying the scaffold: Cursor is worth 14.5 points at high effort, and nothing measurable at any other level.
2026-08-16
→
Report14
Grok 4.6 reasoning-effort benchmark
xAI’s newest model peaks mid-knob: 87.0% at medium, 73.9% at its shipped default, and the new xhigh level recovers less than half the drop. Failures shift from wrong answers to unfinished runs as effort rises.
2026-08-13
→

How would Grok 4.6 do on your stack?

Public results answer "which model is stronger on this suite." Whether Grok 4.6 belongs in your routing table depends on your languages, your codebases, and your task mix, that gets measured, not guessed. Grading rules are in the methodology.