Model results · xAI

Grok 4.5

Report08
Kimi K3 vs Grok 4.5, Claude Fable 5 & GPT-5.6 Sol benchmark
Kimi K3 joins the VulcanBench v3 field: Grok 4.5 leads all eleven configurations at 91.3% with the lowest cost per solved task; K3 debuts at 74%, rising to 87% under an extended-budget ablation.
2026-07-19
Report07
Grok 4.5 vs Claude Fable 5 vs GPT-5.6 Sol benchmark
Grok 4.5 leads at 91% from medium effort onward on the v3 suite; only GPT-5.6 Sol rewards the effort knob; two frontier-hard tasks go 0-for-27.
2026-07-12

How would Grok 4.5 do on your stack?

Public results answer "which model is stronger on this suite." Whether Grok 4.5 belongs in your routing table depends on your languages, your codebases, and your task mix — that gets measured, not guessed. Grading rules are in the methodology.