Model results · Moonshot AI

Kimi K3

Report08
Kimi K3 vs Grok 4.5, Claude Fable 5 & GPT-5.6 Sol benchmark
Kimi K3 joins the VulcanBench v3 field: Grok 4.5 leads all eleven configurations at 91.3% with the lowest cost per solved task; K3 debuts at 74%, rising to 87% under an extended-budget ablation.
2026-07-19

How would Kimi K3 do on your stack?

Public results answer "which model is stronger on this suite." Whether Kimi K3 belongs in your routing table depends on your languages, your codebases, and your task mix — that gets measured, not guessed. Grading rules are in the methodology.