The published record

Benchmarks

Looking for the ranking? The Eval Suite 3 leaderboard holds every model measured on the current suite in one table, ranked by pass@1 at its best-scoring effort, with cost, speed, and effort-curve cards you can download and share. The reports below carry the per-task detail behind each row.
Claude Fable 5Anthropic Claude Opus 4.8Anthropic Claude Sonnet 5Anthropic Claude Opus 5Anthropic Claude Haiku 4.5Anthropic GPT-5.5OpenAI GPT-5.6 SolOpenAI GPT RealtimeOpenAI Grok 4.6xAI Grok 4.5xAI Grok Voice Think Fast 2.0xAI GLM-5.2Z.ai Kimi K3Moonshot AI Qwen3.8-MaxAlibaba DeepSeek V4 ProDeepSeek DeepSeek V4-FlashDeepSeek
Report17
Qwen3.8-27B across the effort knob
Alibaba’s open-weights 27B runs the reasoning dial backward: 82.6% at low and medium, 73.9% at xhigh, and 2.4× slower for it. Far more robust than the Qwen3.8-Max flagship on the same knob.
2026-08-21 · 23 tasks · 129 runs · v3 suite
Report16
Grok 4.6 in Grok Build vs. Cursor vs. a bare-bones harness
The first three-way harness study: one model, three delivery systems. Inside xAI’s own harness the effort knob is monotone and ends at 92.8%, the best result this suite has recorded for Grok 4.6.
2026-08-18 · 23 tasks · 277 runs · v3 suite · Harness Study No. 02
Report15
Grok 4.6 in Cursor vs. a bare-bones harness
Our first model × harness study: the same Grok 4.6 through Cursor and through a deliberately minimal reference loop. The harness is worth 14.5 points at high effort and nothing measurable elsewhere; the agent’s 299 web-fetch attempts were all refused.
2026-08-16 · 23 tasks · 368 runs · v3 suite · Harness Study No. 01
Report14
Grok 4.6 across the effort knob
xAI’s newest model peaks mid-knob: 87.0% at medium, 73.9% at its shipped default, and the new xhigh level recovers less than half the drop. Failures shift from wrong answers to unfinished runs as effort rises.
2026-08-12 · 23 tasks · 92 runs · v3 suite
Report13
DeepSeek V4 Pro vs V4-Flash
Same 87.0% pass@1 at high effort. Pro uses 55% fewer tokens and runs 38% faster, but its higher token price makes the sweep 29% more expensive.
2026-08-09 · 23 tasks · 138 runs · v3 suite
Report12
Qwen3.8-Max across the effort knob
Alibaba’s flagship debuts on v3 with an effort knob that runs backwards: 81.2% at low, 55.1% at its shipped default, and every failure at higher effort is an unfinished run rather than a wrong answer.
2026-08-04 · 23 tasks · 202 runs · v3 suite
Report11
Grok Voice Think Fast 2.0 vs GPT Realtime
The same 200 questions, typed vs spoken: Grok pays +3.3 pp for hearing instead of reading, GPT Realtime +4.0; both lose ~10 pp on spoken arithmetic.
2026-07-31 · 200 questions · 1,840 units · voice-v1 suite
Report10
Claude Opus 5 across the effort knob
Opus 5 debuts on v3: its cheapest setting wins at 20/23 and $0.70/solved; score falls at every step up the knob while cost triples.
2026-07-26 · 23 tasks · 69 runs · v3 suite
Report09
Does contamination move Claude Opus 5’s score?
A controlled A/B: 13 tasks Opus 5 could have memorised vs 13 it cannot have seen. 11/13 vs 12/13, one genuine miss each — no detectable effect.
2026-07-25 · 26 tasks · 26 runs · contamination study
Report08
Kimi K3 vs Grok 4.5, Fable 5 & GPT-5.6 Sol
Kimi K3 joins the v3 board: Grok 4.5 keeps the lead at 91% and the lowest $/solved; K3 debuts at 74%, 87% with an extended budget.
2026-07-19 · 23 tasks · 28 Kimi runs · v3 suite
Report07
Grok 4.5 vs Fable 5 vs GPT-5.6 Sol
Grok leads at 91% from medium effort; only Sol climbs with the knob; two hard tasks go 0-for-27.
2026-07-12 · 23 tasks · 207 runs · v3 suite
Report06
Opus 4.8 vs Haiku 4.5
Five runs per task: the cheapest model wins on dependability; pass@1 hides coin flips.
2026-07-07 · 15 evals · 135 runs · reliability study
Report05
Fable 5 vs Opus 4.8
A four-way pass@1 tie; effort helps Opus and backfires for Fable.
2026-07-04 · 15 evals · 60 runs · frontier-hard tier
Report04
Fable 5 vs Opus 4.8 vs GPT-5.5
Four configurations, fifteen evals, each misses a different task.
2026-07-02 · 15 evals · 60 runs · corrected costs
Report03
Fable 5 vs Sonnet 5 vs GLM-5.2
Accuracy converges; efficiency spans an order of magnitude.
2026-07-01 · 10 tasks · 30 runs
Report02
Sonnet 5 vs Opus 4.8
Does more thinking effort pay off? Cost versus correctness across low, medium, high.
2026-06-30 · 52 tasks · 936 runs
Report01
Sonnet 5 vs Opus 4.8
The pilot that established cost as the live axis once correctness saturates.
2026-07-01 · 4 tasks · 24 runs