Technical Report No. 08

Kimi K3 vs Grok 4.5, Fable 5 & GPT-5.6 Sol

2026-07-19 · 23 tasks · 28 Kimi runs · $33.89 Kimi spend · deterministic grading
Abstract.

Grok 4.5 leads all eleven configurations at 21/23 (91.3%) with the lowest cost per solved task ($0.32). Kimi K3 debuts at 17/23 (74%) in its standard configuration and 20/23 (87%) under an extended-budget ablation — level with the best Fable 5 and Sol configurations. The two hardest v3 tasks remain unsolved by every model.

Model card: Kimi K3 at 30-minute and 2-hour budgets vs Grok 4.5, Claude Fable 5, and GPT-5.6 Sol across effort levels on 23 v3 tasks
Card. All eleven configurations. Brown = Grok 4.5, blue = Claude Fable 5, gray = GPT-5.6 Sol, teal = Kimi K3; shades darken low → medium → high effort — for Kimi, 30-minute → 2-hour budget.

Grok 4.5 tops the field on accuracy and cost; Kimi K3 lands mid-pack.

Model (effort)Scorepass@1CostTokens/taskTime/task$/solved
Grok 4.5 (medium)21/2391.3%$6.67106 K3.9 min$0.32
Grok 4.5 (high)21/2391.3%$8.76141 K4.9 min$0.42
Claude Fable 5 (low)20/2387.0%$12.1828 K3.4 min$0.61
GPT-5.6 Sol (high)20/2387.0%$15.9085 K4.2 min$0.80
Claude Fable 5 (high)20/2387.0%$20.8244 K4.3 min$1.04
Kimi K3 (max, 2-h budget)20/2387.0%$27.43216 K28.3 min$1.37
Grok 4.5 (low)19/2382.6%$3.3957 K2.5 min$0.18
GPT-5.6 Sol (medium)19/2382.6%$8.8350 K2.8 min$0.46
Claude Fable 5 (medium)19/2382.6%$16.3236 K3.7 min$0.86
GPT-5.6 Sol (low)18/2378.3%$3.8523 K1.5 min$0.21
Kimi K3 (max, 30-min budget)17/2373.9%$16.84134 K18.1 min$0.99

2-hour row: the five tasks that did not finish within the 30-minute budget were re-run with exact budgets of 2 h / 400 steps / $4 per run (runs marked budget_override); the other 18 runs are unchanged. It is an ablation, not a leaderboard config — every other model ran under the standard budget. Fable 5 ran with Opus 4.8 refusal fallback (see Report 07). Grok/Fable/Sol rows reproduced from Report 07. Tokens are billed tokens; time is sandbox wall-clock.

1.

Grok 4.5 wins on accuracy and economics. 21/23 (91.3%) at $0.32 per solved task — no other configuration matches either number. Its medium and high efforts tie, keeping medium the efficiency frontier of the board.

2.

Kimi K3 debuts mid-pack and closes with budget. 17/23 in its standard configuration; the extended-budget ablation converts aiohttp, NetworkX-Leiden, and jiff-strftime for +$10.59, reaching 20/23 — matching the best Fable 5 and Sol configurations.

3.

The v3 difficulty ceiling held. PennyLane Trotter fragmentation (0.2 partial) and SQLGlot internal-name canonicalization (0.17 partial) remain unsolved by every configuration ever run — 0-for-27 in Report 07, and K3’s extended attempts didn’t crack them either.

Three more solves under the extended budget.

Task30-min budget2-h budgetTimeCost
aiohttp-upgrade-deferred× did not finish✓ solved64 min$2.66
networkx-leiden-communities× did not finish✓ solved107 min$4.35
jiff-strftime-negpad× did not finish✓ solved21 min$1.14
pennylane-trotter-fragmented× did not finishpartial (0.2), $4 cap100 min$4.78
sqlglot-canonicalize-internal-names× did not finishpartial (0.17), $4 cap94 min$4.12
sqlglot-iso8601-nanospartial (0.7)— not re-run

The two remaining tasks ended at the $4 per-run cost cap with partial progress.

Same 23 v3 tasks as Report 07 (real merged post-cutoff PRs across Python, Rust, TypeScript, JavaScript, Go), graded by deterministic hidden tests in a network-isolated Docker sandbox, one attempt per task per configuration. kimi-k3 via the Moonshot API ($3 / $15 per million tokens, cache-hit input billed here at the cache-miss rate), run in its single available API configuration. The extended-budget ablation used the --override-budgets flag, which stamps runs so they cannot be mixed into standard comparisons.

← All benchmarks