Model results · Alibaba

Qwen3.8-Max

Report12
Qwen3.8-Max reasoning-effort benchmark
Qwen3.8-Max debuts on VulcanBench v3 with an effort knob that runs backwards: 81.2% at low, 71.0% at medium, 55.1% at its shipped default. Every failure at medium and xhigh is an unfinished run, not a wrong answer.
2026-08-04

How would Qwen3.8-Max do on your stack?

Public results answer "which model is stronger on this suite." Whether Qwen3.8-Max belongs in your routing table depends on your languages, your codebases, and your task mix — that gets measured, not guessed. Grading rules are in the methodology.