Model results · Alibaba

Qwen3.8-27B

Report17
Qwen3.8-27B across the effort knob
The open model’s reasoning dial runs flat then backward: 82.6% at low and medium, 73.9% at xhigh, and 2.4× slower for the worse score. Far more robust across the knob than the Qwen3.8-Max flagship, and it matches the flagship at low effort.
2026-08-21 · 23 tasks · 129 runs · v3 suite

How would Qwen3.8-27B do on your stack?

Public results answer "which model is stronger on this suite." Whether Qwen3.8-27B belongs in your routing table depends on your languages, your codebases, and your task mix — that gets measured, not guessed. Grading rules are in the methodology.