Every published VulcanBench measurement of Qwen3.8-27B, Alibaba’s open-weights 27B, newest first: accuracy, cost, tokens, and time, graded by deterministic hidden tests on real tasks. It enters the Eval Suite 3 board at 82.6% pass@1 (best: medium).
Public results answer "which model is stronger on this suite." Whether Qwen3.8-27B belongs in
your routing table depends on your languages, your codebases, and your task mix — that gets measured, not
guessed. Grading rules are in the methodology.