Every published VulcanBench measurement of GLM 5.3, Z.ai’s flagship coding model, newest first: accuracy, cost, tokens, and time, graded by deterministic hidden tests on real tasks. It enters the Eval Suite 3 board at 78.3% pass@1 through its raw API (best: low); the same report measures it inside Z.ai’s own ZCode harness, where the identical weights reach 87.0% at max effort.
Public results answer "which model is stronger on this suite." Whether GLM 5.3 belongs in
your routing table depends on your languages, your codebases, and your task mix: that gets measured, not
guessed. Grading rules are in the methodology.