Model results · OpenAI

GPT Realtime

Report11
Grok Voice Think Fast 2.0 vs GPT Realtime voice benchmark
First results from the VulcanBench Voice Eval Suite: the same 200 questions typed vs spoken. Grok Voice Think Fast 2.0 pays a +3.3 pp voice tax (99.0% text, 95.7% audio); GPT Realtime +4.0 pp. Both lose ~10 pp on spoken arithmetic.
2026-07-31

How would GPT Realtime do on your stack?

Public results answer "which model is stronger on this suite." Whether GPT Realtime belongs in your routing table depends on your languages, your codebases, and your task mix — that gets measured, not guessed. Grading rules are in the methodology.