SpaceXAI just dropped Grok Voice Think Fast 2.0, the successor to its April 2026 flagship voice model, and the numbers are hard to ignore. On the Artificial Analysis Speech-to-Speech Index, the High variant debuts at #2 with an 82.9% composite score , up 7.3 percentage points from its predecessor , and takes the #1 spot on Tau Voice, the agentic benchmark that actually matters for production deployments.
The latency gap nobody else has closed
The headline number is 0.70 seconds time-to-first-audio. That makes Grok Voice Think Fast 2.0 the only model in the top five of the Speech-to-Speech Index under 1 second . For context: GPT-Realtime-2 High comes in at 1.14s, GPT-Realtime-2.1 High at 1.21s, and the overall index leader Qwen Audio 3.0 Realtime Plus at 4.02s . Sub-second response time is the threshold where a voice agent starts to feel like a real conversation rather than a walkie-talkie.
This is not just a speed story, though. The 1.0 model already had low latency , 1.25s , but its conversational dynamics score was a weak point. On the Full Duplex Bench subset, the model scores 95.1%, up from 77.8% for its predecessor , which was the largest single driver of the index gain. Full Duplex Bench (FDB) measures the things that make or break a real phone call: knowing when to stop talking, how to handle a user who cuts in mid-sentence, and whether the model correctly ignores filler words like "yeah" or "uh-huh" without treating them as turn-taking signals.