Frontier Models
SpaceXAI Releases Grok Voice Think Fast 2.0 as Speech-to-Speech Leader
The successor model from SpaceXAI delivers benchmark gains in quality and reasoning while cutting token usage and latency, with a planned default rollout that positions it ahead of GPT-Realtime-2.1 and Gemini 3.1 Flash variants.
Grok Voice Think Fast 2.0 is the successor speech-to-speech model from SpaceXAI that improves on version 1.0 across quality, reasoning, transcription, and efficiency metrics.
SpaceXAI launched Grok Voice Think Fast 2.0 as the successor to its initial voice model. The release focuses on measurable gains in overall quality and agentic performance while introducing efficiency improvements that reduce operational overhead. These changes support more natural and responsive voice interactions across multiple languages and environments.
What benchmark results does Grok Voice Think Fast 2.0 report?
The model achieves 82.9 percent on the Overall AA Speech-to-Speech Quality Index. This represents a 7.3 percentage point increase from the 75.7 percent score of Grok Voice Think Fast 1.0. The company attributes the gain to refinements in model training and architecture that enhance speech-to-speech fidelity. Additional evaluations show 97.2 percent on the Speech Reasoning Big Bench Audio assessment, indicating strong performance on audio-based reasoning tasks.
On the Tau Voice agentic benchmark the model records 56.5 percent. Time to first audio measures 0.70 seconds. These figures establish a performance baseline that the company presents as competitive within the current landscape of speech models.
How does Grok Voice Think Fast 2.0 compare to previous and competing models?
Transcription accuracy improves by a factor of 1.5 to 2.0 over dedicated models such as Deepgram Nova 3 and ElevenLabs Scribe v2. In noisy settings the improvement reaches up to 10 times better than prior dedicated systems. The model operates across 24 languages and delivers 1.4 times better transcription than Grok Voice Think Fast 1.0. It ranks number two overall, ahead of GPT-Realtime-2.1 and Gemini 3.1 Flash variants.
| Benchmark | Grok Voice Think Fast 2.0 | Grok Voice Think Fast 1.0 |
|---|---|---|
| Overall AA Speech-to-Speech Quality Index | 82.9% | 75.7% |
| Speech Reasoning Big Bench Audio | 97.2% | Not reported |
| Tau Voice agentic benchmark | 56.5% | Not reported |
| Time to First Audio | 0.70s | Not reported |
What efficiency and feature enhancements are included?
Reasoning token usage drops to 0.4 times the amount required by Grok Voice Think Fast 1.0. The model supports tool calls during mid-speech in real-time conversations. This capability improves tool use reliability within ongoing voice exchanges. Low latency at 0.70 seconds to first audio further supports fluid conversational flow.
- The model supports improved tool use reliability during conversations.
- It enables mid-speech tool calls in real-time interactions.
- Reasoning token usage is reduced to 0.4 times previous levels.
- Transcription performance improves across 24 languages including noisy conditions.
Today, we're announcing Grok Voice Think Fast 2.0, our next-generation voice model with improved intelligence, transcription accuracy, and conversational capabilities.SpaceXAI
What is the pricing structure for the model?
The model is priced at 0.08 dollars per minute of audio. This rate corresponds to 4.80 dollars per hour for input processing. The pricing applies to API usage and is designed to support broad adoption across developer and enterprise applications.
When will the model become the default for Grok voice services?
On August 5, 2026, grok-voice-latest will transition from Grok Voice Think Fast 1.0 to Grok Voice Think Fast 2.0. The change will occur automatically for API users and those utilizing the Voice Agent Builder. No manual intervention is required for existing implementations to access the updated model.
What are the implications for developers and the voice AI market?
The combination of benchmark gains, reduced token consumption, and real-time tool integration offers developers a more efficient platform for building voice applications. Support for 24 languages broadens accessibility for global deployments. The default rollout simplifies adoption for users already integrated with the Grok voice ecosystem.
What developments are expected following this launch?
The model is positioned to serve as the standard for Grok voice interactions. Future iterations may extend the observed efficiency and reasoning improvements. SpaceXAI has framed the release as part of ongoing updates to its voice product line.
Frequently asked
What is the price of Grok Voice Think Fast 2.0?
The model is priced at $0.08 per minute of audio.
When does Grok Voice Think Fast 2.0 become the default model?
It becomes the default grok-voice-latest on August 5, 2026.
How does transcription performance compare to prior models?
Transcription accuracy improves 1.5 to 2 times over dedicated models and up to 10 times in noisy settings.