Wednesday, July 29, 2026

Today’s Edition

AI Intel Report

MARKETS

Frontier Models

SpaceXAI Releases Grok Voice Think Fast 2.0 as Speech-to-Speech Leader

The successor model from SpaceXAI delivers benchmark gains in quality and reasoning while cutting token usage and latency, with a planned default rollout that positions it ahead of GPT-Realtime-2.1 and Gemini 3.1 Flash variants.

4 MIN READ
A wide angle view inside a spacious modern technology control center with floor to ceiling windows overlooking a desert landscape dotted with multiple large white circular Starlink satellite dishes mounted on low concrete pads connected by thick black cables running into the building. Inside the room several anonymous figures wearing plain dark business casual attire stand with their backs to the viewer at various sleek black desks equipped with arrays of flat rectangular hardware units glowing with soft blue indicator lights and compact microphone arrays on stands positioned near their mouths. One figure gestures toward a central rack of server like equipment while another holds a small handheld device near their face as if engaging in natural conversation. On the desks are multiple identical matte black rectangular voice interface modules with circular speaker grilles and subtle ventilation slots but no visible markings or displays. Through the windows the clear sky shows faint contrails and the distant horizon features additional Starlink ground terminals aligned in rows. The floor consists of polished light gray concrete with subtle cable management channels and the walls are lined with tall narrow cabinets housing additional hardware racks emitting faint status lights. Scattered across the workspace are loose bundles of fiber optic cables in neutral colors and a few plain notebooks lying closed beside the hardware. The overall environment conveys a functional high technology installation focused on satellite enabled real time voice processing with emphasis on the physical presence of communication hardware and anonymous operators interacting through speech interfaces. Additional details include rows of identical black equipment towers with ventilation grilles and indicator LEDs in the background a large wall mounted flat panel showing abstract waveform patterns without any text a water cooler station in the far corner and subtle reflections of the desert light on the glossy surfaces of the desks creating a cohesive real world operational scene representing advanced speech to speech model deployment integrated with satellite networks.
Illustration: AI Intel Report

Grok Voice Think Fast 2.0 is the successor speech-to-speech model from SpaceXAI that improves on version 1.0 across quality, reasoning, transcription, and efficiency metrics.

SpaceXAI launched Grok Voice Think Fast 2.0 as the successor to its initial voice model. The release focuses on measurable gains in overall quality and agentic performance while introducing efficiency improvements that reduce operational overhead. These changes support more natural and responsive voice interactions across multiple languages and environments.

What benchmark results does Grok Voice Think Fast 2.0 report?

The model achieves 82.9 percent on the Overall AA Speech-to-Speech Quality Index. This represents a 7.3 percentage point increase from the 75.7 percent score of Grok Voice Think Fast 1.0. The company attributes the gain to refinements in model training and architecture that enhance speech-to-speech fidelity. Additional evaluations show 97.2 percent on the Speech Reasoning Big Bench Audio assessment, indicating strong performance on audio-based reasoning tasks.

On the Tau Voice agentic benchmark the model records 56.5 percent. Time to first audio measures 0.70 seconds. These figures establish a performance baseline that the company presents as competitive within the current landscape of speech models.

How does Grok Voice Think Fast 2.0 compare to previous and competing models?

Transcription accuracy improves by a factor of 1.5 to 2.0 over dedicated models such as Deepgram Nova 3 and ElevenLabs Scribe v2. In noisy settings the improvement reaches up to 10 times better than prior dedicated systems. The model operates across 24 languages and delivers 1.4 times better transcription than Grok Voice Think Fast 1.0. It ranks number two overall, ahead of GPT-Realtime-2.1 and Gemini 3.1 Flash variants.

Comparison of key performance metrics between Grok Voice Think Fast versions.
BenchmarkGrok Voice Think Fast 2.0Grok Voice Think Fast 1.0
Overall AA Speech-to-Speech Quality Index82.9%75.7%
Speech Reasoning Big Bench Audio97.2%Not reported
Tau Voice agentic benchmark56.5%Not reported
Time to First Audio0.70sNot reported

What efficiency and feature enhancements are included?

Reasoning token usage drops to 0.4 times the amount required by Grok Voice Think Fast 1.0. The model supports tool calls during mid-speech in real-time conversations. This capability improves tool use reliability within ongoing voice exchanges. Low latency at 0.70 seconds to first audio further supports fluid conversational flow.

  1. The model supports improved tool use reliability during conversations.
  2. It enables mid-speech tool calls in real-time interactions.
  3. Reasoning token usage is reduced to 0.4 times previous levels.
  4. Transcription performance improves across 24 languages including noisy conditions.
Today, we're announcing Grok Voice Think Fast 2.0, our next-generation voice model with improved intelligence, transcription accuracy, and conversational capabilities.SpaceXAI

What is the pricing structure for the model?

The model is priced at 0.08 dollars per minute of audio. This rate corresponds to 4.80 dollars per hour for input processing. The pricing applies to API usage and is designed to support broad adoption across developer and enterprise applications.

When will the model become the default for Grok voice services?

On August 5, 2026, grok-voice-latest will transition from Grok Voice Think Fast 1.0 to Grok Voice Think Fast 2.0. The change will occur automatically for API users and those utilizing the Voice Agent Builder. No manual intervention is required for existing implementations to access the updated model.

What are the implications for developers and the voice AI market?

The combination of benchmark gains, reduced token consumption, and real-time tool integration offers developers a more efficient platform for building voice applications. Support for 24 languages broadens accessibility for global deployments. The default rollout simplifies adoption for users already integrated with the Grok voice ecosystem.

What developments are expected following this launch?

The model is positioned to serve as the standard for Grok voice interactions. Future iterations may extend the observed efficiency and reasoning improvements. SpaceXAI has framed the release as part of ongoing updates to its voice product line.

Frequently asked

What is the price of Grok Voice Think Fast 2.0?

The model is priced at $0.08 per minute of audio.

When does Grok Voice Think Fast 2.0 become the default model?

It becomes the default grok-voice-latest on August 5, 2026.

How does transcription performance compare to prior models?

Transcription accuracy improves 1.5 to 2 times over dedicated models and up to 10 times in noisy settings.

Sources

  1. SpaceXAI — Announcement of Grok Voice Think Fast 2.0 including benchmark scores, pricing, and rollout date.
  2. SpaceXAI — Company news updates on the introduction of the Grok Voice Think Fast 2.0 model.