Frontier Models
Google's Gemini 3.8 Live Models Lead in Real-Time Voice Agent Benchmarks
Released on September 15, 2026, the new Gemini audio models support background tool use and parallel reasoning, achieving leading positions on speech-to-speech quality and agentic benchmarks while integrating with Search and Workspace products.
Gemini 3.8 Live is Google's voice model for low-latency real-time dialogue and agent experiences that supports interleaved reasoning and asynchronous function calling.
Google released the Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking models on September 15, 2026. These models are available in the Gemini API and Google AI Studio. They bring advancements in near real-time reasoning to enable more effective voice agents. The base model is built for scale and cost efficiency while the Extended Thinking version targets high-complexity tasks. Gemini 3.5 Transcribe was released on August 26, 2026, to provide highly precise transcription across 85 plus languages. The models support a range of features that allow for more dynamic conversations including background tool use. Integration with Search Live shows immediate application in consumer facing products. Developers can begin using the models to build advanced voice applications with the provided tools. The watermarking with SynthID ensures that audio can be identified as AI generated for transparency purposes. This release builds on previous Gemini models by adding specific audio focused capabilities that address latency and reasoning challenges in voice systems.
What capabilities allow Gemini 3.8 Live to handle real-time conversations effectively?
The Gemini 3.8 Live model supports interleaved reasoning which permits the system to reason and respond in a continuous flow without creating noticeable pauses for the user. Asynchronous function calling allows tool use to occur without halting the audio output or breaking the conversational rhythm. Visual context can be incorporated by providing images as part of the input stream during ongoing dialogue. The model can transition between 97 languages during a single conversation without restarting the session or losing context. These features make it suitable for low-latency voice agent experiences and real-time dialogue in production environments. According to the model documentation from the Gemini API site, it serves as the default option for most such applications where speed is critical. The design avoids delays that reasoning might otherwise introduce in traditional setups. This results in more natural interactions for users who expect immediate responses. The model is optimized to handle the demands of ongoing conversations in various contexts including customer support and personal assistance scenarios.
In practice, these capabilities mean that voice agents can maintain context across multiple turns while performing background operations that enhance response quality. The support for visual context opens possibilities for multimodal interactions where users can show images while speaking to provide additional information. Language switching supports global users who may use multiple languages within the same session. The overall architecture emphasizes efficiency to keep costs low for high volume use cases. This positions the model as a practical choice for developers building production voice systems that require reliability. The release notes highlight its role in scaling voice applications across different industries without requiring extensive infrastructure changes.
How does the Extended Thinking model enable deeper reasoning during voice interactions?
Gemini 3.8 Live Extended Thinking enables background reasoning and asynchronous tool calls while streaming continuous audio responses to the user. This allows the model to handle high-complexity tasks without disrupting the flow of conversation or creating awkward silences. The model can perform parallel reasoning in the background to improve the quality and depth of responses over time. It leads in agentic task completion on the τ-Voice benchmark with a score of 68.6 percent according to data from the Google announcement. On the Sierra’s τ-Voice-banking benchmark, it scores 35.1 percent which indicates strong performance in financial and planning related scenarios. These results indicate strong performance in scenarios that require planning and tool use alongside spoken output. The model maintains the audio stream while processing complex queries in parallel threads. This addresses a key challenge in previous systems where thinking time caused noticeable pauses that reduced user satisfaction. The feature set makes it ideal for tasks that benefit from extended analysis such as troubleshooting or detailed recommendations.
The benchmark leadership demonstrates the effectiveness of the extended thinking approach in delivering high quality speech outputs. Developers can leverage this for applications that require both responsiveness and depth in their voice agents. The model is rolling out to Gemini Live and Google Workspace for subscribers who need advanced capabilities. This provides enterprise users with access to advanced reasoning in voice interfaces for productivity and collaboration tools. The combination of streaming audio and background processing represents a technical advancement in the field of real-time AI systems.
What features does Gemini 3.5 Transcribe offer for speech-to-text applications?
Gemini 3.5 Transcribe provides speech-to-text across 85 plus languages with smart transcription features like disfluency removal to clean up spoken input. It was released on August 26, 2026, as a dedicated model for transcription needs that complements the voice generation models. The average word error rate in streaming mode is 4.0 percent which supports high precision in real-time transcription scenarios. This low error rate supports high precision in real-time transcription scenarios where accuracy matters for downstream processing. The model complements the voice models by providing accurate text output from audio inputs for further analysis or logging. Features such as disfluency removal help in cleaning up transcripts for better readability and usability in business contexts. It can be used in conjunction with the other models for complete voice to text pipelines in agent workflows. The release expands the tools available for developers working on multilingual applications across global markets.
How do the pricing and product integrations work for the new models?
Pricing for Gemini 3.8 Live and Extended Thinking is set at $0.005 per minute for audio input and $0.018 per minute for output according to the developer tools announcement. This structure supports cost efficient deployment at scale for both small and large applications. Gemini 3.8 Live is powering Search Live on the Google app for users who interact with voice search features. The Extended Thinking variant is rolling out to Gemini Live and Google Workspace for subscribers seeking more advanced functionality. These integrations allow users to experience the new capabilities in familiar Google products without needing separate setup. Subscribers to Workspace can access the advanced features for their productivity tools including document and meeting assistance. The pricing model encourages experimentation by developers through the API with predictable costs. The availability in AI Studio provides an accessible entry point for testing before full integration.
| Model | Release Date | Key Capabilities | Top Benchmark | Languages |
|---|---|---|---|---|
| Gemini 3.8 Live | September 15, 2026 | Interleaved reasoning, async function calling, visual context | N/A | 97 |
| Gemini 3.8 Live Extended Thinking | September 15, 2026 | Background reasoning, async tool calls, streaming audio | 82.6 Speech to Speech Quality Index | 97 |
| Gemini 3.5 Transcribe | August 26, 2026 | Smart transcription, disfluency removal | 4.0% WER streaming | 85+ |
What are the market and stakeholder implications of these model releases?
The release of these models has implications for developers seeking to build sophisticated voice agents that can handle complex interactions. The combination of low latency and advanced reasoning opens new possibilities in customer service, personal assistants, and interactive applications across sectors. Stakeholders in the enterprise sector can benefit from the integration with Workspace tools for enhanced collaboration. The benchmark performance suggests competitive advantages in speech quality and task completion rates compared to prior offerings. Market observers may see increased adoption of voice interfaces as a result of these performance gains. The pricing makes it accessible for a range of project sizes from startups to large organizations. This could influence how companies approach AI voice solutions in their operations and customer engagement strategies.
For the broader AI community, these models demonstrate progress in handling real-time constraints while incorporating complex processing in the background. The support for multiple languages expands the potential user base globally and reduces barriers for non-English speakers. Integration with existing Google services provides a pathway for widespread use and rapid iteration by third parties. Developers are encouraged to explore the API for custom solutions tailored to specific industry needs. The focus on watermarking addresses concerns about AI generated content authenticity and supports responsible deployment practices.
- Sign up for access to the Gemini API or Google AI Studio to begin testing.
- Review the documentation for model parameters and capabilities including language support.
- Implement asynchronous function calling in voice agent code for background tool use.
- Test language transition features in sample conversations to verify seamless switching.
- Apply SynthID verification to generated audio outputs for compliance and traceability.
How have experts and Google representatives responded to the Gemini 3.8 Live announcement?
New Gemini audio models just dropped – 3.8 Live is now powering real-time conversations in Search Live.Rajan Patel, VP, Engineering for Search at Google
The statement from Rajan Patel highlights the immediate impact on Google Search through the integration of the new model. It underscores the practical deployment of the technology in a major product that reaches millions of users daily. Other reactions from the industry are likely to focus on the benchmark achievements and the technical innovations in reasoning. The models position Google competitively in the voice AI space against other providers. The emphasis on real-time performance addresses user expectations for seamless interactions in daily use.
What developments can be anticipated following this release?
Future updates may include further improvements in reasoning speed and accuracy based on user feedback and internal testing. Expansion of language support beyond the current 97 could occur to cover additional dialects and regions. Additional integrations with other Google services are possible as the models mature and demonstrate reliability. Developers may see more tools and examples for building with these models in the coming months. The field of voice agents is expected to evolve with these foundational releases serving as building blocks. Continued focus on efficiency and cost will likely remain a priority to encourage broad adoption. The success of these models could influence subsequent generations of Gemini audio capabilities in meaningful ways.
Overall, the release marks a notable advancement in making voice agents more capable and reliable for everyday applications. The technical features address key pain points in previous implementations such as latency during reasoning. With the provided benchmarks and pricing, the models offer a compelling option for new and existing projects seeking to incorporate voice. The combination of the base and extended versions caters to different needs from simple queries to complex agentic workflows. This approach allows for flexible application across various scenarios in both consumer and enterprise settings.
Frequently asked
When were the Gemini 3.8 Live models released and where can developers access them?
The Gemini 3.8 Live and Extended Thinking models were released on September 15, 2026, and are available through the Gemini API and Google AI Studio.
What benchmark scores did Gemini 3.8 Live Extended Thinking achieve?
Gemini 3.8 Live Extended Thinking achieved 82.6 on the Speech to Speech Quality Index, 68.6 percent on the τ-Voice benchmark, and 35.1 percent on the Sierra’s τ-Voice-banking benchmark.
Sources
- Google — Details on model capabilities, release dates, benchmark scores including 82.6 on Speech to Speech Quality Index, 68.6 percent on τ-Voice, and 35.1 percent on Sierra’s τ-Voice-banking benchmark for Gemini 3.8 Live and Extended Thinking.
- Google — Information on Gemini 3.5 Transcribe release date, 4.0 percent average word error rate, pricing of $0.005 per minute input and $0.018 per minute output, and language support for 85 plus languages.
- Google — Technical documentation on Gemini 3.8 Live features including interleaved reasoning, asynchronous function calling, visual context, 97 languages, and default use for low-latency voice agents.
- Search Engine Land — Quote from Rajan Patel on the release and integration with Search Live.
- @thecircuitry_ — Google released Gemini 3.8 Live and 3.8 Live Extended Thinking voice models in Gemini API and AI Studio, supporting background tool use and deeper reasoning. Also released Gemini 3.5 Transcribe for speech-to-text across…