Frontier Models
Gemini 3.8 Live Advances Google's Real-Time Voice AI With Visual and Search Features
Google releases Gemini 3.8 Live for immediate developer use via API and AI Studio, powering Search Live with camera input, background lookups, and 97-language support while the Extended Thinking variant adds multi-step reasoning.
Gemini 3.8 Live is Google's most advanced live dialogue model that combines conversational intelligence with fluid dialogue and visual grounding.
The release of Gemini 3.8 Live introduces a native audio model that processes voice input and output in a single stream while incorporating visual context from camera feeds. This design eliminates the need for separate transcription and synthesis stages that often introduce delays in earlier systems. The model maintains conversation continuity even when performing background operations such as web searches. Users receive answers that include direct web links for further exploration. Language switching occurs without restarting the session, supporting 97 languages in total. These elements combine to produce responses that feel more integrated and responsive during live exchanges.
What background and context surround the release of Gemini 3.8 Live?
Google developed Gemini 3.8 Live to address constraints in cascaded voice pipelines that separate speech recognition from language modeling and output generation. Previous approaches required multiple handoffs between components, leading to higher latency and reduced coherence when users introduced visual or search elements mid-conversation. The new model integrates these functions natively, allowing interleaved reasoning that occurs without pausing audio output. Extended Thinking adds parallel background processing for tasks that require multiple steps. Both variants draw from the same core architecture but differ in reasoning depth and resource allocation. The immediate availability through existing developer platforms reflects an emphasis on production deployment rather than research previews.
Search Live integration demonstrates how the model operates at scale for consumer-facing applications. Real-time responses now include grounding from current web data without disrupting dialogue flow. The pricing model of $0.005 per minute for input and $0.018 per minute for output supports cost-effective use in extended sessions. Asynchronous function calling permits tool interactions to run in parallel, preserving conversational momentum. Full session client content updates enable dynamic adjustments to context during ongoing calls. These technical choices position the model for enterprise and consumer voice agents that demand both speed and accuracy.
What new capabilities does Gemini 3.8 Live introduce in detail?
Gemini 3.8 Live accepts camera input to provide visual grounding during conversations, allowing the model to reference objects or scenes the user shows in real time. Background web lookups occur through search grounding, delivering up-to-date information while the audio stream continues uninterrupted. Answers can embed web links so users can verify or expand on the provided information. Mid-chat language switching lets participants change languages without ending the session or losing prior context. The model handles these transitions across 97 languages with maintained fluency. Native audio output ensures the response voice matches the input language and tone without additional conversion steps.
The Extended Thinking variant extends these capabilities for longer tasks by running parallel reasoning processes while still streaming audio responses. This allows handling of multi-step problems such as planning sequences or analyzing complex queries without forcing the user to wait. Asynchronous function calling supports API calls that execute in the background, enabling the model to gather data from external tools while the conversation proceeds. Interleaved reasoning keeps the primary dialogue responsive even during intensive computations. These additions target use cases where depth and continuity must coexist.
What are the technical specifics of the model variants?
Gemini 3.8 Live serves as the default option for low-latency voice agent experiences where reasoning delays must be avoided. It supports interleaved reasoning that weaves task execution into the ongoing dialogue. Search grounding integrates directly into the response generation process. The Extended Thinking variant increases intelligence and multi-step reasoning capacity for high-complexity tasks. Both models feature built-in audio streaming and full session client content updates that allow real-time modifications to the conversation context. The architecture supports native audio output without intermediate text stages.
| Feature | Gemini 3.8 Live | Gemini 3.8 Live Extended Thinking |
|---|---|---|
| Core Strength | Real-time low-latency dialogue | High-complexity multi-step reasoning |
| Reasoning Mode | Interleaved reasoning | Parallel background reasoning |
| Task Handling | Fluid conversations with visual context | Longer tasks with streaming audio responses |
| Performance Metric | Top speech quality index position | 68.6 percent on τ-Voice benchmark |
| Availability | Immediate via API and AI Studio | Immediate via API and AI Studio |
Pricing remains consistent across both variants at $0.005 per minute for audio input and $0.018 per minute for audio output. This structure facilitates predictable budgeting for applications that involve extended voice sessions. The models also enable full session client content updates, which permit external systems to modify context without restarting the audio stream. These specifications derive from the native speech-to-speech design that reduces overhead compared with cascaded alternatives.
How does this affect the market and stakeholder implications?
Developers gain immediate access to production-ready voice tools through the Gemini API and Google AI Studio, reducing the time required to build and deploy real-time agents. The combination of visual context and search grounding enables applications in areas such as customer support, education, and personal assistance where users benefit from multimodal input. The 97-language coverage expands potential reach to global audiences without requiring separate model deployments for each language. Pricing at $0.005 per minute input supports high-volume usage in commercial settings.
Enterprise stakeholders can integrate asynchronous function calling to connect voice agents with existing backend systems without interrupting user conversations. Search Live users receive responses that include web links and maintain natural flow across language switches. The Extended Thinking variant targets scenarios requiring deeper analysis while preserving audio continuity. These changes lower barriers for organizations seeking to replace legacy voice systems with more capable alternatives.
What are expert reactions to the release?
New Gemini audio models just dropped – 3.8 Live is now powering real-time conversations in Search Live. You’ll get: - More helpful responses, complete with web links to dive deeper - Fluid multilingual support (you can switch languages mid convo) - More natural, free-flowing interactionsRajan Patel, VP, Engineering for Search and Co-founder of Google Lens
The statement from Rajan Patel emphasizes practical benefits for Search Live users, including web links and language flexibility. These features align with the technical capabilities described in the model documentation, confirming that production deployment has already begun. The focus on natural interactions reflects engineering priorities around reducing friction in voice exchanges.
What comes next for Gemini voice models?
Continued iteration on the Extended Thinking variant may expand the range of multi-step tasks that can run in parallel with live audio. Developers are expected to explore new agentic applications that leverage asynchronous function calling for background data retrieval. Integration with additional Google services could further enhance visual grounding and search capabilities. The current pricing and access model provide a foundation for scaling these features across more use cases.
- Access the models through Google AI Studio for initial testing and prototyping.
- Integrate via the Gemini API for production voice applications.
- Enable camera input to add visual context during live sessions.
- Configure search grounding to include web links in responses.
- Activate asynchronous function calling for background tool use without dialogue interruption.
Frequently asked
How can developers begin using Gemini 3.8 Live today?
Developers access Gemini 3.8 Live and the Extended Thinking variant immediately through Google AI Studio for experimentation and the Gemini API for production integration.
What performance metrics distinguish the Extended Thinking variant?
Gemini 3.8 Live Extended Thinking achieves an 82.6 score on the Speech to Speech Quality Index and 68.6 percent on the τ-Voice benchmark for agentic tasks.
Does the model support language changes during a single conversation?
Yes, Gemini 3.8 Live enables mid-chat language switching across 97 languages while maintaining context and response quality.
Sources
- Google — Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking are our most advanced live dialogue models yet. ... Gemini 3.8 Live : Built for scale and cost efficiency, combining conversational intelligence with fluid dialogue and visual grounding. Gemini 3.8 Live Extended Thinking : Built for high-complexity tasks, with increased intelligence and multi-step reasoning.
- Google — Gemini 3.8 Live brings a step change to our native speech-to-speech models, capable of performing tasks while maintaining dialogue. Key capabilities include: Asynchronous function calling... Visual context... Multilingual support: Reach global audiences with coverage for 97+ languages
- Search Engine Roundtable — New Gemini audio models just dropped – 3.8 Live is now powering real-time conversations in Search Live. You’ll get: - More helpful responses, complete with web links to dive deeper - Fluid multilingual support (you can switch languages mid convo) - More natural, free-flowing interactions
- Google — Gemini 3.8 Live is the default option for most low-latency voice agent experiences and real-time dialogue without reasoning-induced delays. It supports interleaved reasoning, asynchronous function calling... Search grounding Supported
- Google — Gemini 3.8 Live is the default option for most low-latency voice agent experiences and real-time dialogue without reasoning-induced delays. It supports interleaved reasoning, asynchronous function calling... Search…