Frontier Models
Gemini 3.8 Flash Delivers Frontier Agentic Performance at Flash Pricing
Google's latest Flash model targets long-horizon software engineering and autonomous agents with enhanced reasoning, a one-million-token context window, and introductory pricing available through multiple enterprise platforms until the end of 2026.
Gemini 3.8 Flash is Google's most intelligent Flash model engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows.
Google launched Gemini 3.8 Flash on September 2, 2026. The release marks the third Flash model in six weeks and builds directly on the 3.7 version released three weeks earlier. The model focuses on sustained reasoning for software engineering and agent tasks. It maintains the speed and cost profile of prior Flash releases while increasing intelligence.
Availability extends across the Gemini API with model ID gemini-3.8-flash. Additional access points include Google AI Studio, the Gemini app for Pro and Ultra subscribers, AI Mode in Google Search, Gemini in Google Sheets, Gemini Enterprise, and Google Antigravity. This broad distribution supports production use in varied environments from individual developers to large organizations.
Background on the Gemini Flash Series
The Flash series has accelerated in development pace over recent months. Each release refines reasoning and coding capabilities while preserving low latency and cost. The 3.7 Flash established baseline performance for agentic workflows before the 3.8 iteration advanced those metrics further.
Google DeepMind emphasizes efficient models that deliver high utility without scaling compute proportionally. This strategy addresses enterprise demand for reliable agents that handle extended tasks. The series differentiates from larger frontier models by prioritizing cost efficiency alongside performance gains.
The Fairwind Program introduces the Gemini 3.8 Flash Cyber variant alongside the standard release. This variant targets trusted defenders with specialized capabilities in vulnerability detection and patching. It extends the core model for security-focused applications.
What's New in Gemini 3.8 Flash
The model introduces greater diligence on complex tasks through additional reasoning steps and iterative tool calls. This design choice improves outcomes on long-horizon problems compared to the predecessor. Evaluations indicate completion of more than three times as many tasks as Gemini 3.7 Flash in document-heavy scenarios.
Pricing remains at the Flash level during an introductory period. The rate holds through December 31, 2026, before adjusting to standard levels. This window allows organizations to test capabilities at reduced cost before committing to ongoing use.
The Cyber variant adds targeted security features under the Fairwind Program. It maintains the same core architecture while optimizing for defender workflows. Performance in vulnerability management reaches frontier levels according to internal assessments.
Technical Specifications
The context window extends to 1 million tokens to support extended documents and conversations. Maximum output reaches 64,000 tokens for detailed responses. Thinking levels adjust across low, medium, and high settings to balance depth and speed.
Multimodal inputs accept text, images, video, audio, and PDF files. This range enables processing of diverse data types in enterprise workflows. The architecture supports sustained operation across long-running agent sessions without degradation.
| Aspect | Gemini 3.7 Flash | Gemini 3.8 Flash |
|---|---|---|
| Launch Date | Mid-August 2026 | September 2, 2026 |
| Context Window | Standard | 1 million tokens |
| Max Output Tokens | Not specified | 64,000 |
| Input Token Price (Intro) | Previous rate | $0.75 per million |
| Output Token Price (Intro) | Previous rate | $3.75 per million |
| DeepSWE v1.1 Score | Baseline | 73.7% |
| Thinking Modes | Fixed | Tunable low/medium/high |
| Multimodal Inputs | Limited | Text, image, video, audio, PDF |
These specifications enable deployment in agent platforms where context retention and output length matter. Tunable thinking provides control over computational effort per query. The combination supports both rapid responses and deep analysis as needed.
Performance on Benchmarks
Google reports strong results on the DeepSWE v1.1 Long-Horizon Software Engineering benchmark. The score reflects improved handling of extended coding and reasoning chains. Sustained performance distinguishes the model in evaluations of real-world tasks.
These performance gains stem from a core design choice: 3.8 Flash works harder. On complex tasks, it exhibits greater diligence — executing extra reasoning steps, and calling tools iteratively.Tulsee Doshi and Raluca Ada Popa, Senior Director, Product Management, Google and Gemini Security Lead, Google DeepMind
Market and Stakeholder Implications
Enterprises gain access to advanced agentic tools at accessible pricing. The introductory rates reduce barriers for testing in production environments. Organizations can integrate the model into existing workflows without immediate budget strain.
Developers benefit from the broad platform availability. Integration through the Gemini API and Google AI Studio simplifies adoption. The model supports both prototyping and scaled deployment across the Gemini Enterprise Agent Platform.
- Assess current workflows for long-horizon tasks that could benefit from enhanced reasoning.
- Test the model in Google AI Studio to evaluate fit for specific use cases.
- Plan for the pricing transition after December 31, 2026, to budget accordingly.
- Explore the Cyber variant if security and vulnerability management are priorities.
- Monitor updates from Google DeepMind on further model improvements.
Expert Reactions
Thai Tran, AI Product Lead at Glean, highlighted performance in document-heavy environments. The model converts complex requests into finished artifacts more reliably than prior versions. This capability addresses practical needs in enterprise knowledge management.
Industry observers note the balance of speed and increased diligence. The iterative tool use improves reliability on multi-step problems. Feedback centers on the model's suitability for autonomous agent applications.
Gemini 3.8 Flash excels at long-running, document-heavy workflows, completing more than three times as many tasks as Gemini 3.7 Flash in our evaluations. We're excited to bring its sustained reasoning capabilities to Glean customers who need to turn complex requests into finished artifacts.Thai Tran, AI Product Lead, Glean
What's Next
Google plans continued iteration on the Flash series with emphasis on agent capabilities. Future releases may refine the tunable thinking mechanisms and multimodal handling. The strategy maintains focus on cost-efficient frontier performance.
Expansion of the Fairwind Program could introduce additional specialized variants. Deeper integration with the Gemini Enterprise Agent Platform will likely follow. Organizations should evaluate the current model during the introductory pricing window.
The launch reinforces Google's position in accessible frontier models. Sustained releases at this cadence suggest ongoing improvements in reasoning depth and workflow support. Stakeholders across development and enterprise sectors stand to gain from the expanded options.
Frequently asked
How does Gemini 3.8 Flash differ from Gemini 3.7 Flash in performance?
Gemini 3.8 Flash achieves higher scores on long-horizon benchmarks and completes more than three times as many tasks in document-heavy evaluations. It introduces tunable thinking levels and greater iterative reasoning on complex problems.
What is the pricing structure for Gemini 3.8 Flash?
The model launches at an introductory rate of $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026. Standard pricing of $1.50 per million input and $7.50 per million output applies afterward.
Which platforms support Gemini 3.8 Flash?
Access occurs through the Gemini API, Google AI Studio, the Gemini app for Pro and Ultra subscribers, AI Mode in Google Search, Gemini in Google Sheets, Gemini Enterprise, and Google Antigravity.
Sources
- Google — Building on the momentum of 3.7 Flash from three weeks ago and marking our third Flash release in only six weeks, today we’re introducing Gemini 3.8, our best reasoning and coding model yet, at the same speed and low cost of 3.7.
- Google — Gemini 3.8 Flash (gemini-3.8-flash) is generally available (GA) and ready for production use. It is our most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows.
- Google DeepMind — Our most intelligent workhorse model yet for coding and agents.