Frontier Models
Gemini 4 Argon: Google DeepMind Retakes Frontier Lead With 1M-Output-Token Model
Google DeepMind's Sept. 30 launch pairs 1M-token outputs and $2/$10 pricing with benchmark wins in coding and cyber — but access starts inside the Fairwind Program.
Gemini 4 Argon is Google DeepMind's most powerful proprietary frontier model, announced Sept. 30, 2026, with a 1 million-token context window that supports up to 1 million output tokens through a capability the company calls Long Decode Continuation.
Google DeepMind announced Gemini 4 Argon on Sept. 30, 2026, positioning it as the company's most powerful proprietary model and its answer to OpenAI's GPT-6 Astra and Anthropic's Claude Opus 5.5. The initial rollout is limited to a set of trusted cyber defenders through the Fairwind Program, Google said, with access to widen over time. The model posts the highest score, or ties for the highest score, in more benchmark categories than either rival, according to VentureBeat's coverage of the launch.
The release resets a frontier race that had recently tilted toward OpenAI and Anthropic. Gemini 4 Argon enters with an introductory price of $2 per million input tokens and $10 per million output tokens, cached input at 95% off, and a 1 million-token context window that also supports up to 1 million output tokens, per Google. Those specifications, combined with top scores in coding and cybersecurity evaluations, give Google a credible claim to the lead it ceded in earlier Gemini generations.
What is Gemini 4 Argon, and when did Google DeepMind launch it?
Gemini 4 Argon is the latest entry in Google DeepMind's Gemini series and the company's most powerful proprietary model to date. Google announced it Sept. 30, 2026, and described the model as supporting a 1 million-token context window with up to 1 million output tokens, a capability the company attributes to a mechanism called Long Decode Continuation. That mechanism lets the model generate very long outputs in a single decoding pass rather than assembling them from multiple shorter generations.
The model accepts text, image, video and speech as input and produces text output, according to Google DeepMind's model documentation. The company says Argon delivers frontier performance across real-world software engineering, enterprise knowledge work in legal and finance, and cybersecurity defense. Those are the workloads Google is prioritizing as it stages the release.
How does Gemini 4 Argon perform on coding, cybersecurity and knowledge-work benchmarks?
Google reported a 77.9% score on DeepSWE v1.1, its benchmark for long-horizon software engineering, and called the result a new state of the art. On the Vals Index, which measures economic impact in knowledge work, Google DeepMind reported 68.9%. On AutomationBench, an evaluation of end-to-end business execution, Google reported 51.3%.
77.9% — Gemini 4 Argon's score on DeepSWE v1.1 for long-horizon software engineering, a new state of the art, per Google.
Independent results align with the company's claims in key areas. Artificial Analysis gave Gemini 4 Argon a score of 53 on its Intelligence Index, tied with GPT-6 Astra. VentureBeat reported that Argon's attack success rate on the Gray Swan IPI prompt-injection benchmark was 0.7%, the lowest among peers, a sign of comparatively strong resistance to adversarial prompt manipulation.
| Benchmark | Result | Source |
|---|---|---|
| DeepSWE v1.1 (long-horizon software engineering) | 77.9% | |
| Vals Index (knowledge-work economic impact) | 68.9% | Google DeepMind |
| AutomationBench (end-to-end business execution) | 51.3% | |
| LVBench | 91.7% | Google DeepMind |
| Artificial Analysis Intelligence Index | 53 — tied with GPT-6 Astra | Artificial Analysis |
| Gray Swan IPI (prompt-injection attack success rate) | 0.7% — lowest among peers | VentureBeat |
What does the 1 million-token output capability mean for real-world workflows?
The output ceiling is the technical differentiator. Most frontier models cap generation at tens of thousands of tokens, forcing agents to chain multiple calls when a task requires a long artifact. Argon's 1 million-token output allows a single run to produce an entire refactored codebase, a full security assessment or a long-horizon agent trajectory, which Google says is the point of the design.
Google described Long Decode Continuation as the mechanism behind the 1 million-token output ceiling, and the capability is central to Argon's agentic positioning. Long-horizon tasks — multi-file code changes, extended security analyses, multi-step business workflows — historically required agents to plan, execute and stitch together many model calls, multiplying cost and failure points. A single long decode changes that economics, which is why Google is steering the first deployments toward those workloads.
Google DeepMind frames the capability around specific users: cyber defenders generating defensive playbooks, engineers executing multi-file changes and enterprise teams running legal and finance analysis. The model's text-only output, despite multimodal input, keeps the deployment surface simpler while the agentic features mature.
- Trusted cyber defenders admitted to the Fairwind Program, starting Sept. 30, 2026
- Long-horizon coding and software engineering workflows
- Cybersecurity defense applications
- Enterprise knowledge work in legal and finance
- Broader availability “as soon as possible,” per Logan Kilpatrick of Google AI Studio
How is Google pricing Gemini 4 Argon, and how does that compare with rivals?
Google set introductory pricing at $2 per million input tokens and $10 per million output tokens, with cached input tokens at $0.10 per million, a 95% discount off the input price. Standard rates rise to $4 per million input and $20 per million output after the discount period, Google said. The company did not disclose when the introductory period ends.
The cached-input price matters for exactly the agent economics Google is courting. Agents re-read context on every step, so a 95% discount on cached tokens can cut the effective cost of long-running workflows dramatically. The rate card is structured to make Argon attractive precisely where rivals' pricing penalizes context-heavy agent loops.
The pricing is aggressive for a model at the top of the benchmark tables, and it signals an intent to win usage volume early. Logan Kilpatrick, who leads Google AI Studio and the Gemini API, announced the rates on X on launch day, writing that Argon is “priced at $2 in and $10 out during introductory pricing” and that he is “really excited by the progress we have made here.”
Why is Gemini 4 Argon's rollout limited to the Fairwind Program?
The Fairwind Program is Google's channel for giving vetted users early access to high-capability models before broad availability. Google is using it to put Argon in the hands of trusted cyber defenders first, with coding and long-horizon workflows next. The stated rationale is safety: the model's long outputs and agentic capabilities raise the stakes of misuse, so Google says it is testing with experienced operators before scaling access.
Safely releasing frontier capabilities at this level requires a phased approach.Koray Kavukcuoglu, SVP, Google DeepMind and Chief AI Architect, Google
Koray Kavukcuoglu, Google DeepMind's chief AI architect, gave that explanation in Google's launch announcement. The phased approach follows a pattern OpenAI and Anthropic have also used for their frontier releases, but it carries commercial risk for Google: while Argon is gated, GPT-6 Astra and Claude Opus 5.5 remain broadly accessible to developers.
How do independent analysts assess Gemini 4 Argon against GPT-6 Astra and Claude Opus 5.5?
Artificial Analysis gave Argon an Intelligence Index score of 53, tying GPT-6 Astra. VentureBeat concluded that Argon posts the highest score, or ties for the highest score, in more categories than GPT-6 Astra or Claude Opus 5.5, while noting that the release is limited.
The gap between benchmark results and deployment reality is the main caveat. Google DeepMind's model page emphasizes complex workflows in software engineering, legal, finance and cybersecurity, a pitch aimed at buyers evaluating agents rather than raw text generation. Whether Argon's benchmark dominance survives real-world traffic will depend on how quickly Google widens the Fairwind rollout.
What comes next for Gemini 4 Argon and Google's frontier roadmap?
Google has not announced a date for general availability. The Fairwind rollout gives the company a window to observe failures and refine safety mitigations before the model faces the full traffic that GPT-6 Astra and Claude Opus 5.5 already handle. In the meantime, the discount pricing and benchmark scores give Google a strong pitch to enterprise customers willing to apply for access.
For OpenAI, the Intelligence Index tie is evidence that its lead is not durable. For Anthropic, the DeepSWE 77.9% result is a direct challenge to Claude Opus 5.5's standing in coding evaluations. For enterprises, the near-term question is simpler: whether Argon's limited availability and text-only output justify building on a model that is not yet generally available.
Frequently asked
When was Gemini 4 Argon released?
Google DeepMind announced Gemini 4 Argon on Sept. 30, 2026, beginning a limited rollout to trusted cyber defenders through the Fairwind Program rather than a general release.
How much does Gemini 4 Argon cost per token?
Introductory pricing is $2 per million input tokens and $10 per million output tokens, with cached input at $0.10 per million. Standard rates rise to $4 and $20 after the discount period, per Google.
Which benchmarks does Gemini 4 Argon lead?
Google reported 77.9% on DeepSWE v1.1, 68.9% on the Vals Index and 51.3% on AutomationBench. Artificial Analysis tied it with GPT-6 Astra at 53 on its Intelligence Index, and VentureBeat reported a 0.7% prompt-injection attack success rate on Gray Swan IPI.
Why is Gemini 4 Argon's rollout limited?
Google is releasing the model first to vetted users in the Fairwind Program, prioritizing coding, cybersecurity defense and long-horizon workflows. Koray Kavukcuoglu said safely releasing frontier capabilities at this level requires a phased approach.
What modalities does Gemini 4 Argon support?
The model accepts text, image, video and speech input and produces text output, with up to 1 million output tokens enabled by Long Decode Continuation, per Google DeepMind.
Sources
- Google — Announced Gemini 4 Argon on Sept. 30, 2026; Fairwind Program rollout to trusted cyber defenders; DeepSWE v1.1 77.9% new state of the art; AutomationBench 51.3%; introductory pricing of $2/$10 per million tokens with cached input at 95% off; Kavukcuoglu phased-approach quote.
- Google DeepMind — Vals Index 68.9% and LVBench 91.7% in performance table; frontier performance across software engineering, legal, finance and cybersecurity; text/image/video/speech input and text output modalities.
- VentureBeat — Argon posts the highest score, or ties for the highest score, in more categories than GPT-6 Astra or Claude Opus 5.5; Gray Swan IPI prompt-injection attack success rate of 0.7%, lowest among peers; limited release framing.
- Artificial Analysis — Gemini 4 Argon scores 53 on the Artificial Analysis Intelligence Index, tied with GPT-6 Astra, placing Google among the top three labs.
- X (Logan Kilpatrick) — Argon rolling out to cyber defenders starting Sept. 30, 2026 and more widely as soon as possible; priced at $2 in and $10 out during introductory pricing.