Google announced Gemini 4 Argon on September 30, calling it a frontier model engineered for long, complex workflows – real-world software engineering, enterprise knowledge work across legal and finance, and cybersecurity defense.
It already runs inside Google, where thousands of Googlers use it for specialized coding, deeper research, and writing. Argon is not generally available to developers or consumers yet. Access opens first to a set of trusted cyber defenders through the Fairwind Program, which already counts more than 650 partners, and Google is participating in the U.S. government’s voluntary pre-release process.
Broader availability for paid API customers and Google AI Ultra subscribers is promised: “as soon as possible.” No public release date has been announced. Gemini 4 Argon is not yet generally available to developers or consumers.
What Google Says Gemini 4 Argon Can Actually Do
Google has lifted Argon’s output ceiling to one million tokens, a considerable jump from the 64,000 permitted by earlier Gemini models. Output capacity, it should be noted, differs from the full context window, and that added room, Google says, lets the model carry deep reasoning through multi-step tasks in a single trajectory.
Argon agents recovered more than 300 TiB of memory across company data centers, with projected savings somewhere between 500 TiB and 1 PiB. Quantum researchers leaned on it to beat a published algorithmic baseline by 40 percent within minutes. Argon agents are also working on C/C++-to-Rust migrations ranging from core libraries such as re2 and libgav1 to more than 800,000 lines associated with Fuchsia’s Zircon kernel; Google says the larger rewrites are still undergoing extensive review before production deployment.
For libgav1, Google says Argon replaced 32,000 lines of SIMD code with memory-safe Rust that runs 2.7× faster than the previous Rust port while producing identical video output, bringing performance closer to the optimized C++ implementation.
These figures come from Google itself. Independent confirmation of internal productivity claims will take time.
Gemini 4 Benchmarks Google Published – and Where Caution Is Needed
Google’s published evaluations put Argon at 77.9% on DeepSWE v1.1 for long-horizon software engineering, 68.9% on the Vals Index weighting finance, coding, legal, and tax work, 51.3% on AutomationBench, 91.7% on LVBench for long-video understanding, and 68% on CWE-bench v1 for vulnerability remediation, where it ties with GPT-6 Astra.
Losses are equally documented. On FrontierSWE v2, Argon scores 55.0% against Astra’s 65.5% and Opus 5.5’s 62.3%. Terminal-bench 4.0 places it at 57.4%, behind Astra at 58.2% and Opus 5.5 at 66.4%. Terminal-Bench Science 0.1 shows 57.6% versus Astra’s 68.1%, and OSWorld-2.0 puts Argon at 69.2% against Astra’s 72.6%. Whether the overall picture holds once third parties run the same suites remains an open question.
Argon Initially Matches GPT-6.1 Sol’s Base API Price
Introductory API pricing is set at $2 per million input tokens and $10 per million output tokens. Cached input tokens receive a 95 percent discount. After the introductory period ends, the rates become $4 and $20 per million tokens, and that introductory level aligns with the list price of OpenAI’s GPT-6.1 Sol. Google has not said how long the lower rate will last.
| Model | Input / 1M | Output / 1M | Cached Input | Availability Status |
| Gemini 4 Argon (intro) | $2 | $10 | $0.10 | API rollout upcoming; trusted defenders first |
| Gemini 4 Argon (std) | $4 | $20 | Not separately stated | After the introductory period; timing unspecified |
| GPT-6.1 Sol | $2 | $10 | $0.10 | Available via API |
Note:
$2/$10 reflects GPT-6.1 Sol’s base Standard short-context API rate. OpenAI charges more once requests exceed 272K input tokens, so the two models are not identically priced for very-long-context work.
Why the Staged Release Matters More Than Another Leaderboard
Google trained Argon specifically for defensive cybersecurity. The model represents a marked improvement over 3.8 Flash Cyber, the earlier cybersecurity-focused model Google released this month, with stronger vulnerability detection performance.

Google says Argon uncovered a critical issue that exposes sensitive personal information in healthcare software used by hospitals worldwide and that previous frontier models missed it. That is a Google launch claim, not yet a fully disclosed independent vulnerability report.
Separately, Wiz’s Scan for Good program, which now uses Argon alongside Gemini 3.8 Flash Cyber, published findings on September 24 covering numerous vulnerabilities found primarily by earlier Gemini models and Wiz tooling.
Google says it is deploying mitigations that monitor Argon’s chain-of-thought and actions, while hardening isolated sandbox environments for high-risk training and evaluations.
Details on how long this testing phase will last remain unclear.
Some workflows may benefit more from models already available today, while others may find the eventual combination of capability and cost compelling once the gate opens.
Gemini 4 Argon is best understood as a controlled frontier-model rollout, with cybersecurity defenders and Google’s internal teams providing the first real-world testing ground before wider availability.