Google has announced Gemini 4 Argon, the first announced model in its Gemini 4 line, positioning it for long-running software work, enterprise tasks and cybersecurity defense. The launch is significant less for immediate consumer access than for what Google is withholding: Argon is initially available only to trusted cyber defenders in its Fairwind Program and to internal teams while the company tests safeguards.
In its September 30 announcement, Google said Argon can generate up to one million output tokens in a single response, up from 64,000 in previous Gemini models. It also set introductory API prices of $2 per million input tokens and $10 per million output tokens. Google has not announced a date for broader availability.
A flagship model, but not a public release
The controlled rollout is a notable departure from the familiar pattern of announcing a flagship model alongside a broad web or API launch. Google says it will gather feedback, improve guardrails and participate in the U.S. government’s voluntary pre-release model-access process before expanding availability. Paid API customers and Google AI Ultra subscribers are among the groups Google has identified for early wider access, followed by developers, businesses and consumers.
Independent accounts from Ars Technica and CNBC likewise describe a deliberately limited release, rather than a product that ordinary Gemini users can try now. That leaves the announcement as a statement of technical direction and access policy, not yet a broad market test of how the system performs in everyday hands.
Google’s initial security cohort will have an unusual level of access. The company says selected trusted defenders and internal teams will receive Argon without cyber guardrails. That does not mean unrestricted public access: the distinction is who receives the model and under what controlled program. But it acknowledges the practical tension in cyber AI. Systems capable of helping defenders investigate flaws, analyze code or automate remedial work may also make dangerous technical tasks easier if distributed carelessly.
What one million output tokens changes
An output-token limit is the amount of material a model can generate, rather than a simple measure of how much source material it can read. At one million tokens, Argon could in principle produce lengthy code changes, extensive investigation notes or a large set of structured workflow steps without being forced to stop at the 64,000-token ceiling Google says applied to earlier Gemini models.
The practical benefit depends on whether the model can retain direction, avoid errors and use tools reliably over a prolonged task. A very large generation budget is not, by itself, proof of long-horizon competence. It can also increase review burdens: a team receiving an enormous patch, migration plan or security analysis still needs to validate the output before putting it into production.
Google’s stated API pricing makes the scale more tangible. At the introductory $10 per million output tokens, a maximum-length response would carry a listed output cost of about $10 before input and any other applicable charges. Cached input tokens receive a 95% discount from the stated input rate, according to Google. Reports of higher prices after the introductory period have circulated, but the available Google announcement does not specify those later figures.
Benchmarks show strengths, not a clean sweep
Google says Argon achieved 77.9% on DeepSWE v1.1, a benchmark intended to evaluate long-horizon software-engineering work, and tied for first on CWE-bench v1 at 68% for vulnerability remediation. It also reports a 51.3% result on Zapier’s AutomationBench and 91.7% on LVBench, a long-video-understanding evaluation. These are company-reported results, not independently reproduced measurements of deployment performance.
The company’s published comparison table also makes a more restrained reading necessary. The New Stack’s analysis notes that Argon did not lead every coding evaluation: named competing models posted stronger results on FrontierSWE v2 and Terminal-Bench 4.0. Benchmark suites measure different tasks and can use different tools, harnesses and agent configurations, so their scores are not a universal league table. Still, the result undercuts any claim that Argon has established across-the-board leadership merely by posting a standout DeepSWE figure.
Google also says Argon found a critical vulnerability in healthcare software used by hospitals worldwide that earlier frontier models missed. It did not identify the software in the announcement, and the claim has not been publicly independently verified. The same caution applies to Google’s internal deployment examples: it says Argon agents freed more than 300 TiB of data-center memory after deployment, with estimated eventual savings between 500 TiB and 1 PiB.
Code migration is the clearest internal use case
Among Google’s examples, code modernization is the most concrete. The company says Argon assisted work on a Rust decoder for libgav1, its AV1 video decoder library, producing a memory-safe implementation that ran 2.7 times faster than an earlier Rust port while generating identical video output. Google also cites C and C++ to Rust migrations as an internal application.
Those examples speak to a commercially important target: organizations with mature, sprawling codebases and expensive backlogs of maintenance or security work. AI vendors increasingly sell agents as systems that can carry out multi-step engineering assignments rather than merely suggest a function in a chat window. The claimed performance matters, but the restricted rollout means outside developers cannot yet test whether Argon’s long outputs translate into dependable work across their own repositories and toolchains.
For now, Google is offering a narrow view of its new flagship: a high-capacity model, a defined introductory price and a security-sensitive first deployment. Broader access remains contingent on feedback and further guardrail work, with no public release date announced.
