Google had gone more than seven months without a model above its cheaper Flash line.
A frontier system called Gemini 3.5 Pro was supposed to arrive in June, then slipped inside the company and never shipped. The last non-Flash release, Gemini 3.1 Pro Preview, dated to February 19.
On September 30 the wait ended, at least on paper.
Google introduced 'Gemini 4 Argon,' the first model in a new generation, and did not offer it to the public.
It is going first to vetted cyber defenders through a program called Fairwind, with a voluntary pre-release review by the U.S. government, and only later to paid API customers and Google AI Ultra subscribers.
The design bet is not a new kind of answer. It is a longer one.
Earlier Gemini models stopped at 64,000 output tokens.
Argon can keep going to 1 million in a single trajectory, counting the reasoning it does before it writes anything a user sees. Google’s claim is that the extra room lets the model stay inside one problem instead of being cut off and restarted.
A token is roughly three-quarters of an English word, so the ceiling is on the order of 750,000 words of thinking and text.
The API has a feature called Long Decode Continuation that pauses a run and resumes it on a later call, so the length does not die on a request timeout.
Artificial Analysis, scoring the model outside Google, found that Argon uses that room. On its tasks the model averaged about 62,000 output tokens, more than twice what GPT-6 Astra spent.
The gain, where there is one, comes from staying on the job, not from a shorter path to the same place.
Inside Google the work looks like a loop, not a single completion.
On libgav1, the company’s open-source video decoder, Argon agents took an existing Rust port and replaced 32,000 lines of hand-written SIMD code.
They did it by running repeated profile-guided experiments, reading what the compiler actually emitted, and writing ordinary safe Rust that the compiler could vectorize on its own.
The result Google reports is a memory-safe decoder 2.7 times faster than the prior Rust port, with identical video output, and closer to the optimized C++.
The same pattern is being applied to larger migrations, from libraries such as re2 up to more than 800,000 lines in the Fuchsia operating system’s Zircon kernel.
Those rewrites do not land on their own.
Google says they go through automated and manual audit, emulation, and review before production. A separate team pointed agents at fleet-wide profiling data and had them find and apply memory changes.
Once rolled out, the changes freed more than 300 tebibytes, with an internal estimate of 500 tebibytes to 1 pebibyte and no new hardware.
In a quantum computing demo, the model cut the spacetime cost of a bottleneck subroutine, measured in qubits times gates, by 40% against a published baseline, in minutes. All of this is Google reporting on Google’s own code.
The cyber path is the same idea pointed at flaws.
Google says Argon was trained to find a vulnerability, check that it is real, and propose a patch, including on live web systems where it does not get the source.
On a black-box test run with the security firm Wiz, the company says the model maps the attack surface, names the weakness, and produces proof-of-concept evidence rather than a description. On an internal set it worked across codebases in 20 languages.
In an early demonstration, Google says it found a critical exposure of personal data in healthcare software used by hospitals, a flaw earlier frontier models had missed. That claim has not been independently described.
For the first outside users, the mechanism includes a deliberate absence.
Fairwind defenders and internal teams get the model without the cyber guardrails that would otherwise refuse offensive-looking requests, on the view that finding a flaw and writing an exploit-shaped proof are the same motion.
Everyone else is supposed to hit a refusal on cyberattacks and on chemical, biological, radiological, and nuclear harm, while dual-use research stays open under Google’s Frontier Safety Framework.
Knowledge work is framed the same way: a long trace over a pile of material, not a short summary.
Google points to chart reading, detail pulled from long video, and action across a series of documents, the sort of input a finance or legal workflow actually arrives as. The public tests in that category, including the Vals Index and a legal-agent benchmark from Harvey, are where Argon’s margins over rival frontier models are widest.
The absolute bar is still low.
On the legal test, fewer than one in five tasks are finished. A model that can stay with a file for hundreds of thousands of tokens is not the same thing as a model that is right at the end of it.
The safety apparatus is part of how a run is allowed to continue.
Google says it monitors internal activations for signs of misuse, and that Argon was adversarially trained against indirect prompt injection, the attack where instructions hide in a document or a web page the model was asked to read.
A separate system reads the chain of thought and the actions and can stop the run if they leave what the user asked for.
A similar monitor watched training, with alerts to an incident team and, Google says, no feedback into the training loop, so the model would not learn to hide the signal. High-risk evaluations run in sandboxes that are isolated and sealed first.
On some external computer-use tests, safety filters turn a flagged response into an empty string and let the episode continue, which is a different choice from killing the task.
Price follows the length.
Introductory rates are $2 per million input tokens and $10 per million output tokens, with cached input at 95% off.
Logan Kilpatrick, who works on the Gemini API, confirmed those figures the morning of the announcement. After the introductory period the stated rates are $4 and $20. At the discount, Artificial Analysis puts the cost of one of its index tasks at about $1.99, under GPT-6 Astra.
At list price it rises above Astra, because Argon writes so much more. The discount is doing the work that token thrift is not.
What is still missing is a run that is not Google’s.
The long trace, the experiment loop, the proof-of-concept cyber step, and the monitor that can halt a trajectory are the account the company is giving of how Argon gets its results. Broader access waits on feedback from the first cohort and on guardrails Google says are not finished.
The date it offers is as soon as possible, not a day.























































































































































































































































































































































































