Gemini 4 Argon is here. Almost nobody can use it yet.
Google's most powerful model ties GPT-6 Astra on independent testing and posts the lowest hallucination rate among top models, but at launch it is limited to vetted cyber defenders. Here is what is known, what is first-party, and what it costs once access opens.
October 2, 2026

Google announced Gemini 4 Argon on September 30, calling it its most powerful model yet and aiming it at long-horizon software engineering, legal and finance work, and cyber defense. It is Google's first new top-tier model since Gemini 3 Pro in November 2025, a gap in which OpenAI shipped GPT-6 Astra and Anthropic shipped the Fable and Mythos tiers. We have added it to our model database.
The catch is in the rollout. Argon is going first to "trusted cyber defenders" in Google's Fairwind Program. Paid API customers and Google AI Ultra subscribers are next, and Google has not given a date.
"Safely releasing frontier capabilities at this level requires a phased approach." - Koray Kavukcuoglu, SVP, Google DeepMind
Why the launch is gated
According to The Guardian, Google is holding the model back from the public to avoid misuse by attackers, the same playbook Anthropic used when it kept Claude Mythos Preview restricted. Argon refuses help with cyber and CBRN attacks for general users, while Fairwind partners receive it without the cyber guardrails. Google says it monitors the model's chain of thought and actions for misalignment, and that it is taking part in the U.S. government's voluntary pre-release access process. The launch landed one day after a White House voluntary AI accord was signed.
The numbers, split by who measured them
Launch benchmarks are marketing until someone else reproduces them. Here is what Google reported next to what Artificial Analysis found in independent testing.
| Benchmark | Gemini 4 Argon | Source |
|---|---|---|
| DeepSWE v1.1 | 77.9% | |
| LVBench (long video) | 91.7% | |
| AutomationBench (Zapier) | 51.3%, #1 | |
| Intelligence Index | 53 (GPT-6 Astra 53, Claude Opus 5.5 58) | Artificial Analysis |
| Terminal-Bench 4 | 57% (Sonnet 5.5 64%, Opus 5.5 60%, Astra 59%) | Artificial Analysis |
| Hallucination rate (AA-Omniscience) | 15% (Astra 51%) | Artificial Analysis |
The hallucination figure is the eye-catching one: 15% is the lowest Artificial Analysis has recorded among models at this capability level. Read it with its partner number, though. Argon answered 50% of the knowledge questions correctly against Astra's 63%. It hallucinates less partly because it declines more often, which is a good trade for legal and finance work and a worse one for open-ended research.
Bloomberg, as reported by The Next Web, says some inside Google doubt the model: uneven coding, weak front-end design, and concerns it was tuned for benchmarks. Google disputes the coding criticism.
What it will cost
Google published pricing even though almost no one can buy it yet: $2 per million input tokens and $10 per million output tokens as an introductory rate, rising to $4 and $20 afterwards, with cached input 95% off. Output can now run to 1 million tokens, up from 64K on earlier Gemini models.
The launch price makes Argon cheaper per token than Astra, but it is a verbose model. Artificial Analysis measured about 62K output tokens per task against Astra's 27K. Even so, its cost per index task came in at $1.99 versus $3.26 for Astra at launch pricing. At the standard $4/$20 rate that rises to an estimated $3.98, roughly Astra's level. If you are budgeting, plan on the second number. Our AI Pricing Index will record the change when the introductory period ends.
Should you wait for it?
For most buying decisions this month, nothing changes. If you use the Gemini app today, Argon is not in it. If you are choosing between assistants, our ChatGPT vs Gemini and Claude vs Gemini comparisons cover the products you can actually subscribe to.
The teams that should watch the rollout closely are those running long, document-heavy workflows where a confident wrong answer is expensive. That is where a low hallucination rate and a 1 million token output limit matter more than a few points on a coding benchmark. Everyone else can let the Fairwind partners find the rough edges first.
Some links in this article are affiliate links. Learn more.