Google announced Gemini 4 Argon, its new frontier model, on September 30, 2026. Most teams cannot use it yet. Argon is going first to a vetted group of cyber defenders, and Google says paid API customers and Google AI Ultra subscribers will follow, without giving a date. What is public already is the price, the output limit and a long list of benchmark claims.
Our view: treat the announcement as a planning signal, not a switching signal. Build the evaluation now, budget at the price Argon will settle at rather than the launch offer, and decide how much output a single task is allowed to buy.
- 01Budget at $4/$20, not $2/$10The launch price is introductory. Google has published the later rate but not when the introductory period ends.
- 02A 1M output limit is also a spending limitOne maximum-length response would cost $10 in output at the launch rate and $20 afterwards. Set caps per task.
- 03The benchmarks are Google’sEvery score below comes from Google or its partners. Your own tasks decide whether Argon earns a place.
01 — The releaseAnnounced for everyone, available to a vetted few
According to Google’s announcement, Argon is rolling out to trusted cyber defenders through the Fairwind Program, while Google gathers feedback and refines guardrails. Google also says it is taking part in the U.S. government’s voluntary process for pre-release model access. The wider release is described as coming “as soon as possible”, led by paid API customers and Google AI Ultra subscribers.
The Fairwind Program page says the programme works with more than 650 partners and that a subset of them get exclusive access to Argon. Priority goes to governments, critical infrastructure operators such as healthcare, energy and telecommunications, and core technology platforms. Applicants are vetted. Approved organisations must use phishing- resistant multi-factor authentication, restrict Argon to internal security, incident-response or penetration-testing teams, and track each employee’s use.
For these defenders, and for Google’s own teams, Argon ships without its cyber guardrails. That is the same gated pattern we described when Gemini 3.5 Flash Cyber launched: the most capable security behaviour goes to vetted users first, and the general release comes with guardrails in place.
As of October 1, 2026, Google has not published a public release date, an API model ID, the length of the introductory price period, Argon’s context window, any surcharge for long prompts, cache storage or batch rates, or whether Argon replaces an existing Gemini model. Plans that depend on any of these should wait for the model documentation.
02 — The economicsA launch price that matches GPT-6.1 Sol, then doubles
Google gives Argon an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input 95% below the input rate. A footnote sets the later price at $4 and $20. A token is a unit of text the model processes: input is what it reads, output is what it writes. The table sets those rates beside Google’s current Pro model on the Gemini API price list and OpenAI’s GPT-6.1 Sol, released the day before.
| Model | Input | Cached input | Output |
|---|---|---|---|
| Gemini 4 Argon, introductory | $2.00 | $0.10 | $10.00 |
| Gemini 4 Argon, after introduction | $4.00 | $0.20 | $20.00 |
| Gemini 3.1 Pro Preview | $2.00 | $0.20 | $12.00 |
| GPT-6.1 Sol | $2.00 | $0.10 | $10.00 |
At launch, Argon’s three headline rates are identical to GPT-6.1 Sol’s and its output rate undercuts Gemini 3.1 Pro Preview. After the introductory period its input and output rates double, which puts Argon’s output at $20, above both. The comparison is not complete: OpenAI also bills cache writes at $2.50 per million, and Google charges Gemini 3.1 Pro users $4.50 per million tokens per hour to store a cache. Argon’s equivalents are not yet published.
Google has used this structure before. Gemini 3.8 Flash lists $0.75 and $3.75 through December 31, 2026, then $1.50 and $7.50 from January 1, 2027. The difference with Argon is that no end date has been given, so a cost model built on the launch rate has no known expiry.
Take 100,000 cached tokens, 10,000 new input tokens and 5,000 output tokens. Argon costs $0.08 at the introductory rate and $0.16 afterwards. Gemini 3.1 Pro Preview costs $0.10 before cache storage; GPT-6.1 Sol costs $0.08 before cache writes. The example excludes retries, tools and review time.
Our frontier-model API price index puts these rates in the wider market. For Argon, the sensible planning figure is the $4 and $20 rate: if the business case only works at the launch price, it is a case with an unknown end date.
03 — The capacityA million output tokens per response
Google is raising Argon’s output limit to 1M tokens, up from the 64K of its previous models, and calls the figure industry-leading. For comparison, OpenAI’s model page lists 128,000 maximum output tokens for GPT-6.1 Sol. Google’s argument is that a model given room to reason across hundreds of thousands of tokens can solve hard problems in a single pass.
When the model has the headroom to think deeply and generate hundreds of thousands of tokens in a single trajectory, it adds a new level of depth in reasoning to solve tough problems in one go.Koray Kavukcuoglu, Google DeepMind, announcing Gemini 4 Argon, September 30, 2026
The output limit is not the context window, which is how much the model can read in one request. Google has not published Argon’s context window, and the two numbers should not be compared or added together.
The output limit is also a cost ceiling. At $10 per million output tokens, a single response that used the full allowance would cost $10 in output alone; at $20, it would cost $20. Google has not yet published Argon’s billing rules, but its current Gemini API price list counts thinking tokens as output. An agent that reasons at length on every step can spend that much over a run without producing more useful work. Set a maximum output per task, a budget per run and a stopping rule before giving an agent that much room.
04 — The evidenceStrong reported scores, mostly without rivals beside them
Google’s announcement and the Fairwind page report these results. Where a comparison model appears, it is the one Google chose to show.
| Benchmark | What it measures | Argon | Context Google gives |
|---|---|---|---|
| DeepSWE v1.1 | Long-horizon software engineering | 77.9% | Described as a new state of the art |
| AutomationBench (Zapier) | End-to-end business tasks | 51.3% | Ranked first |
| LVBench | Understanding long videos | 91.7% | Described as state of the art |
| CWE-bench v1 | Fixing security vulnerabilities | 68% | Tied with Grok 4.7 and GPT-6 Astra; Claude Opus 5.5 at 67% |
| Google internal vulnerability benchmark | Finding real-world vulnerabilities | 85.8% | Gemini 3.8 Flash Cyber: 71.0% |
| Wiz penetration-test benchmark | Probing live web systems without code | 70.9% | Gemini 3.8 Flash Cyber: 58.2% |
| Gray Swan indirect prompt injection | Attack success over 15 attempts; lower is better | 0.7% | Lowest of the models Google charted |
Google also calls Argon the leading model on the Vals Index, which weights finance, coding, legal and tax work by each sector’s share of U.S. GDP, and reports leading results on Vals Finance Agent v2 and Harvey’s Legal Agent Benchmark. The announcement gives no scores for those three.
Two cautions apply. First, these are vendor-reported results, not Digital Applied tests, and most rows show Argon alone or beside an older Google model. Second, the one row with named rivals is informative: on CWE-bench v1, Argon ties two competitors and leads Claude Opus 5.5 by a single point. Expect a capable model in a close field, not a model in a class of its own.
Google’s internal examples point the same way. It says Argon agents freed more than 300 TiB of memory across its data centres and produced a memory-safe Rust video decoder that runs 2.7 times faster than the earlier Rust port. Those are useful indications of long-running engineering work, reported by the company that built the model.
05 — The controlsBetter prompt-injection resistance still needs permissions
Before the wider release, Google lists four areas of safeguards: refusing cyber and chemical, biological, radiological and nuclear misuse under its Frontier Safety Framework, with monitoring of the model’s internal activations; resistance to indirect prompt injection; monitors that watch Argon’s reasoning and actions and stop execution when it goes beyond the user’s intent; and more tightly sealed testing environments.
Indirect prompt injection is the risk most relevant to business agents. It means instructions hidden in a web page, email or document the agent reads, which then try to redirect what the agent does. A 0.7% success rate after 15 attempts on the benchmark Google reported is encouraging. It does not make an agent safe to hold permissions it should not have.
Keep authority in the application. An agent that drafts a refund should not be able to issue it without approval; an agent that reads customer email should not also be able to send external messages unchecked. Google’s own monitors stop execution inside its systems; your tools need equivalent limits inside yours.
06 — The decisionPrepare the evaluation before access opens
Our recommendation is to do the preparation that does not depend on access. Choose a recurring task with a clear acceptance check, such as a code change that must pass review or a research answer that must cite the right document. Record what your current model achieves, what it costs and how long review takes. When Argon reaches the API, the comparison takes days rather than weeks.
Write the test while you wait
Gemini 4 Argon arrives with credible capability claims, a launch price level with GPT-6.1 Sol and an output allowance large enough to need a budget. None of it can be tested on your work until access opens. Use the wait to fix the task sample, the acceptance check and the full-rate cost model, so the first week of access produces a decision rather than an impression.