AI DevelopmentNew Release6 min readPublished September 30, 2026

Google’s new frontier model, priced before it is public

Gemini 4 Argon: Price, Benchmarks and When You Can Use It

Gemini 4 Argon launches at $2/$10 per million tokens, rising to $4/$20 later. Who can use it now, what Google's benchmarks show and how to prepare.

DA
Digital Applied Team
Research and practical guidance
AnnouncedSeptember 30, 2026
Sources checkedOctober 1, 2026

Google announced Gemini 4 Argon, its new frontier model, on September 30, 2026. Most teams cannot use it yet. Argon is going first to a vetted group of cyber defenders, and Google says paid API customers and Google AI Ultra subscribers will follow, without giving a date. What is public already is the price, the output limit and a long list of benchmark claims.

Our view: treat the announcement as a planning signal, not a switching signal. Build the evaluation now, budget at the price Argon will settle at rather than the launch offer, and decide how much output a single task is allowed to buy.

Key takeaways
  1. 01
    Budget at $4/$20, not $2/$10The launch price is introductory. Google has published the later rate but not when the introductory period ends.
  2. 02
    A 1M output limit is also a spending limitOne maximum-length response would cost $10 in output at the launch rate and $20 afterwards. Set caps per task.
  3. 03
    The benchmarks are Google’sEvery score below comes from Google or its partners. Your own tasks decide whether Argon earns a place.

01 — The releaseAnnounced for everyone, available to a vetted few

According to Google’s announcement, Argon is rolling out to trusted cyber defenders through the Fairwind Program, while Google gathers feedback and refines guardrails. Google also says it is taking part in the U.S. government’s voluntary process for pre-release model access. The wider release is described as coming “as soon as possible”, led by paid API customers and Google AI Ultra subscribers.

The Fairwind Program page says the programme works with more than 650 partners and that a subset of them get exclusive access to Argon. Priority goes to governments, critical infrastructure operators such as healthcare, energy and telecommunications, and core technology platforms. Applicants are vetted. Approved organisations must use phishing- resistant multi-factor authentication, restrict Argon to internal security, incident-response or penetration-testing teams, and track each employee’s use.

For these defenders, and for Google’s own teams, Argon ships without its cyber guardrails. That is the same gated pattern we described when Gemini 3.5 Flash Cyber launched: the most capable security behaviour goes to vetted users first, and the general release comes with guardrails in place.

Not yet published

As of October 1, 2026, Google has not published a public release date, an API model ID, the length of the introductory price period, Argon’s context window, any surcharge for long prompts, cache storage or batch rates, or whether Argon replaces an existing Gemini model. Plans that depend on any of these should wait for the model documentation.

02 — The economicsA launch price that matches GPT-6.1 Sol, then doubles

Google gives Argon an introductory price of $2 per million input tokens and $10 per million output tokens, with cached input 95% below the input rate. A footnote sets the later price at $4 and $20. A token is a unit of text the model processes: input is what it reads, output is what it writes. The table sets those rates beside Google’s current Pro model on the Gemini API price list and OpenAI’s GPT-6.1 Sol, released the day before.

USD per million tokens, standard rates, checked October 1, 2026. Argon’s cached rates are derived from Google’s “95% off”, assumed to continue after the introductory period. Gemini 3.1 Pro Preview rates apply to prompts up to 200,000 tokens; GPT-6.1 Sol rates to inputs up to 272,000 tokens.
ModelInputCached inputOutput
Gemini 4 Argon, introductory$2.00$0.10$10.00
Gemini 4 Argon, after introduction$4.00$0.20$20.00
Gemini 3.1 Pro Preview$2.00$0.20$12.00
GPT-6.1 Sol$2.00$0.10$10.00

At launch, Argon’s three headline rates are identical to GPT-6.1 Sol’s and its output rate undercuts Gemini 3.1 Pro Preview. After the introductory period its input and output rates double, which puts Argon’s output at $20, above both. The comparison is not complete: OpenAI also bills cache writes at $2.50 per million, and Google charges Gemini 3.1 Pro users $4.50 per million tokens per hour to store a cache. Argon’s equivalents are not yet published.

Google has used this structure before. Gemini 3.8 Flash lists $0.75 and $3.75 through December 31, 2026, then $1.50 and $7.50 from January 1, 2027. The difference with Argon is that no end date has been given, so a cost model built on the launch rate has no known expiry.

Worked example — illustrative, not a measured run

Take 100,000 cached tokens, 10,000 new input tokens and 5,000 output tokens. Argon costs $0.08 at the introductory rate and $0.16 afterwards. Gemini 3.1 Pro Preview costs $0.10 before cache storage; GPT-6.1 Sol costs $0.08 before cache writes. The example excludes retries, tools and review time.

Our frontier-model API price index puts these rates in the wider market. For Argon, the sensible planning figure is the $4 and $20 rate: if the business case only works at the launch price, it is a case with an unknown end date.

03 — The capacityA million output tokens per response

Google is raising Argon’s output limit to 1M tokens, up from the 64K of its previous models, and calls the figure industry-leading. For comparison, OpenAI’s model page lists 128,000 maximum output tokens for GPT-6.1 Sol. Google’s argument is that a model given room to reason across hundreds of thousands of tokens can solve hard problems in a single pass.

When the model has the headroom to think deeply and generate hundreds of thousands of tokens in a single trajectory, it adds a new level of depth in reasoning to solve tough problems in one go.Koray Kavukcuoglu, Google DeepMind, announcing Gemini 4 Argon, September 30, 2026

The output limit is not the context window, which is how much the model can read in one request. Google has not published Argon’s context window, and the two numbers should not be compared or added together.

The output limit is also a cost ceiling. At $10 per million output tokens, a single response that used the full allowance would cost $10 in output alone; at $20, it would cost $20. Google has not yet published Argon’s billing rules, but its current Gemini API price list counts thinking tokens as output. An agent that reasons at length on every step can spend that much over a run without producing more useful work. Set a maximum output per task, a budget per run and a stopping rule before giving an agent that much room.

04 — The evidenceStrong reported scores, mostly without rivals beside them

Google’s announcement and the Fairwind page report these results. Where a comparison model appears, it is the one Google chose to show.

Reported by Google on September 30, 2026, in its announcement and on the Fairwind Program page. Not independently reproduced. Internal benchmarks are not public.
BenchmarkWhat it measuresArgonContext Google gives
DeepSWE v1.1Long-horizon software engineering77.9%Described as a new state of the art
AutomationBench (Zapier)End-to-end business tasks51.3%Ranked first
LVBenchUnderstanding long videos91.7%Described as state of the art
CWE-bench v1Fixing security vulnerabilities68%Tied with Grok 4.7 and GPT-6 Astra; Claude Opus 5.5 at 67%
Google internal vulnerability benchmarkFinding real-world vulnerabilities85.8%Gemini 3.8 Flash Cyber: 71.0%
Wiz penetration-test benchmarkProbing live web systems without code70.9%Gemini 3.8 Flash Cyber: 58.2%
Gray Swan indirect prompt injectionAttack success over 15 attempts; lower is better0.7%Lowest of the models Google charted

Google also calls Argon the leading model on the Vals Index, which weights finance, coding, legal and tax work by each sector’s share of U.S. GDP, and reports leading results on Vals Finance Agent v2 and Harvey’s Legal Agent Benchmark. The announcement gives no scores for those three.

Two cautions apply. First, these are vendor-reported results, not Digital Applied tests, and most rows show Argon alone or beside an older Google model. Second, the one row with named rivals is informative: on CWE-bench v1, Argon ties two competitors and leads Claude Opus 5.5 by a single point. Expect a capable model in a close field, not a model in a class of its own.

Google’s internal examples point the same way. It says Argon agents freed more than 300 TiB of memory across its data centres and produced a memory-safe Rust video decoder that runs 2.7 times faster than the earlier Rust port. Those are useful indications of long-running engineering work, reported by the company that built the model.

05 — The controlsBetter prompt-injection resistance still needs permissions

Before the wider release, Google lists four areas of safeguards: refusing cyber and chemical, biological, radiological and nuclear misuse under its Frontier Safety Framework, with monitoring of the model’s internal activations; resistance to indirect prompt injection; monitors that watch Argon’s reasoning and actions and stop execution when it goes beyond the user’s intent; and more tightly sealed testing environments.

Indirect prompt injection is the risk most relevant to business agents. It means instructions hidden in a web page, email or document the agent reads, which then try to redirect what the agent does. A 0.7% success rate after 15 attempts on the benchmark Google reported is encouraging. It does not make an agent safe to hold permissions it should not have.

Keep authority in the application. An agent that drafts a refund should not be able to issue it without approval; an agent that reads customer email should not also be able to send external messages unchecked. Google’s own monitors stop execution inside its systems; your tools need equivalent limits inside yours.

06 — The decisionPrepare the evaluation before access opens

Our recommendation is to do the preparation that does not depend on access. Choose a recurring task with a clear acceptance check, such as a code change that must pass review or a research answer that must cite the right document. Record what your current model achieves, what it costs and how long review takes. When Argon reaches the API, the comparison takes days rather than weeks.

You run Gemini 3.1 Pro or 3.8 Flash in production
Build the task sample now. Price Argon at $4/$20 and compare against the model you already pay for.
Prepare the test
Your agents run long coding or migration jobs
Argon is a strong candidate on Google’s evidence. Test it with output caps and a per-run budget in place.
Shortlist it
You defend government or critical infrastructure systems
Read the Fairwind terms and apply. Access requires vetting, strong authentication and per-employee tracking.
Apply now
Your work is hard to check
Wait. A more capable model does not help if you cannot tell whether its answer is right.
Build the check first
Next step

Write the test while you wait

Gemini 4 Argon arrives with credible capability claims, a launch price level with GPT-6.1 Sol and an output allowance large enough to need a budget. None of it can be tested on your work until access opens. Use the wait to fix the task sample, the acceptance check and the full-rate cost model, so the first week of access produces a decision rather than an impression.

Agentic AI implementation

Choose a model around the work it must finish

Digital Applied helps teams evaluate and build AI workflows with clear acceptance checks, cost visibility and controlled tool access.

Task evaluationsCost visibilityControlled access
Before access opens

What to prepare

  • →A recurring task with a pass/fail check
  • →Current accepted-result cost and time
  • →A budget at the full $4/$20 rate
  • →Output caps and stopping rules
Questions and answers

The questions we get about Gemini 4 Argon

Only through the Fairwind Program, which is limited to vetted cyber defenders such as governments, critical infrastructure operators and core technology platforms. Google says paid API customers and Google AI Ultra subscribers will be next, without a date.
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.

Related dispatches

Continue reading

AI Development

GPT-6.1 Sol: Pricing, Benchmarks and the Upgrade Case

GPT-6.1 Sol approaches Astra on key agent tasks at $2/$10 per million tokens. Compare pricing, cache costs, benchmark limits and the case for switching.

September 29, 2026 · 5 minRead
AI Development

Decision Models: AI That Returns Probabilities, Not Text

Four publishers now list decision models on OpenRouter that return probabilities for typed questions, not prose, at free to $0.05 per million input tokens.

September 28, 2026 · 6 minRead
AI Development

Fireworks Ember-1: Kimi K3 Quality With Fewer Tokens?

Fireworks tuned Kimi K3 into Ember-1 and says it matches K3 with 35 to 50% shorter reasoning at the same price. The rows it loses, and the preview caveat.

September 23, 2026 · 5 minRead
AI Development

Claude Sonnet 5.5: Pricing, Benchmarks and Safeguards

Claude Sonnet 5.5 keeps Sonnet 5's $2/$10 price, scores 70.6% on Terminal-Bench 4.0 and adds cyber fallbacks. What the 30% saving is measured on.

September 28, 2026 · 9 minRead
AI Development

What People Let AI Agents Access: 2026 Survey Numbers

A July 2026 survey of 5,067 US adults: 41% of AI users have tried an agent and 32% have let AI act without a final sign-off. Every figure with its base.

September 16, 2026 · 6 minRead
AI Development

After AI Context Compaction, Which Instructions Survive?

Check whether an AI agent follows the right instructions after context compaction. Use behavioral probes for task scope, permissions, evidence and progress.

September 12, 2026 · 6 minRead
Google Search

See more Digital Applied analysis in your Google results by adding us as a preferred source.

Add as a preferred source