Three major model announcements now share a $2 input and $10 output price per million tokens. That is a useful comparison point, not an industry-wide default or a promise of equal bills. Argon’s rate is introductory and access is restricted; task length, caching and retries can outweigh a matching list price.
Editorial note: Prepared October 1 as a September 30, 2026 dispatch. Launch announcements are dated through September 30; GPT-6.1 Sol and Gemini 3.1 pricing documentation was checked October 1. Later product developments are outside this article’s scope.
- 01Compare the same billing modeStandard short-prompt rates are not long-context or accelerated rates.
- 02Price the accepted resultRetries, tools and review belong in the workload cost.
- 03Budget beyond a promotionArgon’s announced later rate is $4 input and $20 output.
01 — The evidenceThe launch-rate convergence, with two counterexamples
Anthropic’s Sonnet 5.5 announcement lists $2/$10 and $0.20 cache reads. OpenAI’s GPT-6.1 Sol documentation lists $2/$10 and $0.10 cached input. Google’s Argon announcement specifies introductory $2/$10 and a 95% input-cache discount, with $4/$20 after the introductory period. Argon begins with vetted access; an announced price is not general API availability.
| Model | Input / output | Cache read | Scope |
|---|---|---|---|
| Sonnet 5.5 | $2 / $10 | $0.20 | Standard launch rate. |
| GPT-6.1 Sol | $2 / $10 | $0.10 | Standard; long prompts cost more. |
| Gemini 4 Argon | $2 / $10 | $0.10 derived | Introductory; later $4 / $20. |
| Gemini 3.1 Pro Preview | $2 / $12 | $0.20 | Prompts up to 200K. |
| Grok 4.7 | $2 / $6 | $0.50 | Prompts below 200K. |
The last two rows prevent an overbroad conclusion. Google’s pricing page and xAI’s September 21 release notes show different output rates at the same input price. Our maintained price index carries the wider model catalogue; this article addresses the purchasing implication of the launch cluster.
02 — Practical implicationsA small workload shows why the token mix matters
Consider an illustrative task with 10,000 uncached input tokens, 100,000 cache-read tokens and 5,000 output tokens. At $2/$10 with $0.10 cache reads, its token cost is $0.02 + $0.01 + $0.05 = $0.08. At the same input/output rates with $0.20 cache reads, it is $0.09. This is arithmetic on an assumed workload, not a measured benchmark.
The example excludes cache creation, storage, tools, regional premiums and accelerated service. It also assumes the cached input qualifies for the vendor’s discount. A prompt that looks similar to a previous prompt does not establish a cache hit. Use actual usage records to distinguish cached input from input billed at the ordinary rate.
Divide the complete workload bill by the number of outputs that passed the same acceptance check. A low per-attempt cost can lose its advantage if the workflow needs more retries or more human correction.
03 — Practical implicationsHold the acceptance standard constant
How much work does the model do?
Record input, cached input and output separately for every attempt.
How often is the result accepted?
Count failed attempts and fallback calls in the same workload.
What happens around inference?
Add tools, review time and the cost of maintaining the integration.
Use the same task set and output contract when comparing candidates. A concise answer that omits a required check should not win because it costs less. Conversely, a longer response is not more valuable simply because it consumes more output tokens. Judge whether the work meets the requirement before interpreting the bill.
Keep reasoning settings explicit. A model may offer several effort levels, and a default setting can change the amount of computation used. Our routing guide describes a task-based comparison that can include a cheaper first attempt and a stronger fallback without hiding the fallback cost.
04 — Practical implicationsKeep access and later pricing in the forecast
For Argon, calculate a later-price scenario before designing a workflow around the introductory offer. On uncached input and output, the announced $4/$20 is twice $2/$10. Do not infer an introductory end date or additional pricing tiers that Google has not published. The Argon launch guide preserves those availability and pricing limits.
A budget should also identify the operational condition that changes the rate: long prompts, faster service, regional processing or a different serving provider. Write that condition alongside the estimate. Otherwise, the number can remain in a planning spreadsheet after the workload no longer qualifies for it.
05 — Practical implicationsUse price convergence to widen the trial
Matching headline rates make it easier to justify testing alternatives; they do not justify switching without evidence. Choose one representative workload, retain the same permission boundary and compare accepted-result cost. Our AI transformation service helps teams connect that evaluation to a controlled routing decision.
Choose from accepted work and the complete bill
Treat $2/$10 as a useful launch-price cluster. Keep each model’s access, cache treatment and billing conditions attached, then choose the configuration that finishes your work reliably at an acceptable total cost.