AI DevelopmentAnalysis3 min readPublished September 30, 2026

Equal token rates can still produce different bills

Three AI Models at $2/$10: Compare the Full Task Cost

Sonnet 5.5, GPT-6.1 Sol and Argon’s introductory rate share $2/$10 token pricing. Compare cache treatment, access limits and the full cost of accepted work.

DA
Digital Applied Team
Research and practical guidance
CoverageSeptember 30, 2026

Three major model announcements now share a $2 input and $10 output price per million tokens. That is a useful comparison point, not an industry-wide default or a promise of equal bills. Argon’s rate is introductory and access is restricted; task length, caching and retries can outweigh a matching list price.

Editorial note: Prepared October 1 as a September 30, 2026 dispatch. Launch announcements are dated through September 30; GPT-6.1 Sol and Gemini 3.1 pricing documentation was checked October 1. Later product developments are outside this article’s scope.

Key takeaways
  1. 01
    Compare the same billing modeStandard short-prompt rates are not long-context or accelerated rates.
  2. 02
    Price the accepted resultRetries, tools and review belong in the workload cost.
  3. 03
    Budget beyond a promotionArgon’s announced later rate is $4 input and $20 output.

01 — The evidenceThe launch-rate convergence, with two counterexamples

Anthropic’s Sonnet 5.5 announcement lists $2/$10 and $0.20 cache reads. OpenAI’s GPT-6.1 Sol documentation lists $2/$10 and $0.10 cached input. Google’s Argon announcement specifies introductory $2/$10 and a 95% input-cache discount, with $4/$20 after the introductory period. Argon begins with vetted access; an announced price is not general API availability.

USD per million tokens, September 30 launch comparison. Vendor sources linked in this section; Argon cache read is 5% of $2. Tool and other charges excluded.
ModelInput / outputCache readScope
Sonnet 5.5$2 / $10$0.20Standard launch rate.
GPT-6.1 Sol$2 / $10$0.10Standard; long prompts cost more.
Gemini 4 Argon$2 / $10$0.10 derivedIntroductory; later $4 / $20.
Gemini 3.1 Pro Preview$2 / $12$0.20Prompts up to 200K.
Grok 4.7$2 / $6$0.50Prompts below 200K.

The last two rows prevent an overbroad conclusion. Google’s pricing page and xAI’s September 21 release notes show different output rates at the same input price. Our maintained price index carries the wider model catalogue; this article addresses the purchasing implication of the launch cluster.

02 — Practical implicationsA small workload shows why the token mix matters

Consider an illustrative task with 10,000 uncached input tokens, 100,000 cache-read tokens and 5,000 output tokens. At $2/$10 with $0.10 cache reads, its token cost is $0.02 + $0.01 + $0.05 = $0.08. At the same input/output rates with $0.20 cache reads, it is $0.09. This is arithmetic on an assumed workload, not a measured benchmark.

The example excludes cache creation, storage, tools, regional premiums and accelerated service. It also assumes the cached input qualifies for the vendor’s discount. A prompt that looks similar to a previous prompt does not establish a cache hit. Use actual usage records to distinguish cached input from input billed at the ordinary rate.

The useful comparison unit

Divide the complete workload bill by the number of outputs that passed the same acceptance check. A low per-attempt cost can lose its advantage if the workflow needs more retries or more human correction.

03 — Practical implicationsHold the acceptance standard constant

Token volume
How much work does the model do?
Measure usage

Record input, cached input and output separately for every attempt.

Billing evidence
Success
How often is the result accepted?
Use fixed checks

Count failed attempts and fallback calls in the same workload.

Quality evidence
Operations
What happens around inference?
Include the workflow

Add tools, review time and the cost of maintaining the integration.

Total cost

Use the same task set and output contract when comparing candidates. A concise answer that omits a required check should not win because it costs less. Conversely, a longer response is not more valuable simply because it consumes more output tokens. Judge whether the work meets the requirement before interpreting the bill.

Keep reasoning settings explicit. A model may offer several effort levels, and a default setting can change the amount of computation used. Our routing guide describes a task-based comparison that can include a cheaper first attempt and a stronger fallback without hiding the fallback cost.

04 — Practical implicationsKeep access and later pricing in the forecast

For Argon, calculate a later-price scenario before designing a workflow around the introductory offer. On uncached input and output, the announced $4/$20 is twice $2/$10. Do not infer an introductory end date or additional pricing tiers that Google has not published. The Argon launch guide preserves those availability and pricing limits.

A budget should also identify the operational condition that changes the rate: long prompts, faster service, regional processing or a different serving provider. Write that condition alongside the estimate. Otherwise, the number can remain in a planning spreadsheet after the workload no longer qualifies for it.

05 — Practical implicationsUse price convergence to widen the trial

Matching headline rates make it easier to justify testing alternatives; they do not justify switching without evidence. Choose one representative workload, retain the same permission boundary and compare accepted-result cost. Our AI transformation service helps teams connect that evaluation to a controlled routing decision.

Next step

Choose from accepted work and the complete bill

Treat $2/$10 as a useful launch-price cluster. Keep each model’s access, cache treatment and billing conditions attached, then choose the configuration that finishes your work reliably at an acceptable total cost.

Agentic AI implementation

Build a workflow you can evaluate and control

Digital Applied helps teams connect AI capabilities to useful work, clear acceptance checks and responsible operating limits.

Task evaluationsCost visibilityControlled access
Start with one task

Define the pilot

  • →Approved source material
  • →A named reviewer
  • →A clear acceptance check
  • →Spending and permission limits
Questions and answers

Practical questions

No. The table includes nearby counterexamples, and providers retain different model and service tiers.
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.

Related dispatches

Continue reading

AI Development

Gemini 4 Argon: Price, Benchmarks and When You Can Use It

Gemini 4 Argon launches at $2/$10 per million tokens, rising to $4/$20 later. Who can use it now, what Google's benchmarks show and how to prepare.

September 30, 2026 · 6 minRead
AI Development

Gemini 2.5 Flash Image Retires October 2 on the API: Act Now

The Gemini API shuts down gemini-2.5-flash-image on October 2, 2026; Vertex says March 15, 2027. Which date applies, three replacements priced, and a test plan.

September 25, 2026 · 4 minRead
AI Development

What an Hour of Audio or Video Costs to Process With AI

List prices for one hour of audio or video across nine providers, from OpenAI and Google to Alibaba and the speech vendors, with every conversion shown.

September 18, 2026 · 8 minRead
AI Development

Gemini 3.8 Live: Should a Voice Agent Think While Talking?

Google split its live voice model in two on September 15: one answers at once, one reasons while it speaks. Which to pick, and where each is available.

September 15, 2026 · 7 minRead
AI Development

AI Agent Governance: Policy and Compliance 2026 Guide

AI agent governance framework for enterprises — access control, audit trails, data residency, and compliance with EU AI Act and SOC 2 requirements.

May 23, 2026 · 20 minRead
AI Development

Google AI Plans: Free vs Plus vs Pro vs Ultra 2026

Google's AI subscription tiers after I/O 2026 — AI Plus $7.99, AI Pro $19.99, AI Ultra $100 (new), AI Ultra $200 (was $250). Feature matrix and decision tree.

May 23, 2026 · 14 minRead
Google Search

See more Digital Applied analysis in your Google results by adding us as a preferred source.

Add as a preferred source