When an agency runs AI agents for a client, the model bill is often the smallest number in the retainer. With the prices Anthropic published on September 28, 2026, a task that reads a long brief and writes a draft costs about 17 cents on Claude Sonnet 5.5 and 31 cents on Opus 5.5. Four hundred of those cost $67.60 and $125.60 a month.
That does not make the model cost irrelevant to pricing. It makes the level unimportant and the variance important. This post works one illustrative retainer through the arithmetic, shows the three things that can multiply the bill, and compares four ways to write it into a contract. The task volumes and token counts are invented for the example; the rates are Anthropic’s published prices.
- 01400 agent tasks a month cost $67.60 on Sonnet 5.5 and $125.60 on Opus 5.5.About 1% and 2% of an illustrative $6,000 retainer. Batch processing halves either figure for work that can wait.
- 02Reviewing those tasks costs about 35 times the Sonnet 5.5 bill.At an illustrative six minutes per task and $60 an hour of internal cost, review is $2,400 a month.
- 03Effort level, volume and vendor price changes can each double the bill.Together they can take the Opus 5.5 line past $500 in the same month, without anyone changing the scope.
- 04Price the model cost in bands, and name the model and effort level in the scope.A small line that can quadruple needs a rule for who absorbs the swing, written down before it happens.
01 — The scenarioOne retainer, worked through at published prices
Picture a content and reporting retainer. Each month an agent runs 400 tasks for the client: drafting posts from briefs, summarising campaign data, checking pages against a style guide. Each task reads about 150,000 tokens, most of it a cached brief and brand guide, and writes 8,000 tokens including its thinking. The rates are from Anthropic’s pricing page: Sonnet 5.5 at $2 per million input tokens and $10 per million output, Opus 5.5 at $4 and $20, both with $0.20 cache reads.
| One task | Sonnet 5.5 | Opus 5.5 | Sonnet, batch |
|---|---|---|---|
| Cache reads · 120,000 tokens | $0.024 | $0.024 | $0.012 |
| Uncached input · 20,000 tokens | $0.040 | $0.080 | $0.020 |
| Cache writes, 5 min · 10,000 tokens | $0.025 | $0.050 | $0.013 |
| Output incl. thinking · 8,000 tokens | $0.080 | $0.160 | $0.040 |
| Per task | $0.169 | $0.314 | $0.085 |
| Per month · 400 tasks | $67.60 | $125.60 | $33.80 |
Output is the largest line on both models, even though it is the smallest number of tokens, because it costs five times as much as uncached input. Cache reads are the largest volume and the smallest cost, which is why a stable, cached brief matters more to the bill than any per-token price cut. On a $6,000 monthly retainer, the model line is 1.1% on Sonnet 5.5 and 2.1% on Opus 5.5.
02 — The real costThe line that is bigger than the tokens
Every task still needs a person to read it
At six minutes of review per task, 400 tasks take 40 hours a month. At an illustrative internal cost of $60 an hour, that is $2,400, or 40% of the retainer. The model is not where the margin goes.
This changes which saving is worth chasing. Moving from Opus 5.5 to Sonnet 5.5 saves $58 a month in this scenario. Cutting review from six minutes to five saves 400 minutes, which is $400. If the stronger model produces drafts that need less correction, it can be the cheaper choice overall. Anthropic’s own charts show Opus 5.5 scoring higher than Sonnet 5.5 at their default settings, which our effort-level comparison works through.
It also changes the conversation with clients who ask for a share of the AI saving, the pressure covered in our post on clients wanting a share of AI savings. The model does make drafting cheaper. What the client pays for is the judgment and checking around it, and that has not become cheap.
03 — The riskThree things that multiply the model bill
A 2% line is safe to absorb until it moves. Three things move it, and none of them needs the scope to change.
- Effort level. Higher effort settings make the model think longer, and thinking is billed as output. On Anthropic’s published Terminal-Bench chart, Sonnet 5.5 costs $1.94 per attempt at
highand $12.54 atmax. A team that turns effort up to fix one hard task and leaves it there can multiply the output line. - Volume. A campaign launch or a reporting season can double the task count for a month.
- Vendor price changes. Prices move in both directions. In the week to September 28, the OpenRouter route sold as GPT-5.6 Sol “Pro” (pro is a reasoning mode on GPT-5.6 Sol, not a separate OpenAI model) doubled from $2 / $10 to $4 / $20. Earlier in September, Sonnet 5’s introductory $2 / $10 became its standard price instead of rising to $3 / $15 on September 1. Launch discounts end too: ElevenLabs’ new voice models are 72% off only until October 12.
| Month | Sonnet 5.5 | Opus 5.5 |
|---|---|---|
| Base case, 400 tasks | $67.60 | $125.60 |
| Output and thinking triple | $131.60 | $253.60 |
| Task volume doubles | $135.20 | $251.20 |
| Both at once | $263.20 | $507.20 |
In the worst row the Opus 5.5 line reaches $507, 8.5% of the retainer. That is still less than the review cost, but it is no longer a rounding error, and it arrived without any change the client asked for. The price index in our frontier model price index is the easiest way to see a vendor change before the invoice does.
04 — The contractFour ways to write the model cost into a retainer
Bundled, with a usage band
The model cost sits inside the retainer up to a stated task volume or spend, say three times the expected bill. Above it, the overage is billed. Simple for the client, and the agency keeps the saving when costs fall.
Pass-through at cost
Transparent, and the client carries the swing. It also invites the client to compare token prices line by line, which pulls the discussion toward the smallest cost in the retainer.
Pass-through with a margin
Covers the time spent monitoring usage and handling price changes. The margin needs to be stated, or it reads as a hidden fee.
Price per output
The client buys finished work, and the agency carries both model and review cost. The highest margin when review time falls, and the highest risk when a vendor raises prices.
Our earlier agency pricing guide compares retainer, value-based and outcome-based models in general. The point here is narrower: whichever model you use, the AI line needs its own rule, because it behaves differently from staff time. It is small, it can change by several times in a month, and the vendor sets the price.
05 — ConclusionThe model bill is small; the rule for who pays when it moves is not
Write the model, the effort level and a usage band into every AI retainer, and price the review time as the main cost
At cents per task, arguing about the token price wastes time. The questions that protect the margin are who pays when usage triples, which model and setting the price assumed, and how much human time each task needs. Put those in the scope before the first invoice. If you want help setting up the cost model, our AI transformation team builds them.