AI DevelopmentNew Release11 min readPublished October 7, 2026

One release · two price tiers · the 100K-token line every bill now depends on

Claude Haiku 5.5: Price, Long-Context Tier and Who Upgrades

Anthropic released Claude Haiku 5.5 on October 7, 2026. For prompts up to 100,000 tokens it costs $0.10 per million input tokens and $0.50 per million output tokens, a tenth of Haiku 4.5. Above that line every rate is five times higher. The context window grows to 1M tokens and the output limit to 128K. This post prices real requests on both sides of the line and says who should move.

DA
Digital Applied Team
Senior strategists · Published October 7, 2026
PublishedOct 7, 2026
Read time11 min
SourcesAnthropic launch post · pricing docs
Input, prompts ≤100K
$0.10/M
Haiku 4.5 was $1.00 · Anthropic pricing page
Output, prompts ≤100K
$0.50/M
Haiku 4.5 was $5.00
Above 100K tokens
5×
$0.50 in · $2.50 out · $0.05 cache read
Context · max output
1M· 128K
Haiku 4.5: 200K · 64K

Claude Haiku 5.5 shipped on October 7, 2026 with a price list that has two columns, and the column you land in depends on how long your prompt is. Up to 100,000 tokens, input costs $0.10 per million tokens and output $0.50. Over 100,000 tokens, input costs $0.50 and output $2.50. Haiku 4.5, the model it replaces, charged $1 and $5 at every length. The API ID is claude-haiku-5-5, and Anthropic says it is live on the Claude API, Amazon Web Services, Google Cloud and Microsoft Azure from day one.

Two things changed at once, and they pull in different directions. The headline rate fell by 90% for the short prompts that make up most small-model traffic. At the same time the model gained a 1M-token context window, five times the 200K that Haiku 4.5 had, and a 128K output limit, double the old 64K. The long-context tier is the price of that extra room. A team that reads the first number and ignores the second will budget wrong.

This post does four things. It lays out both price tiers next to Haiku 4.5 and Sonnet 5.5. It prices two concrete requests, a 50,000-token prompt and a 200,000-token prompt, on each model. It separates what Anthropic measured from what you can expect, and labels every benchmark as the vendor's own. And it ends with a router for who should switch this week. For where Haiku sits in the whole lineup, our frontier model guide is the maintained reference; for the dated ledger of October releases, see the October 2026 release tracker.

Key takeaways
  1. 01
    Short prompts cost a tenth of Haiku 4.5.For prompts up to 100,000 tokens, Haiku 5.5 is $0.10 in and $0.50 out per million tokens against Haiku 4.5's $1 and $5. Anthropic says prompts that size were around 90% of requests to the previous Haiku.
  2. 02
    Above 100K tokens, every rate is five times higher.Input $0.50, output $2.50, cache read $0.05, cache write $0.625. The prompt length that decides the tier counts all input tokens, including cache reads and writes, and each request is priced on its own.
  3. 03
    The window is 1M tokens and the output limit 128K.Haiku 4.5 stopped at 200K and 64K. Haiku 5.5 matches Sonnet 5.5 and Opus 5.5 on both limits, supports adaptive thinking with an effort setting, and has a reliable knowledge cutoff of June 2026.
  4. 04
    The benchmarks are Anthropic's own.Terminal-Bench 4.0 at 39.2% against Haiku 4.5's 0.0%, OSWorld 2.1 offline at 72.4% against 15.7%: these are figures from the launch post, not independent runs. Anthropic itself says Sonnet 5.5 and Opus 5.5 remain the better choice for complex agentic coding.
  5. 05
    Sonnet 5.5 got cheaper the same day.Its cache-read price halved to $0.10 per million tokens. Anthropic puts the effect at around 20% off most agentic work. Max and Team subscribers also start receiving monthly API credits this week.

01 — List RatesWhat the new prices are, both tiers.

Anthropic publishes two full rows for Haiku 5.5 on its pricing page, one for prompts up to 100,000 tokens and one for prompts over it. No other current Claude model is priced this way. Opus 5.5, Sonnet 5.5 and Fable 5.1 charge the same per-token rate across the whole 1M window. Haiku 5.5 is the exception, and the reason is plain enough: it is the first Haiku with a 1M window, and Anthropic has chosen to charge for using the far end of it.

The table puts the two Haiku 5.5 tiers next to Haiku 4.5. The short-prompt tier is a tenth of the old price on every line. The long-prompt tier is half the old input price and half the old output price, so even a 500,000-token request on Haiku 5.5 costs less per token than any request on Haiku 4.5 did. It just costs five times more than the same tokens would have cost under 100K.

Per million tokensHaiku 5.5 ≤100KHaiku 5.5 >100KHaiku 4.5
Input$0.10$0.50$1.00
Output$0.50$2.50$5.00
Cache read$0.01$0.05$0.10
Cache write, 5 minutes$0.125$0.625$1.25
Cache write, 1 hour$0.20$1.00$2.00
Batch input / output$0.05 / $0.25$0.25 / $1.25$0.50 / $2.50
Anthropic's published Claude API list prices, October 7, 2026. Batch is a flat 50% off; the 1-hour cache write is 2× base input.

One row is easy to misread. The cache-read price of $0.01 is 10% of the $0.10 input price, which is the standard Claude ratio. It is not the 5% ratio that Opus 5.5 and Sonnet 5.5 now carry, and not the 2.5% on Fable 5.1. Haiku 5.5's cache reads are cheap because its input is cheap, not because the multiplier moved. Section 04 works through what that means for a cached agent loop.

02 — The ThresholdHow the 100K line is drawn.

Anthropic's long-context pricing note states three rules, and each one matters for a budget. First, the prompt length that decides the tier is the whole input of the request: every token you send, including the tokens served from cache and the tokens written to it. A 95,000-token cached prefix plus a 6,000-token question is a 101,000-token prompt, and it pays the higher rates even though most of it was a cache hit. Second, each request is priced on its own. Third, requests that were billed at the low tier stay there; crossing the line later does not re-price them.

The practical consequence is that the line is not about a conversation's length. It is about the size of the single largest request in it. An agent that accumulates context and re-sends it every turn will cross 100K somewhere in the middle of a long task, and from that turn on its input costs five times more, cache hits included. A pipeline that sends one 40,000-token document per call never gets near it. Anthropic says prompts up to 100,000 tokens were about 90% of requests to the previous Haiku, which is its argument that most users will never see the second column.

One more thing moves the line without anyone noticing. Haiku 5.5 uses the newer tokenizer that arrived with Claude Opus 4.7, and Anthropic's model page says the same text counts as roughly 30% more tokens than it did on Haiku 4.5. A prompt that measured 80,000 tokens on Haiku 4.5 is about 104,000 tokens on Haiku 5.5, and it has crossed the threshold before you changed a word of it.

A failure case worth planning for
A support assistant on Haiku 4.5 keeps a 75,000-token knowledge prefix cached and appends each customer's thread. On the old tokenizer it never exceeded 90,000 tokens. Re-tokenized for Haiku 5.5 the same prefix is close to 98,000, so a thread of a few thousand tokens tips every request over 100K. Input is then $0.50 per million instead of $0.10, and cache reads $0.05 instead of $0.01. The fix is to trim the prefix below the line with headroom, not to switch models.

03 — Worked CostsTwo requests, three models.

Rates per million are hard to feel. A single request is not. The table prices two requests, each with a 2,000-token answer, on Haiku 5.5, Haiku 4.5 and Sonnet 5.5, using the list rates above and no caching or batch discount. The arithmetic is tokens divided by one million, times the rate, summed for input and output. A third row shows the cliff itself: a prompt one token over 100,000.

RequestHaiku 5.5Haiku 4.5Sonnet 5.5
50,000-token prompt, 2,000-token answer$0.006$0.060$0.120
The same request, 1,000 times$6$60$120
200,000-token prompt, 2,000-token answer$0.105Does not fit the 200K window$0.420
Input only: 100,000 tokens vs 100,001 tokens$0.0100 → $0.0500$0.1000 either way$0.2000 either way
Our arithmetic from Anthropic's October 7, 2026 list rates. No caching, no batch discount, no thinking tokens counted. Token counts are held equal across models; the same text is about 30% more tokens on Haiku 5.5 than on Haiku 4.5.

Three readings follow. For the short request Haiku 5.5 is a tenth of Haiku 4.5 and a twentieth of Sonnet 5.5, and at a thousand requests the gap is the difference between six dollars and a hundred and twenty. For the long request Haiku 5.5 is still a quarter of Sonnet 5.5, so the long-context tier is not a penalty relative to the alternatives; it is a penalty relative to Haiku 5.5's own short-prompt rate. And the cliff row shows the shape of the risk: one token changes the input bill by a factor of five, with nothing in the response to tell you it happened.

Thinking tokens are the thing this table leaves out, and they are billed as output. Haiku 5.5 has adaptive thinking on by default at the medium effort level, so a request that looks like 2,000 output tokens may carry more. For a classification or extraction task, set effort low and the number stays close to the table. For a task that needs the model to reason, the output column grows and so does the gap to Haiku 4.5, which had no effort parameter at all.

04 — DiscountsCache and batch rates, and the Sonnet change.

Prompt caching on Haiku 5.5 uses the standard Claude multipliers: a 5-minute cache write at 1.25× input, a 1-hour write at 2× input, and reads at 10% of input. The minimum cacheable prompt is 512 tokens. Batch processing takes a flat 50% off input and output in both tiers. The figures below are the ones to copy into a cost model.

Cache read, prompts ≤100K10% of base input
$0.01 / MHaiku 4.5 $0.10
Cache read, prompts >100Ksame 10% ratio on the higher base
$0.05 / MHaiku 4.5 $0.10
Cache write, 5 minutes1.25× input · ≤100K / >100K
$0.125 / $0.625Haiku 4.5 $1.25
Cache write, 1 hour2× input · ≤100K / >100K
$0.20 / $1.00Haiku 4.5 $2.00
Batch API, input / output50% off · ≤100K tier
$0.05 / $0.25Haiku 4.5 $0.50 / $2.50
Sonnet 5.5 cache read, from October 7halved · now 5% of its $2 input
$0.10 / Mwas $0.20

Work a cached loop through. An agent re-sends a 60,000-token cached prefix with a 3,000-token new turn, fifty times. On Haiku 5.5 each turn costs 60,000 tokens of cache read at $0.01 per million, which is $0.0006, plus 3,000 tokens of fresh input at $0.10, which is $0.0003: under a tenth of a cent for input, or about 4.5 cents across the fifty turns before output. On Haiku 4.5 the same loop was 45 cents. The saving is real, but notice what carries it: the base input rate, not the ratio.

The Sonnet 5.5 change is the one line in this launch that is not about Haiku. From October 7, Sonnet 5.5's cache reads cost $0.10 per million tokens instead of $0.20, which is 5% of its $2 input. Anthropic's estimate, stated in the launch post, is that this cuts the cost of Sonnet 5.5 on most agentic tasks by around 20%. The per-token input and output prices did not move, so the saving lands only where cache reads dominate the bill. Our cache-first agent architecture guide covers how to find out whether yours does.

05 — SpecificationsContext, output and cutoff against Haiku 4.5.

The price is only half the upgrade. Anthropic's model overview lists a 1M-token context window, a 128K-token synchronous output limit, 300K output on the Batch API with a beta header, adaptive thinking with the effort parameter, and a reliable knowledge cutoff of June 2026. Haiku 4.5 had a 200K window, a 64K output limit, no effort parameter, and a reliable cutoff of February 2025, which is now twenty months stale. On context window, output limit and knowledge cutoff Haiku 5.5 now matches Sonnet 5.5 and Opus 5.5.

SpecificationHaiku 4.5Haiku 5.5
Context window200K tokens1M tokens
Max output, sync64K tokens128K tokens
Max output, Batch API beta—300K tokens
ThinkingExtended thinking (manual budget)Adaptive, effort default medium
Reliable knowledge cutoffFeb 2025Jun 2026
TokenizerPre-4.7Current (≈30% more tokens for the same text)
temperature / top_p / top_kAcceptedSome values return a 400 error; omit them
Retirement commitmentNot sooner than Oct 15, 2026Not sooner than Oct 7, 2027

Two rows are migration work rather than marketing. The sampling row means a client that still passes temperature to Haiku can start receiving 400 errors the moment the model ID changes; Anthropic's migration guide lists the values that fail. The retirement row means Haiku 4.5's floor date is eight days away. Anthropic has not announced a retirement, and a floor is a commitment not to retire before that date rather than a plan to retire on it. But a team still on Haiku 4.5 in November is running a legacy model with no stated runway, and should treat this launch as the deadline it did not have last week.

The thinking row carries a quieter change. Haiku 5.5's thinking blocks work only in the account that produced them, or one linked to it, which matters for anyone who logs full conversations and replays them from a different account. The pattern is the same one we documented for the Sonnet 5.5 migration.

06 — Vendor FiguresWhat Anthropic measured, and what it did not claim.

Anthropic's launch post calls Haiku 5.5 "the cheapest, fastest, and most capable small model we've ever released", and backs it with a benchmark table. Every figure in that table is Anthropic's own run, on Anthropic's chosen settings, with no independent replication on launch day. That does not make them wrong. It makes them a vendor's figures, and the honest way to read them is against the vendor's own larger models, which Anthropic helpfully includes.

Terminal-Bench 4.0 is the clearest example. Haiku 5.5 scores 39.2% where Haiku 4.5 scored 0.0% and GPT-6 Luna, the OpenAI model behind ChatGPT's Free and Go tiers, scored 16.4%. Sonnet 5.5 scores 70.6%. Anthropic's own commentary on that chart says Sonnet 5.5 and Opus 5.5 remain the better choices for complex agentic coding, and that Haiku 5.5 is suited to narrowly scoped tasks that were cost-prohibitive before: compaction, summarisation, subagent work. The benchmark table and the positioning agree, which is not always the case in a launch.

Terminal-Bench 4.0, agentic coding, as reported by Anthropic

Source: Anthropic's Haiku 5.5 launch post, October 7, 2026. Vendor-run; no independent replication at publication.
Sonnet 5.5the larger model, for reference
70.6%
Haiku 5.5the new small model
39.2%
GPT-6 LunaOpenAI's Free and Go tier model, as run by Anthropic
16.4%
Haiku 4.5the model it replaces
0.0%

The other rows tell the same story at different altitudes. OSWorld 2.1 on the offline subset: 72.4% against Haiku 4.5's 15.7% and Sonnet 5.5's 83.9%. Humanity's Last Exam with no tools: 45.9% against 10.2% and 56.9%. GDPval-AA v2.1, a knowledge-work rating: 1620 against 735 and 1840. In each case the new Haiku lands closer to Sonnet than to its predecessor. Whether that holds on your own tasks is the question only your own evaluation answers, and it is the one a vendor table cannot.

The cost claim deserves the same label. Anthropic says Haiku 5.5 costs around 75% less to run than Haiku 4.5 on average. That is a calculation, not a list-price comparison. Anthropic's footnote blends the 90% cut for prompts up to 100,000 tokens with the 50% cut above it, weights them by Haiku 4.5's request mix (90% short), and adjusts for the new tokenizer using more tokens per task. Your blend will differ. The one customer figure in the post comes from Asana, quoted on Anthropic's page:

Compared with the model we use today, we saw over a 30% reduction in latency for task completions and up to 2.5x faster inference per agent turn.Aaron Vinh, Staff Software Engineer, Asana — customer statement on Anthropic's Haiku 5.5 launch page, October 7, 2026

A testimonial chosen by the vendor is weaker evidence than a benchmark the vendor ran, and both are weaker than your own trace. It is quoted here because speed is the claim Anthropic leads with and the only one a named customer attests to on Anthropic's launch page.

07 — RolloutWhere it runs, and what else shipped with it.

Haiku 5.5 is available on launch day across every platform Anthropic sells through: the Claude API, Amazon Bedrock, Google Cloud, Microsoft Foundry and Claude Platform on AWS. The model ID is claude-haiku-5-5 everywhere except Bedrock, which prefixes it as anthropic.claude-haiku-5-5. Anthropic describes no staged rollout and no waitlist.

Three other things shipped in the same announcement, and two of them cost nothing to claim.

Max and Team subscribers
Monthly API credits, rolling out this week
$100–$500 /mo

Max 5x gets $100 a month, Max 20x gets $200, and Team plans receive up to $500 pooled across users, usable on any model through the Claude Platform. Anthropic frames it as a budget for building agents and tools that call the API.

Anthropic launch post
Sonnet 5.5
$0.20 to $0.10 per million, from October 7
50% off cache reads

Anthropic's own estimate is around 20% off most agentic work. Input and output prices for Sonnet 5.5 are unchanged at $2 and $10.

Anthropic pricing page
Python and TypeScript SDKs
Computer use and browser use helpers
2betas

Anthropic says Haiku 5.5's speed and price suit these tasks, and its launch post links the browser-use SDK documentation. Beta status means the interface can change.

Anthropic platform docs
Safeguards
Cyber limits sit between Haiku 4.5 and Sonnet 5.5
3tiers

Anthropic states that Haiku 5.5's cybersecurity safeguards are stricter than Haiku 4.5's but permit more defensive work than Sonnet 5.5's, and that its biology safeguards match Sonnet 5, Sonnet 5.5 and Opus 5.

Anthropic launch post, safety section

08 — DecisionWho should switch this week, and who should wait.

The decision has three inputs: what you run on Haiku 4.5 today, how long your prompts are after re-tokenization, and whether you run anything on Sonnet that only needed Sonnet because Haiku 4.5 was weak. The router below is our reading of Anthropic's published numbers. It is not a quality judgement about Haiku 5.5 on your tasks; run your own evaluation before any move that touches customers.

You run classification, extraction, routing or summarisation on Haiku 4.5 with prompts well under 100K tokens
Switch now. Same job at a tenth of the price, a current knowledge cutoff, and a retirement floor that is a year out instead of eight days. Remove temperature and top_p first.
Haiku 5.5
Your Haiku 4.5 prompts sit between 75K and 100K tokens
Measure before switching. The new tokenizer adds about 30%, and crossing 100K multiplies input by five. Trim the cached prefix to leave headroom, then switch.
Haiku 5.5, after a token count
You send whole documents or long transcripts, 100K to 1M tokens, and today split them for Haiku 4.5
Switch and stop splitting. The long tier is half Haiku 4.5's old rate and a quarter of Sonnet 5.5, and the 1M window removes the chunking code.
Haiku 5.5, long tier
You moved a narrow task from Haiku 4.5 to Sonnet 5 or 5.5 earlier this year because Haiku was not good enough
Re-test on Haiku 5.5 at medium effort. If it passes your evaluation, the saving is twentyfold on short prompts.
Trial Haiku 5.5
You run multi-step agentic coding, long tool loops or anything Terminal-Bench-shaped
Stay on Sonnet 5.5 or Opus 5.5. Anthropic says so itself. Use Haiku 5.5 as the subagent for compaction and summarisation inside that loop.
Sonnet 5.5 / Opus 5.5
You replay stored conversations across accounts, or need US-only inference
Check two things first: thinking blocks are bound to the producing account, and the 1.1× US-only multiplier applies to the long-tier prices as well.
Haiku 5.5, after review

The one move we would not make is the one the headline invites: moving everything onto Haiku 5.5 because it is cheap. Anthropic's own table puts it thirty points behind Sonnet 5.5 on agentic coding, and a small model that fails a task three times costs more than a larger one that passes once. The price list makes the small tasks almost free. It does not make the large ones small. If you want help deciding which of your workloads fall on which side, that is work our AI transformation team does with clients every month.

09 — ConclusionThe cheap tier is real, and so is the line.

What to do, October 7, 2026

Count your tokens on the new tokenizer, move the short-prompt work now, and keep agentic coding where Anthropic says it belongs.

Claude Haiku 5.5 is a tenth of Haiku 4.5's price for prompts up to 100,000 tokens, five times that above the line, with a 1M window, a 128K output limit and a June 2026 knowledge cutoff. Those are Anthropic's published rates and specifications as of October 7, 2026, and every one of them is better than the model it replaces.

The threshold is the detail that decides whether the saving shows up on your bill. It counts every input token including cache reads, it applies per request, and the new tokenizer moves prompts toward it by about 30% before you change anything. A team that measures its largest request on the new tokenizer, trims the cached prefix below 100K with headroom, and drops the sampling parameters will see roughly the 90% the list price promises on short work. A team that does none of that may see a fifth of it.

On capability, take Anthropic at its word in both directions. Its benchmarks are vendor-run and we have labelled them so, but its positioning is candid: Sonnet 5.5 and Opus 5.5 stay in charge of complex agentic coding, and Haiku 5.5 takes the compaction, summarisation, classification and subagent work that was too expensive to automate before. That is a narrower claim than the headline, and a more useful one.

Model routing, measured

A price list is a start — your traffic is the answer.

We price model changes against your actual traffic: token counts on the new tokenizer, cache-hit ratios, the share of requests near a pricing threshold, and a pass-rate test before anything moves. The result is a routing table you own, not a vendor's headline.

Free consultationExpert guidanceTailored solutions
What we work on

Model cost engineering

  • →Token-count audits across a model migration
  • →Cache-prefix sizing against pricing thresholds
  • →Pass-rate evaluations before a model swap
  • →Routing tables that send each task to the cheapest model that passes
  • →Monthly spend review against vendor price changes
FAQ · Claude Haiku 5.5

The questions we get every week.

For prompts up to 100,000 tokens, $0.10 per million input tokens and $0.50 per million output tokens on the Claude API, with cache reads at $0.01 and 5-minute cache writes at $0.125. For prompts over 100,000 tokens, every rate is five times higher: $0.50 input, $2.50 output, $0.05 cache read and $0.625 cache write. Batch processing is 50% off in both tiers. These are Anthropic's published list prices as of October 7, 2026; Amazon Bedrock and Google Cloud set their own regional pricing.
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.

Related dispatches

Continue exploring model pricing and selection.

AI Development

Fable 5 Cost Engineering: Cache, Batch and Spend Caps

Claude Fable 5 lists at $10/$50 per million tokens and meters after July 7. How prompt caching, the Batch API, and spend caps cut the bill — with worked math.

July 2, 2026 · 12 minRead
AI Development

AI Model API Price Changes, Q3 2026: Five Labs, One Table

API price changes from Anthropic, OpenAI, Google, DeepSeek and xAI, July to September 2026: launches, cuts, a rise and promotions per million tokens, dated.

October 3, 2026 · 7 minRead
AI Development

Anthropic: Open GLM-5.3 Nearly Matches Mythos at Exploits

In Anthropic's tests the open GLM-5.3 model succeeded at 50 of 410 exploit attempts, close to Claude Mythos Preview at 56. What the report shows and omits.

September 29, 2026 · 7 minRead
AI Development

Migrating to Claude Sonnet 5.5: Every Breaking Change

Claude Sonnet 5.5 rejects thinking disabled, forced tool choice and three more settings. The exact errors and fixes, grouped by the model you run today.

September 28, 2026 · 6 minRead
AI Development

Deleting AI Agent Memory: Where Stored Copies Survive

Deleting AI agent memory takes more than clearing a chat. Map stored copies, retrieval indexes and backups, then verify what your system can still recover.

September 4, 2026 · 7 minRead
AI Development

Preview, Beta, GA: What Vendors Said vs What Coverage Said

A 36-row comparison of announcement and coverage wording, including control records, date exceptions and cases where independent coverage was not located.

August 22, 2026 · 27 minRead
Google Search

See more Digital Applied analysis in your Google results by adding us as a preferred source.

Add as a preferred source