Data residency pricing for AI APIs is no longer a negotiated mystery: as of August 25, 2026, four major vendors — OpenAI, Azure OpenAI, Google Vertex AI, and Mistral — publish an explicit, isolable surcharge for pinning inference to a region or data zone, and all four converge on the same number: 10%, a 1.1× multiplier on standard list pricing.
That convergence is the finding, but the outlier is the story. AWS Bedrock’s own price list shows Anthropic’s Claude 3.5 Sonnet at an identical rate in Virginia, Ohio, Oregon, Frankfurt, Ireland, Zurich and Paris — no residency line item at all — while Mistral and Gemma models on that same list carry non-uniform, double-digit regional deltas that AWS never labels “residency”. Whatever the residency premium is, it is not a law of nature.
This is a billing reference, not an architecture guide. Every rate below comes from a primary pricing or documentation page retrieved on August 25, 2026, with its unit and its baseline stated. For region-pinned inference patterns, redaction, sovereign overlays and the compliance map, see the architecture side of data residency — the implementation companion to this page.
- 01Four vendors publish the same 10% residency surcharge.OpenAI states a 10% uplift for regional processing endpoints and Mistral states 1.1× list pricing for its EU/US regional endpoints. Azure OpenAI (Data Zone and Regional deployments) and Google Vertex AI (non-global endpoints) publish separate price rows that work out to the same ~10%. All four as of August 25, 2026.
- 02AWS Bedrock’s price list is the counter-example.Claude 3.5 Sonnet costs $6.00 / $30.00 per 1M tokens across the seven US and EU regions listed for it — a flat 0% delta. Mistral Large 3 (+17-18% EU) and Gemma 3 27B (+55-57% London) vary on the same list, and AWS labels none of it residency pricing.
- 03The 10% surcharge is scoped, not universal.OpenAI’s uplift applies only to models released on or after March 5, 2026 that are eligible for data residency. Google’s non-global pricing took effect July 1, 2026. Mistral’s regional endpoints exclude Agents, Batch and the Files API, and leave billing metadata outside the region.
- 04Do not confuse a 10% feature with a 75% tier.Mistral’s Enterprise APIs (residency controls bundled with SLAs, rate limits and support) run 75% above list pricing on select APIs, per Mistral’s own wording. AWS Bedrock’s Priority tier is also +75% over its Standard tier — and is not residency-specific at all. Anthropic’s direct API publishes no regional rate whatsoever.
- 05The 1.1× figure now lives inside the tooling.Claude Code 2.1.239 (August 21, 2026) bakes a 1.1× US-only-inference premium for data-residency workspaces into /cost, the status line and --max-budget-usd — the same 10% figure the cloud price lists publish, surfacing in a coding agent’s own budget math.
01 — The FindingOne number, four price lists.
Ask each vendor what it costs to keep inference inside a named geography and the answers arrive in different vocabularies — “regional processing” at OpenAI, “Data Zone” at Azure, “non-global endpoints” at Google, “regional inference” at Mistral. Convert them all to the same unit and the vocabulary differences dissolve: each one is a 10% uplift on the per-token list rate. OpenAI’s pricing page states it plainly: “Regional processing (data residency) endpoints are charged a 10% uplift for models released on or after March 5, 2026, that are eligible for data residency.” Mistral’s developer docs use the identical arithmetic: “Regional inference is billed at 1.1× standard list pricing (a 10% upcharge) for input tokens, output tokens, cached reads, and cache writes.”
To be precise about what this convergence is: an observation, not a standard. No source states that these vendors are matching each other. Four companies independently priced the same constraint and landed on the same multiplier — which is exactly what makes the one price list that did not land there worth reading closely.
Four vendors, one number
OpenAI, Azure OpenAI, Google Vertex AI and Mistral each publish a 10% (1.1×) surcharge for region-pinned inference on their own pricing or docs pages — independently, in four different vocabularies.
Claude on AWS Bedrock
$6.00 / $30.00 per 1M tokens for Claude 3.5 Sonnet in Virginia, Ohio, Oregon, Frankfurt, Ireland, Zurich and Paris alike. A published price table with zero delta between US and EU — and no residency line item.
The enterprise-shaped premium
Mistral’s Enterprise APIs (residency controls + SLAs + rate limits + support) run 75% above list pricing on select APIs, and AWS Bedrock’s Priority tier runs 75% above its Standard tier. A different product shape entirely — and not an isolated residency price.
02 — Pricing ReferenceThe residency tax, by vendor.
The table below is the reference this post exists to provide: every published residency or regional price differential we could source from a primary vendor page, in one unit, with its baseline named. All rates were retrieved on August 25, 2026, from the linked pages.
| Vendor / product | Baseline (per 1M in / out) | Region-pinned (per 1M in / out) | Delta | Named a residency charge? |
|---|---|---|---|---|
| Published, isolated surcharges — the vendor states the number | ||||
| OpenAI API — models released ≥ Mar 5, 2026 | Standard global endpoint | Regional processing endpoint | +10% (1.1×) | Yes — “regional processing (data residency)” uplift, stated as a percentage on the pricing page |
| Azure OpenAI — GPT-5.5 | Global $5.00 / $30.00 | Data Zone $5.50 / $33.00 | +10% | Yes — Data Zone (EU or US) and Regional deployment types are priced as distinct rows |
| Azure OpenAI — GPT-5 | Global $1.25 / $10.00 | Data Zone $1.38 / $11.00 | ~+10% | Yes — same ~10% factor holds on cache-write and cache-read rates for the same rows |
| Azure OpenAI — o1 | Global $15.00 / $60.00 | Data Zone $16.50 / $66.00 | +10% | Yes — the multiplier is uniform across essentially every model row checked, not a flat fee |
| Google Vertex AI — Gemini 3.5 Flash | Global $1.50 / $9.00 | Non-global $1.65 / $9.90 | +10% | Yes — non-global (regional) pricing, in effect for Gemini 3+ models since July 1, 2026 |
| Google Vertex AI — Gemini 3.7 Flash | Global $0.75 / $3.75 | Non-global $0.825 / $4.125 | +10% | Yes — docs state global endpoint requests “are charged at a lower price” as a rule |
| Mistral — regional inference, all supported models | Global api.mistral.ai, no upcharge | EU / US regional endpoint, 1.1× list | +10% | Yes — applies to input, output, cached reads and cache writes, per the docs |
| The same price list, no residency line — AWS Bedrock | ||||
| AWS Bedrock — Claude 3.5 Sonnet | US East (N. Virginia) $6.00 / $30.00 | Frankfurt / Ireland / Zurich / Paris $6.00 / $30.00 | 0% — no delta | No — identical published rate across all seven listed US and EU regions |
| AWS Bedrock — Mistral Large 3 | US East $0.50 / $1.50 | Europe (Ireland, Milan) $0.59 / $1.76 | +18% / +17% | No — the page never uses residency language for these deltas; they read as regional infrastructure pricing |
| AWS Bedrock — Google Gemma 3 27B | US regions $0.23 / $0.38 | Europe (London) $0.36 / $0.59 | +57% / +55% | No — same list, wildly different delta than Claude’s 0%, and still no residency label |
| Inside the tooling — the multiplier as budget math | ||||
| Claude Code CLI — v2.1.239, Aug 21, 2026 | Standard inference cost estimate | Data-residency workspace estimate | 1.1× (+10%) | Yes — applied in-product inside /cost, the status line and --max-budget-usd; not a separate vendor charge |
Sources: the OpenAI API pricing page, the Azure OpenAI Service pricing page, the Vertex AI generative AI pricing page, the Vertex AI data residency documentation, the Mistral regional inference docs, the Mistral API pricing page, the Amazon Bedrock pricing page, Anthropic’s regional compliance page and the Claude Code changelog. All retrieved August 25, 2026.
03 — Fine PrintThe 10% is real — and scoped.
OpenAI, Google and Mistral each attach scoping language to their published surcharge that matters more to a budget than the headline percentage, because it determines whether the model you actually run is covered at all.
Regional processing
The stated uplift covers models released on or after March 5, 2026 that are eligible for data residency — older models sit outside the policy language. OpenAI also disclaims parity with Bedrock-hosted OpenAI models, which are billed through AWS and may differ from direct pricing.
Three deployment types
The ~10% factor holds across essentially every model row checked — GPT-5.5, GPT-5.2, GPT-5.1, GPT-5, GPT-4.1, o3, o1 — and scales cache-write, cache-read and output rates by the same 1.1×. It is a multiplier, not a flat fee. Regional deployments reach up to 27 local regions.
Non-global is new
The non-global surcharge applies to generally available Gemini 3 and later model families, and only since July 1, 2026 — before that date, global endpoint pricing applied to non-global endpoints too. Google’s docs state global endpoints are cheaper as a rule.
Regional, with carve-outs
Three named endpoints: global (no upcharge, no residency guarantee), EU and US (1.1× list, pinned to data centers in the named geography). Function Calling is the only supported regional tool call — Agents, Batch and the Files API are unavailable on regional endpoints.
Two details in that grid deserve to be read twice. First, Mistral’s docs are explicit that regional inference does not make everything regional: “account configuration, API keys, billing, access management, usage analytics, and other operational metadata may still be handled by Mistral systems outside the selected inference geography.” You are buying region-pinned inference, not a region-pinned vendor relationship — a distinction a compliance team will care about even when procurement does not. The wider story of what runs on those endpoints is covered in our post on Mistral’s EU regional endpoints.
Second, the Azure baseline is about to move. The Azure pricing page carries this banner: “Notice: Starting September 1, 2026, pricing increases for Microsoft Foundry EU Data Zone and non-US regional deployments.” No magnitude is disclosed — the page states that an increase is coming, dated, with no percentage, as of the August 25 retrieval. The 10% delta mechanism predates that announcement; what changes on September 1 is the baseline it multiplies.
04 — The OutlierOn Bedrock’s own price list, Claude’s residency delta is zero.
This is the finding that makes the table more than four vendors saying “10%”. AWS Bedrock publishes per-region rates for the models it hosts, and for Anthropic’s Claude 3.5 Sonnet the number is the same everywhere it is listed: $6.00 input / $30.00 output per 1M tokens in US East (N. Virginia), US East (Ohio), US West (Oregon), Europe (Frankfurt), Europe (Ireland), Europe (Zurich) and Europe (Paris). Running Claude in Frankfurt costs exactly what it costs in Virginia. There is no residency line item because there is no difference to itemize.
Now read the rest of the same price list. Mistral Large 3 runs $0.50 / $1.50 in US East but $0.59 / $1.76 in Europe (Ireland, Milan) — an 18% / 17% delta — and $0.61 / $1.82 in São Paulo and Tokyo, yet only $0.5150 / $1.5450 in Sydney, a 3% bump. Google’s Gemma 3 27B is $0.23 / $0.38 in US regions but $0.36 / $0.59 in Europe (London), a 57% / 55% jump — while its Sydney rate is roughly flat to slightly cheaper than the US. DeepSeek v3.2 runs about 19-20% higher in its non-US region group than in US regions.
EU vs US input-rate delta on the same Bedrock price list
Source: Amazon Bedrock pricing page, input-token rates, retrieved August 25, 2026One honest caution about how to read this. AWS never calls any of these deltas “residency pricing” — the page does not use that language, and the variance reads as ordinary regional infrastructure pricing that happens to coincide with residency-relevant regions. It would be a fabrication to relabel Gemma’s London premium a residency tax.
But that is precisely the point. The one hyperscaler price list that shows real per-region rates demonstrates that region-pinned Claude can be sold at zero premium in the EU while other models on identical infrastructure carry double-digit, non-uniform deltas. The 10% figure four vendors converged on is a pricing decision, not a cost pass-through — or at minimum, whatever costs it recovers are costs some vendor-model combinations demonstrably do not incur or do not bill. Anthropic’s own regional-compliance page, for its part, points European buyers at exactly this route: residency guarantees run “through AWS Bedrock, GCP Vertex, and Microsoft Foundry”, with Bedrock covering Frankfurt, Ireland and Paris among others, and Canada plus additional European options listed as coming in 2026.
05 — Two Different PremiumsA 10% feature is not a 75% tier.
Most coverage of “paying more for residency” conflates two things the price lists keep separate. The first is an isolated surcharge: same model, same service level, region-pinned, priced as its own line — that is the 10% club. The second is a bundled enterprise tier where residency controls arrive wrapped in SLAs, rate limits and support, and the residency component cannot be bought or even priced alone. Mistral’s own pricing page advertises exactly that: “Introducing our Enterprise APIs, including regional data processing controls, system-level SLAs, increased rate limits, and premium support. This service is available for 75% above list pricing on select APIs.”
The census vocabulary from this batch’s cost series fits the hard cases here. Some vendors publish an isolable residency price (OpenAI, Azure, Google, Mistral’s regional endpoints). Some bundle without isolating — Mistral’s Enterprise APIs fold residency into a 75% tier, and AWS Bedrock’s Priority tier is the same 75% shape without being about residency at all (it buys latency and throughput, alongside a Flex tier at 50% off Standard). And one vendor is effectively silent: Anthropic’s regional-compliance page for its direct API says regional endpoint pricing “may reflect dedicated regional infrastructure costs” and publishes no rate or percentage at all — while noting its global endpoints carry no pricing premium.
| Offer | Premium | What it bundles | Residency price isolable? |
|---|---|---|---|
| Isolated surcharges — residency as its own line item | |||
| OpenAI regional processing | +10% | Residency only — same model, region-pinned processing | Yes |
| Azure OpenAI Data Zone / Regional | ~+10% | Deployment type only — EU/US zones or up to 27 local regions | Yes |
| Google Vertex AI non-global endpoints | +10% | Endpoint locality only — ML processing stays in the named geography | Yes |
| Mistral regional endpoints | +10% (1.1×) | Regional inference only — with Agents, Batch and Files API unavailable, and control-plane data possibly outside region | Yes, with carve-outs |
| Bundled or unpublished — residency has no price of its own | |||
| Mistral Enterprise APIs | +75% | Regional data processing controls + system-level SLAs + increased rate limits + premium support — on select APIs | No — bundled |
| AWS Bedrock Priority tier | +75% vs Standard | Service quality (latency / throughput) — not residency-specific at all; Flex tier runs 50% below Standard | No — different product |
| Anthropic direct API regional endpoints | Unpublished | Page states pricing “may reflect dedicated regional infrastructure costs” — no rate or percentage given | No — no published number |
Isolated surcharge vs bundled tier
Sources: vendor pricing pages listed above, retrieved August 25, 2026The practical consequence: when a vendor quotes “residency” at something far above 10%, ask what else is in the box. As of the August 25 retrieval, no vendor checked publishes a residency premium as anything other than a percentage of the base rate — no flat monthly fee, no minimum-commitment gate for the plain “same model, region-pinned” case. Minimums and gates do exist, but they attach to the broader Enterprise and Priority tiers that bundle far more than residency. For the structural version of the EU sovereignty question — where the vendor itself, not just the endpoint, is the compliance answer — see Mistral’s sovereign-AI positioning.
06 — The PegThe 1.1× has moved into the tooling.
The clearest sign that the 10% residency premium has hardened from a pricing-page footnote into operational reality is where it showed up next: inside a coding agent’s own budget arithmetic. Claude Code 2.1.239, released August 21, 2026, added residency-aware cost accounting to the CLI itself.
Read what that means in practice. If your organization runs Claude Code in a data-residency workspace, the /cost command, the status line, and the --max-budget-usd spend cap all now multiply their estimates by 1.1 — the same 10% figure that OpenAI, Azure, Google Vertex and Mistral publish on their pricing pages. A budget cap set before upgrading to 2.1.239 buys less residency-pinned work after it, because the estimate now includes the premium. Teams budgeting multi-hour autonomous runs should fold this into their math — our companion piece on long-horizon agent run budgeting covers what spend limits do to a job mid-flight.
One discipline note: nothing ties these numbers together causally. No source states that vendors are matching each other, and no source states that Anthropic derived its tooling multiplier from competitors’ pages. What the record supports is narrower and more interesting — five independently published 1.1× figures, from four cloud pricing pages and one CLI changelog, all live on the same August 2026 retrieval date.
07 — Procurement PlaybookWhat a buyer does with this table.
A published 10% premium is, above all, a negotiating baseline and a budgeting constant. Here is how the reference translates into decisions for the four buyer situations we see most often.
Budgeting residency-pinned inference
Model 1.1× on every per-token rate — input, output, and cache traffic, since Azure and Mistral apply the multiplier to cache rates too. Then check the Bedrock route for your model: if it is Claude, the published EU rate carries no premium at all.
No published number to plan on
Anthropic’s direct API publishes no regional rate — only that regional pricing 'may reflect dedicated regional infrastructure costs' and that residency guarantees run through Bedrock, Vertex and Foundry. Plan the cloud-platform route, where the prices are published.
Separate the 10% from the 75%
If a quote prices 'residency' far above 10%, it is almost certainly a bundled tier — SLAs, rate limits, support. Ask for the isolated regional-endpoint rate as a comparison line before valuing the bundle.
The self-hosting comparison
A permanent 10% cloud surcharge on a large workload is a real number — and the honest alternative is not a cheaper cloud, it is owning the region entirely. Weigh the surcharge against self-hosting TCO before treating 10% as unavoidable.
Two forward-looking notes belong in any plan built on this table. First, the baselines are moving in one direction. Google’s non-global surcharge only took effect on July 1, 2026, and Azure has announced an EU Data Zone increase for September 1, 2026 with no disclosed magnitude. Residency pricing is young, and the pattern so far is premiums being introduced and raised, not competed away — the Bedrock counter-example is the one data point suggesting competition could eventually push the number toward zero, at least for models whose economics allow it. Treat every rate here as a dated snapshot, not a stable constant.
Second, the surcharge compounds with everything else in your token bill. A residency multiplier applies before you hit spend caps, and what happens at those boundaries differs sharply by vendor — see this batch’s companion census on what vendors actually do when a spend cap exhausts. And if the 10% line is what finally makes cloud inference feel expensive, the comparison worth running is self-hosting as the alternative to a residency premium — capex against a permanent percentage. For teams that want help pressure-testing an AI vendor decision like this against their own workloads, our AI transformation engagements start with exactly this kind of comparative pricing eval.
08 — ConclusionA pricing decision wearing a compliance costume.
Residency costs 10% where it is priced at all — and zero where it is not.
As of August 25, 2026, the published state of AI data-residency pricing is this: OpenAI, Azure OpenAI, Google Vertex AI and Mistral each charge an isolated 10% (1.1×) surcharge to pin inference to a region — with OpenAI, Google and Mistral each scoping theirs by model eligibility, effective date, or feature carve-outs. Enterprise bundles that include residency run 75% above list pricing on select APIs at Mistral, and AWS’s similarly shaped Priority tier is not about residency at all. Anthropic’s direct API publishes no regional rate.
The row that keeps the table honest is the one with a zero in it. Claude on AWS Bedrock costs the same in four EU regions as in Virginia, on a published price list, with no residency line item — while other models on that same list carry uneven regional deltas that AWS never calls residency. The 10% premium is real, published, and now embedded in tooling budget math. It is also, demonstrably, a choice.
Use the table as a baseline: budget 1.1× where you are pinned, unbundle any quote that prices residency above 10%, check the cloud-platform route for your specific model before assuming the premium applies, and re-verify every rate against the linked primary pages — this market is young enough that the numbers carry dates for a reason.