AI DevelopmentCost Playbook3 min readPublished October 3, 2026

Which AI bills you can actually stop, and which only email you

AI Spend Controls Compared: Usage Caps Across 8 Vendors

How Anthropic, OpenAI, Google, Microsoft, AWS and three others let teams cap AI spend: hard limits, alerts, per-user budgets and what happens at the cap.

DA
Digital Applied Team
Research and practical guidance
CoverageOctober 3, 2026

Of eight AI vendors we checked on October 3, 2026, six let a team set a cap that stops usage when spend reaches it. Microsoft’s Azure budgets only send alerts, and AWS stops usage only if you wire up a budget action yourself. Even the real caps can overshoot, because billing data arrives minutes to hours late and long-running jobs finish after the line is crossed.

Key takeaways
  1. 01
    Alerts are not capsAzure budgets, OpenAI spend alerts and GitHub seat budgets notify you; they do not stop anything.
  2. 02
    Most APIs can stopAnthropic, OpenAI, Google and OpenRouter can stop usage at a cap you set.
  3. 03
    Caps run lateOpenAI and Google both warn that spend can pass the cap before enforcement catches up.
  4. 04
    Scope mattersPer-project, per-key and per-user caps contain one runaway job; an account cap stops everything.

01 — ContextCaps, alerts and actions are different things

“Spend control” covers three mechanisms, and vendors use the same words for all of them. The difference only shows on the day an agent loops overnight.

A
Hard cap
Anthropic, OpenAI, Google, OpenRouter, Cursor, GitHub metered

Requests fail with an error once spend reaches the limit, until the period resets or someone raises it.

Stops usage
B
Alert
Azure budgets, OpenAI spend alerts, GitHub seats

An email or notification at a threshold. Traffic carries on.

Tells you
C
Automated action
AWS budget actions

A rule that changes permissions when a threshold is hit, which can block further use if you design it to.

You build it

02 — The dataEight vendors compared

The table covers the controls each vendor documents for its AI products or, for Azure and AWS, the account-wide cost tools that cover their AI services.

Sources: each vendor’s spend-limit, billing or budget documentation, read October 3, 2026.
VendorCan it stop usage?ScopeWhat happens at the cap
Anthropic Claude APIYesOrganisation and workspace; per session for Managed AgentsTier cap: HTTP 429 with no retry-after until the 1st of next month. Your own limit: HTTP 400. Agent session pauses, in-flight request finishes
OpenAI APIYes, if you turn on a hard limitOrganisation and projectHTTP 429 with a spend-limit code; can slightly overshoot; resets monthly. Alerts never stop traffic
Google Gemini APIYesBilling-account tier cap and per-project caps (experimental)Overage possible: billing data can lag about 10 minutes, and batch jobs and agent sessions can run past the cap. Not for invoiced accounts
Microsoft Azure budgetsNoSubscription, resource group or wider scopesEmail alerts, or an action group that can call your own automation; Microsoft says consumption is not stopped. Costs evaluated every 24 hours
AWS BudgetsOnly through a budget actionAccount or organisationAn action you configure can apply an IAM policy or service control policy, automatically or after approval
GitHub CopilotFor metered AI credits, not seatsEnterprise, organisation, cost centre, repository or userMetered usage can be stopped at the budget; seat licences only get alerts
Cursor TeamsYes, team-wideTeam; per member on EnterpriseLimits apply to on-demand usage after included usage, which is on by default
OpenRouterYesAccount balance and each API keyHTTP 402 naming the limit that was hit; per-key limits can reset on a schedule

Two newer controls stand out. OpenAI added hard spend limits for organisations and projects on July 22, 2026. Anthropic added session budgets for Claude Managed Agents (in beta) on August 7, a dollar cap on one agent session priced at public list rates. A paused session keeps its history and can resume if the budget is raised.

GitHub’s budgets documentation draws the line most teams miss: a budget on a seat licence such as Copilot only warns, while a budget on metered Copilot AI credits can stop usage, and can be set per user.

03 — The dataCaps you get without setting anything

Anthropic and Google also apply a monthly ceiling by account tier, whether or not you set your own. Reaching one affects every app on the account, not just the one that caused it.

Google Gemini API, Tier 1Billing account linked
$250 / month
Anthropic, Start tierMonthly spend cap
$500 / month
Anthropic, Build tierMonthly spend cap
$1,000 / month
Google Gemini API, Tier 2$100 paid and 3 days since first payment
$2,000 / month
Google Gemini API, Tier 3$1,000 paid and 30 days since first payment
$20,000 to $100,000+
Anthropic, Scale tierCustom tier has no cap
$200,000 / month

Anthropic’s rate limits page says that once an organisation reaches its tier cap, the API returns HTTP 429 until midnight UTC on the first of the next month, with no retry-after header, so automatic retries keep failing. OpenAI separately assigns each organisation an approved monthly usage limit by tier, distinct from the limits a team sets.

04 — The catchWhy even hard caps overshoot

Caps are only as fast as the billing data behind them. OpenAI says enforcement is not instantaneous and recorded spend can slightly exceed the limit. Google’s Gemini API billing page puts the lag at up to about 10 minutes and warns that batch jobs and agent sessions can run past a project cap. Azure’s budget alerts rely on cost data that typically arrives 8 to 24 hours late and is evaluated every 24 hours.

Anthropic’s session budgets show the same effect at small scale: the check happens between model requests, so the request that crosses the cap still completes. OpenRouter takes the opposite approach for concurrent traffic on smaller prepaid balances, estimating each paid request’s token cost up front and refusing new requests once the running ones fill an in-flight budget set as a fraction of the balance.

Set the cap below the real limit

If the most you can afford to lose in a month is $5,000, do not set the cap at $5,000. Leave room for an hour or a day of overshoot at your peak spending rate, and set alerts well before the cap so a person sees the trend first.

05 — Practical implicationsA spend setup that holds

Several apps share one API account
One project, workspace or key per app, each with its own cap
Contain one runaway
Long-running agents
Per-session budgets where offered, plus a project cap
Anthropic, OpenRouter keys
Models bought through Azure
Budgets only alert; build your own shut-off
Azure
Models bought through AWS
A budget action that applies a deny policy at a threshold
AWS Budgets

A cap is the last line, not the plan. The first is knowing what each workflow should cost, which our agent token budget framework covers, and the second is prices that can change under you, which our Q3 price change tracker follows. For long agent runs specifically, see budgeting long-horizon agent runs. Teams that want caps, alerts and per-workflow budgets designed together can work with our AI transformation team.

Next step

Test what your cap does before you need it

In a sandbox project, set a tiny cap, run a loop past it and watch what your application does with the error. A cap that turns into a silent retry storm, or a 400 your code treats as a bad prompt, is worse than an alert.

AI cost control

Put a ceiling on every AI workflow you run

Digital Applied sets up per-project caps, alerts and agent budgets across your AI vendors, then tests that your systems fail safely when a cap is hit.

Per-project capsAlert thresholdsFailure tests
Before the next invoice

Check four things

  • →Is it a cap or an alert?
  • →What scope does it cover?
  • →How late can it fire?
  • →What does our code do then?
Questions and answers

Practical questions

Yes. Since July 22, 2026, organisations and projects can enforce a monthly hard spend limit. Requests then return a 429 error with a spend-limit code until the limit is raised or the month resets.
Digital Applied newsletter

Deep dives on AI, marketing and development.

Practical guides and fresh insights by email. No recycled takes.

Related dispatches

Continue reading

AI Development

Prompt Caching Changed at OpenAI and Anthropic This Week

OpenAI's 30-minute cache window, breakpoints and effort changes without a miss; Anthropic's diagnostics GA, inline tools and newly billed refusals. What to do.

September 24, 2026 · 4 minRead
AI Development

Which AI Coding Agents Let You Intercept Tool Calls?

Claude Code, Codex, Cursor, Gemini CLI, Copilot and seven more compared: which events each hook system exposes, what it can block or rewrite, where it runs.

October 3, 2026 · 6 minRead
AI Development

Reported FTC AI Safety Probe: What Teams Should Record

The FTC is reportedly investigating AI agent safety at OpenAI and Anthropic. What businesses running agents should document now, from permissions to logs.

September 30, 2026 · 5 minRead
AI Development

GitHub Copilot Features Turn On by Default Oct 22: Check Now

From October 22, unconfigured GA Copilot features, including MCP servers and code review, turn on for Business and Enterprise unless an admin decides first.

September 24, 2026 · 5 minRead
AI Development

Which AI Agent Features Fall Outside Zero Data Retention

Compare 31 agent features across Anthropic, OpenAI, Google and AWS, with dated sources for retention, ZDR eligibility, contract scope and deletion.

September 8, 2026 · 9 minRead
AI Development

Deleting AI Agent Memory: Where Stored Copies Survive

Deleting AI agent memory takes more than clearing a chat. Map stored copies, retrieval indexes and backups, then verify what your system can still recover.

September 4, 2026 · 6 minRead
Google Search

See more Digital Applied analysis in your Google results by adding us as a preferred source.

Add as a preferred source