Skip to main content

Simon Willison argued on 3 October that hard budget caps need to become a default feature of pay-by-usage services: “after $X/month, cut this thing off and return errors”. The post reached 552 points on Hacker News inside a day, which tells you how many people have been quietly frightened of their own infrastructure. AWS shipped project-level spend limits in September, and Google Cloud added spend cap budgets over the summer. The feature is arriving.

So we read the billing documentation for thirteen providers we actually deploy on, and recorded two things: whether a hard cap exists, and what your application sees at the moment it fires. The second question is the one nobody is discussing, and it is the one that will page you at 3am.

TL;DR

  • Of thirteen providers checked, ten can hard-stop on spend and one offers it only on credit-based subscriptions. Cloudflare’s own docs state its budget alerts “are informational only and do not pause or cap usage”.
  • The same event, hitting your ceiling, arrives as HTTP 400, 402, 429 or 503 depending on supplier. Two of the biggest return 429, the status every SDK retries by default, and Anthropic’s docs confirm those retries simply fail until the next month.
  • Every hard cap documents an overshoot. Google Cloud’s own wording: “enforcement of spend caps isn’t instant and any cost overages are billed as normal.”
  • Nothing resumes by itself. AWS goes further: a paused project whose owner takes no action for 90 days has its data permanently deleted.
  • The cap rarely covers the whole bill. Supabase excludes Compute, read replicas and PITR; Google Cloud limits each cap to one project and one eligible service.

Who actually has one

The hard cap has gone from rare to common in about a year, and no two implementations behave alike.

Provider Hard stop On by default? What happens
Anthropic API Yes Yes, a tier cap Requests pause until 00:00 UTC on the 1st
OpenAI API Yes, opt-in toggle No, alerts only Requests fail for the org or project
OpenRouter Yes, prepaid credits Yes, structurally Requests rejected, in-flight spend counted
AWS (AWS Settings) Yes, limited release No Project paused, all resources stopped
Google Cloud Yes, spend cap budgets No One service in one project paused
Azure Credit subscriptions only Yes, on those Subscription disabled at the credit amount
Vercel Yes Yes, pauses production Production deployments stop serving
Railway Yes Agent spend only Workloads taken offline
Supabase Yes, Spend Cap Configurable, Pro plan Capped usage items disallowed
Neon Yes, API only No All computes for the project suspended
GitHub Yes, budget option Only for user budgets Metered usage blocked
Cloudflare No n/a Email alert, usage continues
Hetzner Per-server price cap Structural Email alert on overage

One event, four status codes

A spend cap is not a billing setting. It is a new failure mode in your request path, and the industry has not agreed on how to signal it.

Anthropic’s tier cap returns HTTP 429 with error type rate_limit_error, no retry-after header, and the discriminator buried in error.details.error_code as enforced_spend_limit_reached. The documentation is blunt about the consequence: “Retrying, including the SDK’s automatic retries, fails until access resumes.” A limit you set yourself returns something different again, HTTP 400 with type invalid_request_error.

OpenAI’s hard limit also returns 429, with organization_spend_limit_exceeded or project_spend_limit_exceeded in the error code. OpenRouter returns 402 and tells you to branch on error.metadata.limit_source rather than the message text. Vercel’s paused projects serve visitors a 503 DEPLOYMENT_PAUSED.

So: 400, 402, 429 and 503, for one logical event. And in the infrastructure cases, AWS, Google Cloud, Railway, Neon, the service is simply not there any more, so your caller sees connection failures and timeouts with no error body to inspect at all.

The 429 cases are the dangerous ones. Every HTTP client, SDK retry policy, job runner and queue consumer in your estate is configured to treat 429 as transient and back off. Turn a cap on without touching that code and you have built a retry loop against a wall that will not open for days, burning your rate limit budget and filling your logs while the thing you were protecting stays down. The only way to tell the two apart is a vendor-specific string inside the body. None of it is in the status code.

Every hard cap overshoots

“Hard” turns out to be a claim about the reported number, not the real one. Google Cloud’s documentation says enforcement “isn’t instant and any cost overages are billed as normal”, and recommends setting the budget slightly below your absolute limit. OpenAI states that enforcement “is not instantaneous” and recorded spend “can slightly exceed the configured amount”. Vercel checks usage every few minutes and warns that pausing can lag the crossing by several minutes.

For ordinary workloads that is noise. For the scenario caps exist to catch, a runaway agent or a recursive job, the lag is the entire risk, because the burn rate during those minutes is the pathological one. Set the cap meaningfully below the number you genuinely cannot exceed, and alert well below that.

Nothing resumes itself, and AWS starts a clock

None of the hard caps we checked restore service automatically when you raise the limit. Google Cloud pauses the service “until you manually lift the spend cap”. Vercel is explicit that projects will not unpause when you increase the spend amount; each one has to be resumed individually through the dashboard or REST API. Neon’s suspended computes stay suspended until the next billing period unless you reset the quota.

AWS has the line that should make every reader stop and check their own account. Reaching a project spend limit pauses the project and stops all resources, and then: “If you take no action within 90 days of your project being paused, AWS permanently deletes your project data.” A cost control with a data destruction deadline attached is still worth having, but it belongs in your runbook and your calendar, not in a billing page you visited once.

The cap rarely covers the whole bill

“We have a spend cap” and “our spend is capped” are different claims. Google Cloud limits each spend cap to one project and one eligible service, drawn from a short list: the Gemini API, the Gemini Enterprise Agent Platform, Cloud Run and Cloud Run functions. Supabase’s Spend Cap explicitly excludes Compute, Branching Compute, Read Replica Compute, provisioned disk IOPS and throughput, IPv4 addresses and Point-in-Time Recovery, which is to say most of what a production bill is made of. Vercel’s pause does not stop AI Gateway or v0 usage, which keeps accruing against the same amount. Railway caps agent spend by default at $5 on Hobby and $20 on Pro, while compute has no default cap at all.

There is also a floor. AWS sets the minimum spend limit at the greater of $20 or “a conservative estimate of your likely spend”, so a busy account must shut resources down before it can set a lower number. Anthropic’s self-set limit cannot exceed your tier’s cap. You can cap the growth above your current burn. You cannot cap the burn itself.

The argument against, taken seriously

A hard cap is a denial of service you have pre-authorised against yourself. Anyone who can drive your metered usage, through a scraper, a retry storm or a deliberate attack, can convert your own control into an outage, where the alert-only configuration would have degraded to a bill. That is a real trade, and it is why the cap belongs above your realistic ceiling rather than near your average. It is not a reason to go without one. Willison’s framing holds: most teams would rather serve errors than discover a five-figure invoice, and the ones who would not should have to tick a box saying so.

What to do this week

  1. Inventory your metered suppliers and put each into one of three buckets: hard cap available, alerts only, nothing.
  2. Write the cap event into your error handling before you enable anything. Branch on the vendor error code ahead of the status class, and make sure a cap response is never retried.
  3. Set the cap below your true ceiling to absorb the documented lag, with alerts at 50% and 80%.
  4. Write the un-pause runbook now, including which projects need individual resumption and who holds the billing permission at 3am. Diarise the AWS 90-day deletion window.
  5. Check coverage line by line, because the items excluded from a cap are usually the expensive ones.

This is where cost control gets dropped: finance assumes the cap is a technical control, engineering assumes it is a billing setting, and nobody owns the error path between them. Our earlier piece on Uber burning through its AI budget in four months covers the governance half.

REPTILEHAUS builds and runs usage-metered systems across AI agents, Web3 and conventional web platforms, and spend controls belong in the architecture rather than in an afterthought at invoice time. If you are unsure what your ceiling actually is, or what your application does when it hits one, get in touch.

📷 Photo by Mark Kats on Unsplash