Entry 010Comparisons · Guides8 min readBy Mandare Labs

Cloudflare AI Gateway spend limits: what they enforce

Cloudflare AI Gateway spend limits are dollar budgets over a rolling or fixed window, checked before each request goes to the provider; an over-budget request gets a 429. Per Cloudflare's documentation, read on 10 October 2026, enforcement is eventually consistent, so a burst of concurrent requests can briefly pass the limit, and a per-agent budget splits on metadata that the caller sends.

A spend limit is a dollar budget over a window, scoped by dimensions

Cloudflare added spend limits to AI Gateway on 5 June 2026, in a changelog post that calls them "cost-based budgets that track cumulative dollar spend and block requests when the budget is exceeded". The spend limits page, last updated 30 September 2026, gives the mechanics:

When cumulative spend reaches the limit within a time window, AI Gateway blocks further requests with a 429 response until the window resets.

A rule holds a budget in dollars over a rolling or fixed window. Rules are set per gateway in the dashboard (Settings, then Spend limits) or through the API, up to 20 per gateway. Before a request is forwarded, "AI Gateway evaluates all applicable spend limit rules at once. If any individual rule is over budget, the request is blocked with a 429 response."

What a rule counts is set by its dimensions: provider, model, or a custom metadata key. Each dimension is either split by value (every distinct value gets an independent bucket) or filtered to one value. The page's own examples, for a request with model openai/gpt-5.5 and an agent_id of agent_42:

ScenarioDimensionsBudget bucket
Global budgetNoneOne shared bucket
Per-agent budgetagent_id metadata: split by valueSeparate bucket per agent
Per-agent, per-modelagent_id split, model splitSeparate bucket per agent and model
One modelmodel filtered to openai/gpt-5.5Applies to that model's requests

A dimension left out of a rule means all values share one bucket.

A 429 from a spend limit looks like a 429 from a rate limit

The rate limiting page describes a different control with the same status code: "the server will respond with a 429 Too Many Requests status code and your request will not be processed." Rate limits count requests per fixed or sliding window. Spend limits count dollars. Neither page shows a response body that tells the two apart, so a client cannot be sure from the status alone which one it hit.

The retry logic matters, because the two clear on different clocks. A rate limit clears when its request window slides or ends. A spend limit clears when the budget window resets, which is a day if that is the window set. The daily spending limit guide covers how the providers' own caps answer, and the same rule holds here: an agent that retries every 429 without a ceiling keeps pushing against a wall.

// A 429 may be a rate limit (clears with its window) or a spend limit (clears at budget reset).
// Retry a bounded number of times, then stop and report to the operator.
const MAX_ATTEMPTS = 3;
for (let attempt = 1; attempt <= MAX_ATTEMPTS; attempt++) {
  const res = await fetch(url, init);
  if (res.status !== 429) return res;
  await sleep(2 ** attempt * 1000);
}
throw new Error('429 after retries: stop, do not loop');

When a limit is reached, the page offers two behaviours: block the request (the default), or fall back to a cheaper model through a Dynamic Route with a primary model and a fallback. With a fallback the agent does not see a 429 at all, so the budget stops being a stop signal and becomes a routing rule. That suits cost control and does not suit a hard cap.

Spend limits are eventually consistent, so a burst can pass the limit

The page states the timing plainly:

Spend limits are eventually consistent. The current request's cost is recorded after completion, so a burst of concurrent requests can briefly exceed the limit before enforcement catches up.

Read with the sentence on evaluation, the sequence is: each request is checked against recorded spend before it is forwarded, and its own cost is recorded after it completes. Thirty requests started together against a $20 budget, each costing about $1, would each be checked against a recorded spend of $0 and could all be admitted. That example is arithmetic on the two documented sentences, not a measurement. The page gives no bound for the overshoot.

The same page carries a second caveat under Limitations: "Cost tracking is a best-effort estimation based on token counts and model pricing. Refer to your provider's dashboard for exact billing amounts." The number being limited is an estimate, and the limit applies to requests "for models with known pricing". What happens to a request for a model with no known price is not on the page.

None of this makes the feature wrong for its job. A budget that stops a runaway agent after a burst is still useful, and "briefly" is the page's word. It does mean the limit is a ceiling to size with headroom, not an exact figure to reconcile against an invoice.

A per-agent budget splits on a value the request carries

The per-agent row in the table depends on agent_id arriving with each request. The custom metadata page shows how: a JSON header, cf-aig-metadata, on the call. This is the page's own cURL example:

curl -X POST "https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/v1/chat/completions" \
  --header "Authorization: Bearer $CLOUDFLARE_API_TOKEN" \
  --header "Content-Type: application/json" \
  --header 'cf-aig-metadata: {"team": "AI", "user": 12345, "test":true}' \
  --data '{"model": "openai/gpt-4.1", "messages": [{"role": "user", "content": "What should I eat for lunch?"}]}'

For a per-agent budget the header would carry the agent_id key from the spend limits examples. Three facts from that page shape what the budget means:

Together these say where the trust sits. For people behind Access, cf.user_id is set by Cloudflare and a per-user budget rests on a verified value. For an agent authenticating with a service token, there is no cf.user_id, so the per-agent budget rests on a value the agent's own process sends. A caller that sends a value no one has used before starts a fresh bucket by the page's definition. The pages read do not describe Cloudflare checking custom metadata values, and do not say whether a request that lacks the key is limited at all.

That is a fair trade for a budget meant to catch mistakes. It is a weaker guarantee for a budget meant to hold against an agent that has been steered by a prompt injection. A key the agent cannot choose is what makes the second case work, and the stolen-credential entry covers proof of possession, where a request is signed under a key the agent holds.

Where Cloudflare is the better choice

On several points a hosted AI gateway is the right answer, and the Mandare door is not.

For a human-facing app where a modest overshoot is acceptable and users already sit behind Access, spend limits are a sound fit. The comparison of 20 governance tools places Cloudflare and others side by side on four controls.

Where Mandare fits, and where it stops

Mandare is an open-source layer that sits in the request path and signs and records what happens. Its gateway, called the door, enforces a signed mandate. The two systems answer the same questions differently:

QuestionCloudflare spend limits (page read 10 October 2026)Mandare door (v0.1.0)
When is a call counted?Cost recorded after completion; eventually consistentEstimated cost reserved inside the ledger transaction before the call; the true cost settles after
What scopes the budget?Provider, model, or metadata the request carriesThe mandate names one agent key (did:key); the agent signs requests under it
Model with no price?Not stated on the pageRefused with MODEL_UNPRICED, never forwarded
What does a refusal look like?429403 with a code such as PER_DAY_EXCEEDED, recorded as a ledger entry

The first row rests on the concepts page: "an intent reserves its estimated cost inside the ledger transaction before the call; results settle the true cost." The threat model adds that "a 25-way race test through the gateway never lets settled spend pass the cap, and a 40-way driver-level test admits exactly the reservations that fit, on SQLite and Postgres". The second row rests on the same concepts page: agents "sign every request under their key (RFC 9421 HTTP Message Signatures, with a Content-Digest over the exact body bytes)". The reservation is a design choice with a cost: the door must know a price for every model it forwards, which is why an unpriced model is refused.

The no-docker demo, pnpm demo, releases a loop against a €20 per day mandate. This is the capture in docs/demos/S2-runaway-demo.txt, cut with …:

pnpm demolive capture
…
  call # 72  403 DENIED — THE LOOP DIES HERE
    code:   PER_DAY_EXCEEDED
    reason: per-day budget exceeded: reserved 0.00 EUR + settled 19.723587 EUR + estimate 0.278853 EUR > cap 20.00 EUR
    the refusal itself is ledger entry 59034994bfacad09…

[demo] runaway made 71 calls in 0.1s before the mandate killed it.

The limits, as of v0.1.0 (the version npm view @mandarelabs/cli version reported on 10 October 2026):

Questions

What does Cloudflare AI Gateway return when a spend limit is reached?

A 429 Too Many Requests response, per the spend limits page read on 10 October 2026. The default is to block until the budget window resets. The alternative is a Dynamic Route that sends the request to a cheaper fallback model. The page does not show a response body, and the rate-limit page also uses 429.

Can requests pass a Cloudflare AI Gateway spend limit?

Yes, briefly. The documentation says spend limits are eventually consistent: a request's cost is recorded after it completes, so a burst of concurrent requests can exceed the limit before enforcement catches up. It gives no bound for how far. Cost tracking is also described as a best-effort estimate from token counts and model pricing.

Can a spend limit be set per agent in Cloudflare AI Gateway?

Yes, by splitting a rule on a custom metadata key such as agent_id, which each request carries in the cf-aig-metadata header. Behind Cloudflare Access, the reserved cf.user_id key holds the verified user ID instead. The documentation does not describe Cloudflare verifying other metadata values, so the per-agent scope is as trustworthy as the sender.

Do spend limits work with your own provider API keys?

Per the page, spend limits apply to both Unified Billing requests and bring-your-own-keys requests, for models with known pricing. The page does not say what happens to a request for a model without known pricing. Up to 20 rules can be set per gateway.

Sources

  1. Spend limits (Cloudflare AI Gateway documentation) · Cloudflare · 2026-09-30
  2. Custom metadata (Cloudflare AI Gateway documentation) · Cloudflare · 2026-09-24
  3. Rate limiting (Cloudflare AI Gateway documentation) · Cloudflare · 2026-09-30
  4. Control AI costs with spend limits (Cloudflare changelog) · Cloudflare · 2026-06-05

Run it yourself: 3 commands, no API keys.

github.com/mandarelabs/mandare →
← All journal entries