How to set a daily spending limit for an AI agent
The Anthropic and OpenAI consoles document spend limits as monthly, so a daily cap has to come from a layer in front of them. OpenRouter keys take a daily reset, and LiteLLM budgets take a duration such as days. Each trips with a different status code. An agent that retries blindly keeps hitting the wall, so the error handling matters as much as the number.
The provider consoles cap the month, not the day
Anthropic's rate limits page defines two kinds of limit: "Spend limits set a maximum monthly cost an organization can incur for API usage" and rate limits on requests. Each usage tier carries a monthly spend cap, and you can set your own limit below it. The page lists these tier caps:
| Usage tier | Monthly spend cap |
|---|---|
| Start | $500 USD |
| Build | $1,000 USD |
| Scale | $200,000 USD |
OpenAI's rate limits guide has a matching section. It says to consider "spend limits for your organization or projects to control monthly API spend", and its table separates two controls. A spend alert "Sends a notification; API traffic continues". A hard spend limit means "Affected API requests return a 429 error", and the stated use is to "Enforce a monthly organization or project cap".
Neither page describes a daily dollar window. OpenAI does list RPD (requests per day) and TPD (tokens per day) among its rate-limit metrics, but those count requests and tokens. They do not convert to a currency cap, because the price per token differs by model.
For an agent, the month is a long time. At the Start tier a loop that spends $500 in one night exhausts the whole month, and the provider cap is then doing its job as a backstop while the damage is already done. A daily limit has to sit in a layer you control.
Where a daily window does exist
Two widely used gateways document a reset period on a budget. The reset is the part that turns a budget into a daily limit.
OpenRouter keys take a credit limit with a reset type. The create-key reference describes limit_reset as "Type of limit reset for the API key (daily, weekly, monthly, or null for no reset). Resets happen automatically at midnight UTC". Its example body, with the reset changed to daily, looks like this:
{
"name": "research-agent-1",
"limit": 20,
"limit_reset": "daily"
}
That body is sent to POST https://openrouter.ai/api/v1/keys. The limit is in USD. The limits page says GET /api/v1/key returns limit_remaining and usage_daily (credits used in the current UTC day), so an agent or a cron job can read how much of the day is left before a call fails. The page's own advice is to "Monitor proactively".
LiteLLM's proxy takes a max_budget with a budget_duration on keys, users and teams. The proxy documentation says "Budget is reset at the end of specified duration. If not set, budget is never reset" and lists the units: seconds ("30s"), minutes ("30m"), hours ("30h") and days ("30d"). A daily key is therefore the same call with a one-day duration:
curl 'http://0.0.0.0:4000/key/generate' \
--header 'Authorization: Bearer <your-master-key>' \
--header 'Content-Type: application/json' \
--data-raw '{ "max_budget": 20, "budget_duration": "1d" }'
The docs print "30d" in their own example; "1d" follows the documented unit format. One detail affects a daily cap: the same page notes "By default, the server checks for resets every 10 minutes, to minimize DB calls." A key that has hit its budget may therefore stay shut for some minutes after its window ends, and the window is a duration from the last reset, not a calendar day as with OpenRouter's midnight UTC. The interval is configurable (proxy_budget_rescheduler_min_time and proxy_budget_rescheduler_max_time).
Both are the better choice when you route between many providers or already run a proxy for model routing: a daily budget is one field on a key, and nothing else has to be added.
Each cap says no in its own way
The status code is the first thing the agent sees, and it is not consistent across tools.
| Source | What trips | Response |
|---|---|---|
| Anthropic tier cap | Monthly tier spend cap | HTTP 429 rate_limit_error, error_code enforced_spend_limit_reached, no retry-after |
| Anthropic limit you set | Your own spend limit | HTTP 400 invalid_request_error, message begins "You have reached your specified API usage limits" |
| OpenAI hard spend limit | Org or project cap | A 429 error (the page does not show the body) |
| OpenRouter key limit | Per-key credit limit | HTTP 402, limit_source openrouter_key_limit |
| LiteLLM key budget | max_budget crossed | Authentication Error, ExceededTokenBudget: Current spend for token: … |
The Anthropic page spells out the consequence for retry logic: the spend-cap 429 "has no retry-after header. Retrying, including the SDK's automatic retries, fails until access resumes." Access resumes at 00:00 UTC on the first day of the next month. OpenAI uses the same 429 code for temporary rate limits, where its guide recommends retrying with exponential backoff, so code that treats every 429 as transient will spin against a spend limit.
A small guard avoids that. This sketch reads the fields named in the two vendors' documentation and stops instead of retrying:
async function isBudgetStop(res: Response): Promise<boolean> {
const body = await res.clone().json().catch(() => null);
const anthropic = body?.error?.details?.error_code === 'enforced_spend_limit_reached';
const openrouter = res.status === 402 && body?.error?.metadata?.limit_source === 'openrouter_key_limit';
return anthropic || openrouter;
}
OpenRouter marks a different 402 as transient: openrouter_in_flight_budget carries a Retry-After header and the documentation says to wait and retry. Branch on limit_source, as the page advises ("branch on limit_source, not on the hint text"), not on the status code alone.
Count the cost before the call, or accept overshoot
A budget enforced from billing data trails the spend. Whether it can be crossed depends on when a call is counted.
OpenRouter's limits page describes the problem for account balance. It "charges a request when it finishes, so many requests running at the same time could commit more than your balance covers before any of them settles". Its fix is an in-flight spending budget: it estimates each paid request's cost up front from input tokens plus the completion tokens allowed by max_tokens, and holds that estimate while the request runs. The page describes this hold against the account balance. It does not say whether a per-key credit limit is checked the same way, so a daily key limit should not be assumed to be exact under a burst of parallel calls.
Two things follow for an agent budget:
- Set
max_tokenson every call. Estimates are only as tight as the ceiling they are given, and OpenRouter notes it uses a fixed per-request cap whenmax_tokensis not set. - Size the daily cap with the burst in mind. If the agent can run thirty calls in parallel and each can cost a dollar, a $20 cap that is counted after settlement can be passed by a margin of that size.
Give each agent its own key and keep the upstream key out of reach
A daily cap is only as strict as its scope. Anthropic's workspaces let you "set custom spend and rate limits per Workspace", with the note "You can't set limits on the default Workspace". OpenAI projects take their own spend limit. OpenRouter and LiteLLM budgets attach to the key.
The rule is one key per agent, or per task class, so that one runaway agent exhausts its own budget and not the fleet's. A budget that lives in a proxy has a second condition: the agent must not hold the provider's real key. If it can read that key from the environment, it can call the provider directly and the proxy's budget never sees the call. Treat the upstream key as a credential for the gateway, not for the agent.
Where Mandare fits
Mandare, an open-source accountability layer for agents, enforces a daily cap at its door, a gateway that sits in the request path. It does not replace a provider cap. It is one more layer, with the monthly provider limit still behind it.
The owner signs a mandate once. The quickstart shows the command, with --per-day 20 as the daily cap in euros:
./bin/mandare mandate issue --agent <did:key from above> \
--per-tx 5 --per-day 20 --total 100 --approval-above 5 \
--out mandate.sdjwt
The door reserves the estimated cost of each call inside the ledger transaction before forwarding it, and the result settles the true cost. The day is a UTC calendar day: dayBucket in packages/ledger/src/projection.ts slices the first ten characters of a UTC timestamp. When the next call would cross the cap, the door refuses it before it reaches a provider. This is the capture in docs/demos/S2-runaway-demo.txt, cut with …:
…
call # 72 403 DENIED — THE LOOP DIES HERE
code: PER_DAY_EXCEEDED
reason: per-day budget exceeded: reserved 0.00 EUR + settled 19.723587 EUR + estimate 0.278853 EUR > cap 20.00 EUR
the refusal itself is ledger entry 59034994bfacad09…
[demo] runaway made 71 calls in 0.1s before the mandate killed it.The refusal is itself an entry in the hash-linked ledger, which mandare verify checks. The runaway-loop walkthrough covers that record, and the approval-fatigue entry covers the --approval-above threshold.
The limits, as of v0.1.0 (the version npm view @mandarelabs/cli version reports on 5 October 2026):
- The cap holds given a correct price table. A model with no price entry is refused with
MODEL_UNPRICEDand is not forwarded; the price table is extended withMANDARE_PRICING_PATH. We make no claim that a cap cannot be exceeded under any condition. - The door proxies Anthropic
POST /v1/messagesand OpenAI or OpenRouterPOST /v1/chat/completions. It does not route across providers the way LiteLLM or OpenRouter do. - The demo numbers use a mock provider: the enforcement is real, the money is not. The no-docker demo stops at call #72 with €19.72 settled, and the docker quickstart stops at call #24 with €19.17.
- Provider keys in
.envbehind an unauthenticated door are, in the quickstart's words, "a starting point, not a boundary": a process that can read.envcan call the provider and bypass the mandate. The keys belong in the vault, with token or passport auth on the door; see Self-hosting and the threat model. - There is no hosted service, and the release has had two AI-assisted review passes and no external audit.
Questions
Is there a daily spending limit on the Anthropic or OpenAI API?
Not from the pages read on 5 October 2026. Anthropic describes spend limits as a maximum monthly cost, and OpenAI's hard spend limit is documented as a monthly organization or project cap. OpenAI lists requests-per-day and tokens-per-day rate limits, but those count requests and tokens, not currency. A daily window needs a gateway or key-level budget in front of the provider.
What happens to an agent when it hits a spend limit?
It gets an error, and the status code depends on the provider: HTTP 429 for Anthropic's tier cap and for OpenAI's hard spend limit, HTTP 400 for a limit you set yourself on Anthropic, 402 for an OpenRouter key limit. Anthropic states that retrying, including the SDK's automatic retries, fails until access resumes, so the agent has to treat these as stop signals.
Is a per-key budget the same as a per-agent budget?
Only if each agent has its own key and the agent cannot reach the upstream provider key. A shared key pools every agent's spend under one cap, and an agent that can read the real provider key can call the provider directly and skip any budget that lives in a proxy.
Does a budget stop the call that crosses it?
It depends on whether the cost is reserved before the call or counted after. OpenRouter documents that it charges a request when it finishes and holds an estimate for the account balance in the meantime. Its page does not say the same hold applies to a per-key credit limit. Check how a given tool counts in-flight calls before relying on the cap as exact.
Sources
- Rate limits (Claude API documentation) · Anthropic
- Rate limits (OpenAI API documentation) · OpenAI
- API Credit & Rate Limits: Handle 402 and 429 Errors · OpenRouter · 2026-09-17
- Create API key (OpenRouter API reference) · OpenRouter
- Virtual keys, users and teams: budgets and rate limits (LiteLLM proxy) · LiteLLM
Run it yourself: 3 commands, no API keys.
github.com/mandarelabs/mandare →