Entry 005Guides7 min readBy Mandare Labs

How to set a daily spending limit for an AI agent

The Anthropic and OpenAI consoles document spend limits as monthly, so a daily cap has to come from a layer in front of them. OpenRouter keys take a daily reset, and LiteLLM budgets take a duration such as days. Each trips with a different status code. An agent that retries blindly keeps hitting the wall, so the error handling matters as much as the number.

The provider consoles cap the month, not the day

Anthropic's rate limits page defines two kinds of limit: "Spend limits set a maximum monthly cost an organization can incur for API usage" and rate limits on requests. Each usage tier carries a monthly spend cap, and you can set your own limit below it. The page lists these tier caps:

Usage tierMonthly spend cap
Start$500 USD
Build$1,000 USD
Scale$200,000 USD

OpenAI's rate limits guide has a matching section. It says to consider "spend limits for your organization or projects to control monthly API spend", and its table separates two controls. A spend alert "Sends a notification; API traffic continues". A hard spend limit means "Affected API requests return a 429 error", and the stated use is to "Enforce a monthly organization or project cap".

Neither page describes a daily dollar window. OpenAI does list RPD (requests per day) and TPD (tokens per day) among its rate-limit metrics, but those count requests and tokens. They do not convert to a currency cap, because the price per token differs by model.

For an agent, the month is a long time. At the Start tier a loop that spends $500 in one night exhausts the whole month, and the provider cap is then doing its job as a backstop while the damage is already done. A daily limit has to sit in a layer you control.

Where a daily window does exist

Two widely used gateways document a reset period on a budget. The reset is the part that turns a budget into a daily limit.

OpenRouter keys take a credit limit with a reset type. The create-key reference describes limit_reset as "Type of limit reset for the API key (daily, weekly, monthly, or null for no reset). Resets happen automatically at midnight UTC". Its example body, with the reset changed to daily, looks like this:

{
  "name": "research-agent-1",
  "limit": 20,
  "limit_reset": "daily"
}

That body is sent to POST https://openrouter.ai/api/v1/keys. The limit is in USD. The limits page says GET /api/v1/key returns limit_remaining and usage_daily (credits used in the current UTC day), so an agent or a cron job can read how much of the day is left before a call fails. The page's own advice is to "Monitor proactively".

LiteLLM's proxy takes a max_budget with a budget_duration on keys, users and teams. The proxy documentation says "Budget is reset at the end of specified duration. If not set, budget is never reset" and lists the units: seconds ("30s"), minutes ("30m"), hours ("30h") and days ("30d"). A daily key is therefore the same call with a one-day duration:

curl 'http://0.0.0.0:4000/key/generate' \
  --header 'Authorization: Bearer <your-master-key>' \
  --header 'Content-Type: application/json' \
  --data-raw '{ "max_budget": 20, "budget_duration": "1d" }'

The docs print "30d" in their own example; "1d" follows the documented unit format. One detail affects a daily cap: the same page notes "By default, the server checks for resets every 10 minutes, to minimize DB calls." A key that has hit its budget may therefore stay shut for some minutes after its window ends, and the window is a duration from the last reset, not a calendar day as with OpenRouter's midnight UTC. The interval is configurable (proxy_budget_rescheduler_min_time and proxy_budget_rescheduler_max_time).

Both are the better choice when you route between many providers or already run a proxy for model routing: a daily budget is one field on a key, and nothing else has to be added.

Each cap says no in its own way

The status code is the first thing the agent sees, and it is not consistent across tools.

SourceWhat tripsResponse
Anthropic tier capMonthly tier spend capHTTP 429 rate_limit_error, error_code enforced_spend_limit_reached, no retry-after
Anthropic limit you setYour own spend limitHTTP 400 invalid_request_error, message begins "You have reached your specified API usage limits"
OpenAI hard spend limitOrg or project capA 429 error (the page does not show the body)
OpenRouter key limitPer-key credit limitHTTP 402, limit_source openrouter_key_limit
LiteLLM key budgetmax_budget crossedAuthentication Error, ExceededTokenBudget: Current spend for token: …

The Anthropic page spells out the consequence for retry logic: the spend-cap 429 "has no retry-after header. Retrying, including the SDK's automatic retries, fails until access resumes." Access resumes at 00:00 UTC on the first day of the next month. OpenAI uses the same 429 code for temporary rate limits, where its guide recommends retrying with exponential backoff, so code that treats every 429 as transient will spin against a spend limit.

A small guard avoids that. This sketch reads the fields named in the two vendors' documentation and stops instead of retrying:

async function isBudgetStop(res: Response): Promise<boolean> {
  const body = await res.clone().json().catch(() => null);
  const anthropic = body?.error?.details?.error_code === 'enforced_spend_limit_reached';
  const openrouter = res.status === 402 && body?.error?.metadata?.limit_source === 'openrouter_key_limit';
  return anthropic || openrouter;
}

OpenRouter marks a different 402 as transient: openrouter_in_flight_budget carries a Retry-After header and the documentation says to wait and retry. Branch on limit_source, as the page advises ("branch on limit_source, not on the hint text"), not on the status code alone.

Count the cost before the call, or accept overshoot

A budget enforced from billing data trails the spend. Whether it can be crossed depends on when a call is counted.

OpenRouter's limits page describes the problem for account balance. It "charges a request when it finishes, so many requests running at the same time could commit more than your balance covers before any of them settles". Its fix is an in-flight spending budget: it estimates each paid request's cost up front from input tokens plus the completion tokens allowed by max_tokens, and holds that estimate while the request runs. The page describes this hold against the account balance. It does not say whether a per-key credit limit is checked the same way, so a daily key limit should not be assumed to be exact under a burst of parallel calls.

Two things follow for an agent budget:

Give each agent its own key and keep the upstream key out of reach

A daily cap is only as strict as its scope. Anthropic's workspaces let you "set custom spend and rate limits per Workspace", with the note "You can't set limits on the default Workspace". OpenAI projects take their own spend limit. OpenRouter and LiteLLM budgets attach to the key.

The rule is one key per agent, or per task class, so that one runaway agent exhausts its own budget and not the fleet's. A budget that lives in a proxy has a second condition: the agent must not hold the provider's real key. If it can read that key from the environment, it can call the provider directly and the proxy's budget never sees the call. Treat the upstream key as a credential for the gateway, not for the agent.

Where Mandare fits

Mandare, an open-source accountability layer for agents, enforces a daily cap at its door, a gateway that sits in the request path. It does not replace a provider cap. It is one more layer, with the monthly provider limit still behind it.

The owner signs a mandate once. The quickstart shows the command, with --per-day 20 as the daily cap in euros:

./bin/mandare mandate issue --agent <did:key from above> \
  --per-tx 5 --per-day 20 --total 100 --approval-above 5 \
  --out mandate.sdjwt

The door reserves the estimated cost of each call inside the ledger transaction before forwarding it, and the result settles the true cost. The day is a UTC calendar day: dayBucket in packages/ledger/src/projection.ts slices the first ten characters of a UTC timestamp. When the next call would cross the cap, the door refuses it before it reaches a provider. This is the capture in docs/demos/S2-runaway-demo.txt, cut with …:

pnpm demolive capture
…
  call # 72  403 DENIED — THE LOOP DIES HERE
    code:   PER_DAY_EXCEEDED
    reason: per-day budget exceeded: reserved 0.00 EUR + settled 19.723587 EUR + estimate 0.278853 EUR > cap 20.00 EUR
    the refusal itself is ledger entry 59034994bfacad09…

[demo] runaway made 71 calls in 0.1s before the mandate killed it.

The refusal is itself an entry in the hash-linked ledger, which mandare verify checks. The runaway-loop walkthrough covers that record, and the approval-fatigue entry covers the --approval-above threshold.

The limits, as of v0.1.0 (the version npm view @mandarelabs/cli version reports on 5 October 2026):

Questions

Is there a daily spending limit on the Anthropic or OpenAI API?

Not from the pages read on 5 October 2026. Anthropic describes spend limits as a maximum monthly cost, and OpenAI's hard spend limit is documented as a monthly organization or project cap. OpenAI lists requests-per-day and tokens-per-day rate limits, but those count requests and tokens, not currency. A daily window needs a gateway or key-level budget in front of the provider.

What happens to an agent when it hits a spend limit?

It gets an error, and the status code depends on the provider: HTTP 429 for Anthropic's tier cap and for OpenAI's hard spend limit, HTTP 400 for a limit you set yourself on Anthropic, 402 for an OpenRouter key limit. Anthropic states that retrying, including the SDK's automatic retries, fails until access resumes, so the agent has to treat these as stop signals.

Is a per-key budget the same as a per-agent budget?

Only if each agent has its own key and the agent cannot reach the upstream provider key. A shared key pools every agent's spend under one cap, and an agent that can read the real provider key can call the provider directly and skip any budget that lives in a proxy.

Does a budget stop the call that crosses it?

It depends on whether the cost is reserved before the call or counted after. OpenRouter documents that it charges a request when it finishes and holds an estimate for the account balance in the meantime. Its page does not say the same hold applies to a per-key credit limit. Check how a given tool counts in-flight calls before relying on the cap as exact.

Sources

  1. Rate limits (Claude API documentation) · Anthropic
  2. Rate limits (OpenAI API documentation) · OpenAI
  3. API Credit & Rate Limits: Handle 402 and 429 Errors · OpenRouter · 2026-09-17
  4. Create API key (OpenRouter API reference) · OpenRouter
  5. Virtual keys, users and teams: budgets and rate limits (LiteLLM proxy) · LiteLLM

Run it yourself: 3 commands, no API keys.

github.com/mandarelabs/mandare →
← All journal entries