AI agent approval fatigue: a signed limit beats prompts
Approval fatigue is what happens when a human is asked so often that the answer becomes automatic. Anthropic reports that Claude Code users approve 93% of permission prompts. The alternatives are a classifier that decides for you, fixed allow and deny rules, or authority bounded once with a threshold above which a human is asked, out of band, and silence means no.
A prompt that is almost always approved is not a control
Anthropic's engineering post on Claude Code's auto mode opens with a number: "Claude Code users approve 93% of permission prompts." The post names the effect that follows, approval fatigue, where people stop paying close attention to what they approve.
Two cautions on that figure. It describes users of one coding tool, so it says nothing about a platform team's agents with other defaults. And it is a rate of approval, not a rate of mistakes: a prompt answered yes 93 times in 100 can still be correct every time. The problem is the other seven. A reader who has clicked yes ninety times in a row is poorly placed to notice the ninety-first.
For an engineer running agents with real credentials, the design question is therefore not "how do we ask more often". It is which decisions must reach a human at all, and what happens to everything else.
Four ways to decide, and what each one gives up
| Approach | Who decides | What it gives up |
|---|---|---|
| Prompt on every action | The human, each time | Attention. The answer becomes a reflex. |
| Classifier ("auto mode") | A model, per action | A false-negative rate. Anthropic reports 17% on real overeager actions (n=52). |
| Fixed allow and deny rules | A rule written in advance | Coverage. Rules match patterns, and agents find variants. |
| Bounded authority plus a threshold | A signed limit, then a human above it | Judgement on what the action means. The limit sees amounts, not intent. |
Each row is a trade, not a ranking.
What the classifier trades
Anthropic is direct about it. The post puts the false-negative rate on real overeager actions at 17% (n=52), calls it "the honest number", and says auto mode is "not a drop-in replacement for careful human review on high-stakes infrastructure". A classifier reduces the number of prompts a person sees. It does not remove the case where the classifier is wrong.
What fixed rules trade
Claude Code's permission documentation shows the mechanism clearly. Rules are allow, ask or deny, and "Rules are evaluated in order: deny, then ask, then allow." A configuration that lets npm scripts and commits run but blocks pushes looks like this:
{
"permissions": {
"allow": ["Bash(npm run *)", "Bash(git commit *)"],
"deny": ["Bash(git push *)"]
}
}
The same page warns that "Bash permission patterns that try to constrain command arguments are fragile", and gives its own example: a push written as git -C . push is not matched by the deny rule above. Rules are the right tool when the dangerous action has a stable shape. They are weak when it can be spelled many ways.
Ask only the questions a signed limit cannot answer
A fourth route moves the decision earlier. The accountable human writes down, once, what the agent may do: an amount per call, an amount per day, a time window, and a threshold above which a person must be asked. Everything inside the limits runs without a prompt. Everything outside them is refused. The band between "clearly fine" and "clearly too much" reaches a human.
This works for money because money has a number on it. It works less well for actions with no natural scale, such as "delete this file" or "send this email". Say so in the design: a spend threshold answers "how much", not "should this happen". For those actions, allow and deny rules or a sandbox are the better tools, and they combine with a spend limit rather than compete with it.
A held request needs six properties
Once a human is asked above a threshold, the request has to wait somewhere. The pattern is older than agents. OpenID's CIBA specification (Final, 1 September 2021) describes an authentication flow in which, unlike OpenID Connect, there is "direct Relying Party to OpenID Provider communication without redirects through the user's browser". The user starts on one device and approves on another.
CIBA is about logins, not agent calls, but the shape carries over. An approval flow for an agent needs these properties:
- The default is no. If nobody answers within a fixed time, the held call is refused.
- The decision is single-use. An approve link that can be replayed authorises more than the call it was sent for.
- The decision is recorded before the call proceeds. If the record cannot be written, the call does not run.
- State is re-checked on resume. A budget may have moved, or an operator may have revoked the agent, while the call waited.
- Pending holds are capped. A looping agent must not be able to flood the human with pushes, and notification floods are themselves an attack on attention.
- No channel, no pass. If no approval channel is configured, a call above the threshold is refused instead of waved through.
Items 3 to 6 are the ones that matter when something goes wrong.
How Mandare implements the pattern, and where it stops
Mandare is an open-source accountability layer for agents: a signed identity per agent, a spending limit the operator signs once, and a record the infrastructure writes. The concepts page describes the mandate as "spend caps per transaction / day / task / total, validity window, allowed scopes, and an approval threshold above which a human must asynchronously approve", evaluated in a fixed order: identity, window, scope, budget, counterparty, approval.
The demo script issues one mandate with --per-tx 5 --per-day 20 --approval-above 0.25 (euro amounts), then runs eight steps against a mock provider. This is the captured output from docs/demos/S4-mandate-demo.txt, cut with …:
[owner] mandate mnd_955530bd9abe signed ONCE: €5/tx, €20/day, ask me above €0.25.
…
[agent] step 1/6 signed call (~€0.001) → 200 OK — no human involved
…
6 calls, 0 permission prompts. The mandate IS the answer.
[agent] step 7: big synthesis call (~€0.28, above the €0.25 threshold)…
[push] → "Mandare: approve ~0.2788 EUR?"
[human] taps APPROVE on the phone.
[agent] held call resumed → 200 OK (approval recorded on the ledger)
[agent] step 8: ANOTHER big call (agent got ambitious)…
[push] → "Mandare: approve ~0.2788 EUR?"
[human] taps DENY.
[agent] held call refused → 403 APPROVAL_DENIEDThe repository's example folder is titled "one mandate replaces 40 permission prompts". The capture above runs six calls, so the number 40 is an illustration of scale, not a measurement.
The six properties map onto the gateway code in the public repository, read at commit 327442e on 4 October 2026:
- Default no.
MANDARE_APPROVAL_TIMEOUT_MSdefaults to 120000, and the environment reference states "no decision ⇒ deny". The refusal code isAPPROVAL_TIMEOUT. - Single use. Each button carries a 256-bit random token, stored as a SHA-256 hash, dead after first use or timeout. Replay answers
ALREADY_DECIDED. - Recorded first. The decision is appended to the ledger as
approval.granted,approval.deniedor an expiry entry. If an approval cannot be recorded, the gateway returns 503 and does not act. - Re-checked. After an approval the gateway re-checks revocation, then re-runs the full policy order. The approval waives the threshold and nothing else.
- Capped.
MANDARE_MAX_PENDING_APPROVALSdefaults to 8. Past it, the call is refused withAPPROVAL_BACKLOGand written to the ledger as a denial; repeats of the same refusal past a burst are not written again. - No channel, no pass. With no notifier configured, a call above the threshold is refused with
APPROVAL_REQUIRED.
Then mandare verify shows the human decisions inside the same hash-linked trail as the calls (approval.requested, approval.granted, approval.denied). The RFC 9421 entry covers how the agent's identity on each call is checked, and a runaway loop dying at €20 shows the budget side of the same mandate.
Limits of the current release
These apply to v0.1.0, the version npm view @mandarelabs/cli version reports on 4 October 2026.
- The threshold compares against the estimated cost of an LLM call at the door. The door proxies Anthropic
POST /v1/messagesand OpenAI or OpenRouterPOST /v1/chat/completions. It does not see an agent's shell commands or file edits, so Claude Code's own permission rules remain the better tool for those. - Mandare is, in its own words, not a content filter and not a sandbox. It does not judge what a call is for.
- The demo uses a mock provider, and its push channel is a file, so the "tap" is scripted. A phone push needs
MANDARE_NOTIFIER=ntfy. On the public ntfy.sh the topic name is the sole secret, and the threat model advises self-hosting the channel. - The approve button is a capability token, not a login. The ledger attributes the decision to the mandate's principal; it proves a decision arrived through that channel, not who held the phone.
- The release has had two AI-assisted review passes and no external audit.
Spend caps and approvals answer different questions. The cap says how much, whoever asks. The approval says yes or no to one call, above a line the owner drew in advance. The mechanism is described on the demos page.
Questions
What is approval fatigue in AI agents?
It is the drift from reading each permission prompt to clicking approve without reading. Anthropic's engineering post of 25 March 2026 names it and reports that Claude Code users approve 93% of permission prompts. The figure describes one product's users, not agents in general.
Should an agent ask a human before every action?
No. A prompt that is nearly always approved adds delay and hides the rare call that deserves attention. Decide in advance what is always fine and what is never fine, then ask a human only for the narrow band in between, with a clear default when nobody answers.
How can a human approve an agent action asynchronously?
The held request waits while a notification goes to a second device, where the human taps approve or deny. OpenID's CIBA specification describes this shape for logins: an authentication device separate from the consumption device, with no browser redirects. The design needs a timeout, a record of the decision and a limit on pending requests.
What should happen when nobody answers an approval request?
The request should be refused, not allowed. In Mandare v0.1.0 the default window is 120000 ms (MANDARE_APPROVAL_TIMEOUT_MS) and no decision means deny. Eight pending approvals is the default ceiling (MANDARE_MAX_PENDING_APPROVALS); past it the call is refused and recorded.
Sources
- A look at Claude Code's auto mode · Anthropic · 2026-03-25
- Configure permissions (Claude Code) · Anthropic
- OpenID Connect Client-Initiated Backchannel Authentication (CIBA) Flow - Core 1.0 · OpenID Foundation · 2021-09-01
Run it yourself: 3 commands, no API keys.
github.com/mandarelabs/mandare →