Stolen AI agent API key: what proof of possession fixes
A leaked LLM API key is a bearer credential: whoever holds it can use it until someone revokes it. Three measures narrow that. Short lifetimes bound the window, sender-constraining (proof of possession) makes a copied token id useless without a secret, and a local kill switch ends the rest. None of them protects a request body or a key stolen whole.
A bearer key is worth what its holder can do with it
RFC 6750 defines the property that makes a leaked credential dangerous. A bearer token is "a security token with the property that any party in possession of the token (a "bearer") can use the token in any way that any other party in possession of it can." The next sentence names what is missing: "Using a bearer token does not require a bearer to prove possession of cryptographic key material (proof-of-possession)."
An agent's provider key usually works this way. It sits in an environment variable, a config file, or the process memory of something that also reads untrusted text. A copy out of a log line, a crash dump, a prompt-injected tool call or a shared .env is a complete credential. Nothing in the request tells the provider that the sender changed.
What the thief gets is bounded by three things: what the key is allowed to do, how long it lives, and how fast someone notices and revokes it. The rest of this entry is about moving each of those.
Short lifetimes bound the window, not the damage inside it
RFC 6750 section 5.3 recommends: "Token servers SHOULD issue short-lived (one hour or less) bearer tokens, particularly when issuing tokens to clients that run within a web browser or other environments where information leakage may occur." Its reason: "Using short-lived bearer tokens can reduce the impact of them being leaked."
For an agent that means the provider key itself should not be what the agent carries. A static provider key has no expiry at all. A token minted per task, with a lifetime in minutes, turns a leak from "until someone rotates the key" into "until the clock runs out".
A short lifetime does nothing inside the window. A thief who copies a 15-minute token has 15 minutes of full use, and an agent loop can spend a lot in 15 minutes. Expiry limits how long, not how much. How much is the job of a spending limit, covered in setting a daily spending limit for an agent.
Sender-constraining makes a copied token id useless on its own
The standard answer to "a copy is a complete credential" is to make the copy incomplete. RFC 9700, the OAuth 2.0 security best current practice (January 2025), states it in section 2.2.1: "A sender-constrained access token scopes the applicability of an access token to a certain sender. This sender is obliged to demonstrate knowledge of a certain secret as a prerequisite for the acceptance of that token at the recipient." It adds that servers "SHOULD use mechanisms for sender-constraining access tokens, such as mutual TLS for OAuth 2.0 [RFC8705] or OAuth 2.0 Demonstrating Proof of Possession (DPoP) [RFC9449]."
RFC 9449 specifies DPoP. Its abstract calls it "a mechanism for sender-constraining OAuth 2.0 tokens via a proof-of-possession mechanism on the application level", and the introduction says what that buys: "the legitimate presenter of the token is constrained to be the sender that holds and proves possession of the private part of the key pair."
The client signs a small JWT with each request. Section 4.2 requires these claims in it:
| Claim | RFC 9449 wording | Why it is there |
|---|---|---|
htm | "The value of the HTTP method" | Ties the proof to the verb |
htu | "The HTTP target URI ... without query and fragment parts" | Ties the proof to the endpoint |
jti | "Unique identifier for the DPoP proof JWT" | Lets the server detect reuse |
iat | "Creation timestamp of the JWT" | Lets the server reject stale proofs |
Two details matter for anyone designing a similar scheme. First, the signature algorithm must be asymmetric: section 4.2 says alg "MUST NOT be none or an identifier for a symmetric algorithm (Message Authentication Code (MAC))." The server never holds anything that can forge a proof. Second, replay detection is the server's job. Section 11.1 says servers "can store the jti value of each DPoP proof for the time window in which the respective DPoP proof JWT would be accepted to prevent multiple uses of the same DPoP proof." The same reasoning for HTTP signatures, including how long to keep a nonce, is in the RFC 9421 entry.
What a stolen token id, a replay and a full copy each get
Put the three measures against the three ways a credential leaks. The right-hand columns follow from the RFC text above and the demo further down, not from a measurement.
| What the thief has | Bearer key | Short lifetime | Proof of possession |
|---|---|---|---|
| The token or key string only (log line, prompt) | Full use | Full use until expiry | Refused: no proof |
| One captured request | Replay at will | Replay until expiry | Refused if the server remembers the nonce or jti |
| Token and signing secret, copied whole | Full use | Full use until expiry | Full use until expiry or revocation |
| A man in the middle, who can alter the body | Full use | Full use until expiry | Body can be changed under a valid proof |
The last row is a deliberate gap in the standard. RFC 9449 section 11.7: "DPoP does not ensure the integrity of the payload or headers of requests. The DPoP proof only contains claims for the HTTP URI and method, but not the message body or general request headers." For an agent the body holds the model name, the token cap and the prompt, so a scheme for agents has to add a body digest.
The third row is the one that proof of possession cannot close. Section 11.4 of the RFC says that if the private key "is stored in such a way that it cannot be exported ... the adversary cannot exfiltrate the key and use it to create arbitrary DPoP proofs", but it also notes the adversary can still create proofs as long as the client is online and uses the key. A thief who runs code on the agent's host is the agent, as far as the verifier can tell. Expiry and revocation are what help there.
Revocation has to be local, immediate and recorded
The fourth measure is the one that works on the third row. A kill switch has three requirements that a plain "rotate the key" does not meet:
- It must not need the network. A revocation that has to reach a cloud service can be blocked by the same incident that made it necessary.
- It must be checked before authentication on every request, so a revoked agent cannot use a still-valid token in the gap between the kill and the next token refresh.
- The refusal should leave a record. "The agent was refused at 14:03 under revocation entry X" is evidence; a silent 403 is not.
A revocation list also has a standard shape if more than one verifier needs it. Mandare renders its revocation state as an IETF Token Status List bitstring, as its concepts page describes. This entry does not rely on that draft for anything else.
Where Mandare fits, and what it does not do
Mandare's gateway (the "door") has a token mode built on this pattern. As of v0.1.0, which npm view @mandarelabs/cli version reported on 6 October 2026, the vault mints a scoped token for an agent: a public id plus a per-token secret, with a default lifetime of 15 minutes and a ceiling of 30. The CLI prints the secret once (mandare token issue --actor <did> --mandate <id>, see the CLI reference). Each request carries these headers, from the repository's red-team test:
x-mandare-token <token id>
x-mandare-timestamp <ISO-8601 UTC>
x-mandare-nonce <single-use value>
x-mandare-pop <HMAC over the preimage below>
The preimage is five fields joined by newlines, in this order, as defined in packages/vault/src/tokens.ts: token id, upper-cased method, path, timestamp, nonce. The timestamp must be within 120 seconds of the vault's clock, and a nonce is kept for 245 seconds so it outlives every moment a captured proof could still pass the freshness check. The nonce is claimed only after the proof verifies, so a forged probe cannot use up a real client's nonce. The repository's stolen-token demo shows the outcome (docs/demos/S3-dead-paper-demo.txt, cuts marked):
[agent] legitimate signed call → 200 OK
[thief] stolen token id, forged proof → 401 BAD_POP (binding: no pop secret)
[thief] replay of a captured request → 401 REPLAYED_NONCE (single-use nonce)
$ mandare kill did:mandare:dev-agent
KILLED did:mandare:dev-agent
…
vault: 1 live token(s) revoked
…
status: written to the LOCAL ledger — the gateway fails closed on its next request (offline, un-jammable)
[agent] valid signed call AFTER the kill → 403 AGENT_REVOKED
the refusal is ledger entry fc84d99646ba90a5…The provider key sits in the vault and the door uses it. The agent never holds it. The demos page lists the same run. Limits, each next to the claim it limits:
- Token mode is not DPoP. The proof is an HMAC under a per-token secret, a symmetric scheme, which RFC 9449 section 4.2 forbids for DPoP proofs. The vault has to be able to recompute the HMAC, so it stores each secret sealed under its master key (AES-256-GCM, master key in the OS keychain by default). Someone who obtains the vault file and the master key can produce proofs for live tokens. We make no claim of compatibility with DPoP or any OAuth server.
- The body is not bound. The repository's token module says the proof covers method, path, timestamp and nonce and not the payload, and the threat model lists it as a residual. Passport mode, with an Ed25519 agent key and an RFC 9421 signature over a Content-Digest, closes that gap and is the mode for anything off loopback. Python's client supports token mode only.
- A whole credential is good until it dies. Token id plus secret, copied together, works until the 30-minute ceiling or
mandare kill. The kill is local and offline, and the door checks revocation on every request before authentication. - The key moves to the door. The provider key is still a bearer credential between the door and the provider. It is now in the vault on the door's host and not in the agent's environment. A runaway spend there is a separate problem, shown in a runaway loop that dies at €20.
- What the demo is. It runs against a mock provider with no API keys: the refusals are real, the money is not. The red-team tests were written by the team that wrote the code, and no external audit has happened yet.
Questions
What can someone do with a stolen AI agent API key?
Anything the key allows, until the key expires or is revoked. RFC 6750 defines a bearer token as one that any party in possession can use "in any way that any other party in possession of it can", so the provider cannot tell the thief from the agent.
What is a proof-of-possession token?
A token that is useless alone: the presenter must also prove it holds a secret or private key, usually by signing each request. RFC 9700 calls this a sender-constrained access token and says servers SHOULD use such mechanisms, for example mutual TLS or DPoP.
Does proof of possession stop a stolen private key?
No. If the key leaves the host, the thief can sign valid requests until the credential expires or is revoked. RFC 9449 section 11.4 says an adversary running code in the client's context ends the protection: even with a non-exportable key, the adversary can still create proofs while the client is online. Expiry and revocation cover the rest.
Does DPoP protect the request body?
No. RFC 9449 section 11.7 says DPoP "does not ensure the integrity of the payload or headers of requests". The proof carries the HTTP method and URI, not the body. A scheme that must bind the body has to sign a digest of it as well.
Sources
- RFC 6750: The OAuth 2.0 Authorization Framework: Bearer Token Usage · IETF / RFC Editor · 2012-10-01
- RFC 9449: OAuth 2.0 Demonstrating Proof of Possession (DPoP) · IETF / RFC Editor · 2023-09-01
- RFC 9700: Best Current Practice for OAuth 2.0 Security · IETF / RFC Editor · 2025-01-01
Run it yourself: 3 commands, no API keys.
github.com/mandarelabs/mandare →