Entry 007Comparisons9 min readBy Mandare Labs

20 tools that cap, identify or audit AI agents (2026)

Mandare Labs publishes this journal, and Mandare, which is v0.1.x software, ranks #1 on the criteria below. Twenty tools were scored on four controls that must run outside the agent's own process: a signed identity, a hard spending cap, a tamper-evident record and an offline kill switch. Scores come from each vendor's own documentation, read on 7 October 2026. Mandare scores four of four. No other tool scores more than two.

How we ranked

Four controls count, and each only when it runs outside the agent's own process and a documentation page shows it. A tool's score is the number of full "Yes" results. "Part" means a weaker form is documented. "No" means the pages checked did not show it.

  1. Signed identity. Each agent holds its own key or credential, and the enforcement point checks a signature or signed token (RFC 9421, RFC 9449). A shared or per-key bearer secret is Part.
  2. Hard cap. A money limit per agent, key or identity, enforced by a gateway, issuer or provider that refuses calls at the limit. Provider caps sit one level up: Anthropic documents organization and workspace limits, OpenAI organization and project limits. Soft caps and rate limits are Part.
  3. Tamper-evident record. Written by infrastructure, with a documented integrity mechanism such as a hash chain, signatures or write-once storage (RFC 6962, Article 12 of the EU AI Act). An ordinary log is Part.
  4. Offline kill switch. The operator revokes an agent without its cooperation and without a vendor-hosted service. RFC 7009 revocation is a call to the authorization server, so hosted products score Part.

Ties are broken in this order, fixed before scoring: self-hostable, open source (enterprise-licence features do not count), no beta, preview, early-access or maintenance label, lower integration effort, then alphabetical. Inside a tied score the order is a tie-break, not a quality verdict.

These four controls are the ones Mandare is built around. Another list would rank differently. Every vendor statement below is paraphrased from that vendor's own pages, read on 7 October 2026, with the main page linked.

ToolIdentityCapRecordKillYes
1 MandareYesYesYesYes4
2 LiteLLMPartYesPartYes2
3 SPIFFE/SPIREYesNoPartYes2
4 Kong AI GatewayYesYesPartPart2
5 OpenRouterYesYesPartPart2
6 SkyfireYesYesPartPart2
7 PortkeyPartYesPartPart1
8 LangSmithPartYesPartPart1
9 Entra Agent IDYesNoPartPart1
10 Stripe IssuingPartYesPartPart1
11 KeycardYesNoPartPart1
12 Okta for AI AgentsYesNoPartPart1
13 LangfuseNoPartPartPart0
14 NeMo GuardrailsNoNoNoNo0
15 HeliconePartPartPartNo0
16 LakeraNoNoPartNo0
17 Cloudflare AI GatewayPartPartPartPart0
18 Vercel AI GatewayPartPartPartPart0
19 AembitPartNoPartPart0
20 Auth0 for AI AgentsPartNoPartPart0

The ranking

1. Mandare

Open-source gateway for AI agents, v0.1.0 on npm. Covers all four. Each agent signs requests with its own Ed25519 key (concepts). Mandates cap spend per transaction, day, task and total. The ledger is hash-linked and written by the door. mandare kill revokes locally, offline (CLI). Best for: one local gateway enforcing all four. Limits: v0.1.x, no external audit yet, no production-readiness claim, owner verification is a mock, truncation is caught only with a witness on separate infrastructure (threat model). Pick LiteLLM instead if you need 100+ providers.

2. LiteLLM

Open-source proxy with one interface to many providers. Covers: requests fail once a key crosses its max_budget, and a block call runs against your own proxy and Postgres. Identity is a bearer virtual key. JWT auth and audit logs are Enterprise features. Best for: self-hosted budgets per key, team and user. Limit: the default budget is null and unchecked until set, and budgets need a database.

3. SPIFFE/SPIRE

Graduated CNCF workload identity, Apache-2.0. Covers: each registered workload gets short-lived SVIDs, and spire-server entry delete removes an entry on your own server. Default lifetimes are one hour (X.509) and five minutes (JWT). The audit log is off by default and goes to the ordinary log output. Best for: attested identity without a vendor service. Limit: identity only, no spend control.

4. Kong AI Gateway

Gateway for AI traffic, self-hosted or on Konnect. Covers: it verifies an IdP token and maps it to a consumer, one per agent if you set it up that way. Cost-based rate limiting returns 429, though a call's cost lands on the next request. Best for: governing AI traffic per consumer through your IdP. Limit: the cost-based limit is documented as AI Gateway Enterprise only.

5. OpenRouter

Hosted API for hundreds of models behind one endpoint. Covers: an exhausted per-key credit limit returns 402. Workload identity federation checks a signed workload token, on Business and Enterprise plans. Guardrail budgets are enforced per user and per key, with daily, weekly or monthly resets. Disabling the key cuts access through the hosted service. Best for: per-key spend control across models. Limit: no self-hosting documented.

6. Skyfire

Identity and payments layer for agents. Covers: tokens are signed JWTs that name the buyer agent, signed by Skyfire. A token's amount may not exceed the wallet balance (create token), and a charge past the remaining balance fails. Best for: letting a seller verify an agent's payment credential. Limit: deactivating an agent does not void issued tokens, and no periodic per-agent cap is documented.

7. Portkey

Gateway with budgets, guardrails and observability. Its docs now carry the name Prisma AIRS AI Gateway. Covers: key budgets stop usage at the limit, on Enterprise and select Pro plans. Keys are bearer secrets, and its docs say virtual keys migrated to Model Catalog. The audit log tracks administrative actions. Best for: governance across keys and workspaces. Limit: no log integrity mechanism documented.

8. LangSmith

Tracing and evaluation platform with an LLM Gateway. Covers: a spend policy caps an organization, workspace, API key or user and returns 402, for gateway-routed calls. The gateway is beta, and application traces are sent by the SDK. Audit logs are called tamper-resistant, name no mechanism and need Enterprise. Best for: tracing plus optional per-key caps. Limit: the gateway is not in the v0.16.0 self-hosted stable release.

9. Microsoft Entra Agent ID

Identity framework extending Entra to agents, generally available per its docs. Covers: the token subject is the agent identity, and downstream APIs verify its signature. Admins can disable an agent identity in the hosted admin center. Agent identities hold no credentials of their own. Best for: agents inside a Microsoft tenant. Limit: no spend control found in the pages checked.

10. Stripe Issuing for agents

Virtual cards for agents. Covers: spending limits can decline a purchase before authorization, and each agent can hold its own card. Best for: issuer-enforced limits on purchases. Limit: access is by application, aggregation is best-effort with up to 30 seconds of delay, and later tips and fees can exceed a limit. Freezing a card is a call to Stripe's API.

11. Keycard

Identity and access platform for agents, Early Access per its docs. Covers: each agent is a registered application, and Keycard verifies the signature of its runtime-issued token. Audit records export to S3 with no integrity mechanism documented, and local development falls back to a client secret. Best for: policy checks on agent access to tools, outside the agent. Limit: a token already issued keeps working until it expires.

12. Okta for AI Agents

Registration and governance of agents inside an Okta org. Covers: each agent's public key validates its token-exchange requests. Deactivation is a call to Okta's hosted API, and access requires a subscription. Okta's pages disagree on maturity: Beta in the API reference, Early Access and Generally Available in release notes. Best for: least-privilege governance in an Okta tenant. Limit: no spend limit shown.

13. Langfuse

Open-source tracing and cost platform, self-hosted with Docker. Some add-on features need a licence key. Covers: none in full. Cost tracking offers alerts at a spend threshold, not refusal. Audit logs are called immutable but need an Enterprise plan, and no integrity mechanism is described. Agent traces come from the application's own SDK, and project API keys are shared bearer keys. Best for: self-hosted tracing and cost visibility. Limit: measurement, not enforcement.

14. NeMo Guardrails

Apache-2.0 Python library for input, output and tool-call rails. Covers: none of the four. Its FAQ says the application owns identity enforcement. Its deployment page rates the library's API server as limited for production reliability, with no high availability out of the box. Best for: content and tool-call checks inside your own application. Limit: it runs in the application, with no spend cap or credential control.

15. Helicone

Open-source observability, self-hosted on Docker or Kubernetes, with a gateway reached by changing the base URL. Covers: a cost limit in cents over a time window, set by a request header. A blog post of 3 March 2026 says Mintlify acquired it and its services run in maintenance mode. Best for: logging and cost tracking behind one URL. Limit: no key revocation documented.

16. Lakera

Its docs now describe Check Point AI Guardrails, which screens content going into and out of LLMs through an API. Agent security is early access, and self-hosting needs an Enterprise licence. Covers: none of the four. The integration page leaves the response to your application. Best for: screening content and tool allow lists. Limit: it returns a verdict. Enforcement and identity sit with your architecture.

17. Cloudflare AI Gateway

Gateway with analytics, caching and rate limiting, available on all plans and adopted with a base-URL change. Covers: spend limits return 429 and can split by agent_id, a value the caller sends. Limits are eventually consistent. A request header can exclude its own log. Best for: quick budgets in front of an existing app. Limit: per-agent scope rests on a header the agent sets.

18. Vercel AI Gateway

Managed gateway for credentials, logs, budgets and routing. Covers: budgets return 402 for a scope such as one API key. The docs call a budget "a soft cap, not a hard limit", and bring-your-own-key spend is excluded. Best for: one managed endpoint for credentials and budgets. Limit: it is a managed gateway, and deleting a key invalidates it through the hosted dashboard or API.

19. Aembit

Identity and access platform for agents and workloads. Edge components and its MCP gateway can run on your own host, but policy evaluation still calls Aembit Cloud. Covers: none in full. Revocation is a policy change, and both planes run in Aembit Cloud. Best for: replacing static agent keys with short-lived credentials. Limit: no spend limit documented, and revocation depends on the hosted plane.

20. Auth0 for AI Agents

Auth0's hosted identity platform applied to agents, with Apache-2.0 AI SDKs. Covers: none in full. Agent as Principal is Early Access. An agent holds no credentials of its own. After dissociation, existing tokens remain valid until expiry. Best for: user-facing agents calling APIs on a user's behalf. Limit: no spend control shown, and Agent Gateway is listed as shipping next.

Where Mandare is not the right choice

For Mandare's spend limits, read how to set a daily limit and what a stolen key can do. The quickstart runs the demo without API keys.

Questions

Which tool caps an AI agent's spend?

In the documentation read on 7 October 2026, LiteLLM, OpenRouter, Kong AI Gateway, Skyfire, Portkey, LangSmith's gateway, Stripe Issuing and Mandare each document a limit that refuses further use. Scope and caveats differ. Vercel calls its budget a soft cap, and Cloudflare splits budgets by a header the caller sets.

Is there a kill switch that works without a vendor cloud?

In the pages read, LiteLLM (a block call to your own proxy), SPIFFE/SPIRE (deleting a registration entry on your own server) and Mandare (mandare kill, local and offline) revoke without a vendor service. Hosted products revoke through their own API. Tokens already issued can outlive a revocation until they expire.

Do agent tools document a tamper-evident audit log?

Mandare documents hash-chained, door-signed entries and an RFC 6962 Merkle tree. For the other 19 tools, the pages read describe no integrity mechanism for records of agent calls. Kong documents signing for its admin audit log. Mandare catches truncation only with a witness on separate infrastructure.

Does a higher score mean a better tool?

No. The score counts four controls that Mandare is built around, and another list would rank differently. A tool scoring zero here can be the right choice, such as Langfuse for tracing or NeMo Guardrails for content rails. Mandare is v0.1.x and has had no external audit.

Sources

  1. Rate limits (Claude API documentation) · Anthropic
  2. Rate limits (OpenAI API documentation) · OpenAI
  3. RFC 9421: HTTP Message Signatures · IETF
  4. RFC 9449: OAuth 2.0 Demonstrating Proof of Possession (DPoP) · IETF
  5. RFC 6962: Certificate Transparency · IETF
  6. Regulation (EU) 2024/1689 (Artificial Intelligence Act), Article 12 · EUR-Lex
  7. RFC 7009: OAuth 2.0 Token Revocation · IETF
  8. LLM06:2025 Excessive Agency · OWASP Gen AI Security Project
  9. Security Best Practices (Model Context Protocol) · Model Context Protocol project
  10. LiteLLM: virtual keys and key blocking · LiteLLM
  11. LiteLLM: budgets for keys, teams and users · LiteLLM
  12. LiteLLM README · LiteLLM
  13. SPIFFE overview · SPIFFE
  14. SPIRE server configuration reference · SPIFFE
  15. SPIRE audit log · SPIFFE
  16. SPIRE README · SPIFFE
  17. Kong AI Gateway: AI auth strategy · Kong
  18. Kong: AI Rate Limiting Advanced plugin · Kong
  19. OpenRouter: API credit and rate limits · OpenRouter
  20. OpenRouter: workload identity federation · OpenRouter
  21. OpenRouter quickstart · OpenRouter
  22. Skyfire: verify and extract data from tokens · Skyfire
  23. Skyfire: create token · Skyfire
  24. Skyfire: payments and settlement · Skyfire
  25. Portkey: enforce budget and rate limits · Portkey
  26. Portkey: audit logs · Portkey
  27. LangSmith: LLM Gateway spend policies · LangChain
  28. LangSmith: LLM Gateway · LangChain
  29. LangSmith: audit logs · LangChain
  30. Microsoft Entra Agent ID: agent identities · Microsoft
  31. Microsoft Entra Agent ID: manage agent identities as an admin · Microsoft
  32. Microsoft Entra Agent ID: what's new · Microsoft
  33. Stripe Issuing: spending controls · Stripe
  34. Stripe Issuing for agents · Stripe
  35. Keycard: an agent with its own identity · Keycard
  36. Keycard: revoke once · Keycard
  37. Okta for AI Agents API · Okta
  38. Langfuse: token and cost tracking · Langfuse
  39. Langfuse: audit logs · Langfuse
  40. NeMo Guardrails: runtime security FAQ · NVIDIA
  41. NeMo Guardrails: deployment with the microservice · NVIDIA
  42. Helicone: custom rate limits · Helicone
  43. Helicone: joining Mintlify · Helicone · 2026-03-03
  44. Check Point AI Guardrails: integration · Lakera
  45. Check Point AI Guardrails: quickstart · Lakera
  46. Cloudflare AI Gateway: spend limits · Cloudflare
  47. Cloudflare AI Gateway: logging · Cloudflare
  48. Vercel AI Gateway: budgets · Vercel
  49. Aembit: AI agents use case · Aembit
  50. Aembit: planes and responsibilities · Aembit
  51. Auth0 for AI Agents: associate an agent with a client · Auth0

Run it yourself: 3 commands, no API keys.

github.com/mandarelabs/mandare →
← All journal entries