Security
AI Agent Least Privilege: Credentials, Tokens and Secrets
October 202611 min read

It means each tool the agent can call gets only the access it needs for its one job, and nothing more. OWASP's LLM06 Excessive Agency entry describes the failure as excessive functionality, excessive permissions and excessive autonomy. Least privilege caps all three, so a successful prompt injection can only do what that narrow credential allows.
Anything that enters the model's context can be repeated, summarised or exfiltrated by an injected instruction. OpenAI's remote MCP guide warns that a malicious server can exfiltrate sensitive data from anything in context. Let the model name the action and have your executor code, outside the model, attach the credential.
You list a domain, a placeholder name and the secret value inside the shell tool's network_policy. The model and the container only see the placeholder, such as $API_KEY, and a sidecar substitutes the real value only for requests to the approved domain. Outbound access also requires an org allow list in the dashboard plus an explicit network_policy on the request.
No. The Responses API does not store the value of the authorization field and does not show it in the Response object, so you must send it with every request. That lets you keep the token in your own store, scoped to the user who is present.
Token passthrough is when an MCP server accepts a client's token without checking it was issued for that server and forwards it unchanged to a downstream API. The MCP 2026-07-28 specification requires servers to validate the token audience and forbids accepting or transiting other tokens. Passthrough breaks audit trails and lets a stolen token use the server as a proxy.

Key Takeaway
AI agent least privilege means designing credentials on the assumption that a prompt injection will succeed. Keep every secret out of the model's context, substitute credentials outside the model with placeholders, send audience-bound OAuth tokens per request, give each tool its own narrowly scoped identity, and require human approval before any tool writes.
Picture an accounts-payable helpdesk agent wired into an ERP. It reads vendor invoices, answers status questions and drafts payments. One afternoon a vendor PDF arrives with a line of white-on-white text telling the agent to export the vendor master list to an outside address. Whether that line becomes an incident has very little to do with the prompt. It depends on what the agent's credentials allow it to reach, and whether a secret ever sat where the model could read it.
This post covers the identity and secret layer of an agent, the part that limits damage after an injection gets through. It works from the OWASP LLM06 Excessive Agency entry, the OpenAI hosted shell and remote MCP guides, and the MCP 2026-07-28 authorization specification, and ends with a review checklist you can run against any agent before it reaches production. Sandboxing and injection detection are separate topics; here the question is narrower: when the model is fooled, what can it actually do?
OWASP names the failure Excessive Agency and splits it into three root causes: excessive functionality, excessive permissions and excessive autonomy. Each one maps cleanly onto a credential decision, which is why the identity layer is where agent security is won or lost. Filters and classifiers reduce how often an injection lands; credentials decide what it costs when one does.
| OWASP root cause | What it looks like in an agent | Credential control that caps it |
|---|---|---|
| Excessive functionality | A mailbox summariser that can also send mail, or an ERP reader that exposes a generic query tool | Expose only named tools; filter remote MCP servers with allowed_tools so unused tools are never imported |
| Excessive permissions | One admin API key shared by every tool the agent calls | One service account or OAuth scope per tool, read-only by default, write scopes requested only on step-up |
| Excessive autonomy | Payments posted or records deleted without anyone looking | Approval gates on every write tool, decided by code and a person rather than by the model |
The OWASP example is worth remembering because it is so ordinary: an assistant built to summarise email also had permission to send it, and an injected email told it to forward sensitive messages to the attacker. The listed fixes are all permission changes, not prompt changes. Remove the send capability, authenticate with a read-only OAuth scope, or require the user to approve each send.
OpenAI's remote MCP guide states the problem plainly: a malicious server can exfiltrate sensitive data from anything that enters the model's context. The same is true of a poisoned web page, a hostile PDF or a confused model. So the first rule is structural. An API key written into a system prompt, a tool description or a tool argument is a key the model can repeat. The model should name an action; your executor, running outside the model, should decide which credential that action uses.
// Wrong: the API key is part of the prompt. Anything in context can be
// repeated, summarised or sent somewhere by an injected instruction.
const tools = [{
type: "function",
name: "query_erp",
description: `Call https://erp.example.co.id/api with header X-Api-Key: ${process.env.ERP_KEY}`,
parameters: { type: "object", properties: { path: { type: "string" } } },
}];
// Right: the model only names an action and its arguments.
// The executor - your code, outside the model - chooses the credential.
const CREDENTIAL_FOR_TOOL = {
read_invoice: "erp/svc-agent-invoice-reader", // read-only account
draft_payment: "erp/svc-agent-ap-drafter", // may create drafts, never post
} as const;
type ToolName = keyof typeof CREDENTIAL_FOR_TOOL;
export async function executeTool(name: ToolName, args: unknown) {
const secretRef = CREDENTIAL_FOR_TOOL[name];
if (!secretRef) throw new Error(`no credential mapped for tool ${name}`);
// getSecret is your own seam over Vault / a cloud secret manager.
const token = await getSecret(secretRef);
const res = await callErp(name, args, token);
// The tool result goes back into context, so it is a leak path too:
// drop echoed headers, signed URLs and anything shaped like a token.
return redactSecrets(res);
}The mapping from tool name to secret reference is the most useful line in the whole design. It makes the credential a property of the tool, not of the conversation, so no amount of clever input can make read_invoice run with the drafting account. It also gives you a single file to review when someone asks what the agent can touch.
Tool results are a leak path too. An ERP error that echoes the request headers, a debug response that includes the Authorization value or a pre-signed download URL all flow straight back into context, where the next injected instruction can find them. Redact tool output before returning it, and never let a shell tool print its own environment variables.
Code-executing agents make rule one harder, because the code the model writes needs to authenticate somewhere. OpenAI's hosted shell answers this with domain_secrets. Each entry holds a target domain, a friendly name and the secret value. The model and the runtime see only a placeholder such as $ERP_TOKEN; an auth-translation sidecar applies the raw value only for approved destinations, and the documentation states that raw values do not persist on API servers and do not appear in model-visible context.
import OpenAI from "openai";
const client = new OpenAI();
const response = await client.responses.create({
model: process.env.AGENT_MODEL!,
input: "Fetch today's open purchase orders and summarise them by vendor.",
tools: [
{
type: "shell",
environment: {
type: "container_auto",
// No network_policy = no outbound network at all (the default).
network_policy: {
type: "allowlist",
// Must be a subset of the org allow list set in the dashboard,
// or the request fails instead of silently widening access.
allowed_domains: ["erp-api.example.co.id"],
domain_secrets: [
{
domain: "erp-api.example.co.id",
name: "ERP_TOKEN",
// The model and the shell only ever see $ERP_TOKEN.
// The sidecar swaps in the real value for this domain only.
value: process.env.ERP_READONLY_TOKEN!,
},
],
},
},
},
],
});The network side matters as much as the secret side. Hosted containers have no outbound network access by default. To enable it, an admin must first configure the organisation allow list in the dashboard, and the request must also set network_policy explicitly. The org list defines the full set of reachable domains, the request-level list can only narrow it, and a request that names a domain outside the org list fails rather than quietly widening access.
OpenAI also warns that allowlisting itself creates an exfiltration route: only allow domains you trust and that an attacker cannot use to receive data. A general-purpose paste site or an object store that accepts anonymous uploads defeats every other control on this page. The same guide recommends recording the hosts each session requested and the destinations it actually reached, which is the log you will want after an incident.
Remote MCP moves credentials across a trust boundary, and both ends have rules. On the OpenAI side, the Responses API does not store the value you pass in the authorization field and does not show it in the Response object, so you must send it with every request. That is a feature: the token lives in your token store, scoped to the user who is actually present, and nothing on the platform side keeps a long-lived copy.
// Client side: the token belongs to the human user, is bound to ONE
// MCP server (RFC 8707 resource = its canonical URI), and is resent on
// every call because the Responses API does not store it.
const MCP_URL = "https://mcp.erp.example.co.id/mcp";
const userToken = await tokens.getAccessToken({ userId, resource: MCP_URL });
await client.responses.create({
model: process.env.AGENT_MODEL!,
input,
tools: [{
type: "mcp",
server_label: "erp",
server_url: MCP_URL,
authorization: userToken,
allowed_tools: ["read_invoice", "draft_payment"],
require_approval: { never: { tool_names: ["read_invoice"] } },
}],
});
// Server side: accept only tokens minted FOR this server.
import { createRemoteJWKSet, jwtVerify } from "jose";
const JWKS = createRemoteJWKSet(new URL("https://id.example.co.id/.well-known/jwks.json"));
export async function authenticate(req: Request) {
const bearer = req.headers.get("authorization")?.replace(/^Bearer /, "");
if (!bearer) return challenge(401, "invoices:read");
const { payload } = await jwtVerify(bearer, JWKS, {
issuer: "https://id.example.co.id",
audience: MCP_URL, // a token for any other API is rejected here
});
// Never forward 'bearer' to the ERP. Exchange or use this server's own
// credential downstream - token passthrough is forbidden by the spec.
return payload;
}The MCP 2026-07-28 authorization specification adds the audience rules. Clients must send the RFC 8707 resource parameter, set to the server's canonical URI, in both the authorization and token requests. Servers must validate that a token was issued specifically for them, must only accept tokens valid for their own resources, and must not accept or transit any other tokens. The security best practices page calls the forbidden pattern token passthrough: forwarding a client's token unchanged to a downstream API breaks audit trails and turns the server into a confused deputy.
The same release hardens the flow against mix-up attacks. Clients record the expected issuer before redirecting and validate the iss parameter from RFC 9207 before sending the authorization code to any token endpoint, so a code from an honest server cannot be redeemed at a hostile one. PKCE alone does not stop this, because the client would hand its code_verifier to the attacker's token endpoint. Client ID Metadata Documents replace Dynamic Client Registration as the recommended registration path, with DCR kept only for backwards compatibility.
A single agent should not be a single principal. Give each tool its own service account or OAuth scope, sized to the one verb that tool performs. In the accounts-payable example, the identity split looks like this:
| Tool | Identity | Scope | If injected, worst case |
|---|---|---|---|
| read_invoice | User's delegated token | invoices:read | Reads invoices the user could already see |
| draft_payment | User's token after step-up | payments:draft | Creates a draft that a person must still release |
| lookup_vendor | Read-only service account | vendors:read on name and status fields only | Learns vendor names, not bank details |
| post_payment | Not given to the agent | None | Nothing; posting stays a human action in the ERP |
MCP gives you the mechanism to keep the starting scope small. The specification tells servers to put the required scope in the WWW-Authenticate challenge and tells clients to treat it as authoritative. When a token lacks a scope at runtime, the server answers 403 with error insufficient_scope, and the client re-authorises with the union of what it already had and what the challenge asks for.
# The agent's token carries invoices:read only. It tries to draft a payment:
HTTP/1.1 403 Forbidden
WWW-Authenticate: Bearer error="insufficient_scope",
scope="payments:draft",
resource_metadata="https://mcp.erp.example.co.id/.well-known/oauth-protected-resource",
error_description="Drafting a payment needs payments:draft"
# The client re-authorises with the UNION (invoices:read payments:draft),
# the user sees exactly one new permission on the consent screen,
# and the elevation is logged with a correlation id.
# payments:post is never requested by the agent's client at all.The best practices page lists the mistakes that undo this: publishing every scope in scopes_supported, wildcard or omnibus scopes such as all or full-access, bundling unrelated privileges to avoid future prompts, and treating the scopes claimed in a token as sufficient without server-side authorisation logic. That last one matters most for ERP work, where the real rule is usually row-level: this user may only see invoices for their own company code.
Scopes limit what is possible; approvals decide what actually happens. The Responses API exposes require_approval on an MCP tool as never, always, or an object listing the tool names that may skip approval. Pair it with allowed_tools so a server's destructive tools are never imported in the first place. Each pending call then arrives as an mcp_approval_request item, and your code answers with an mcp_approval_response carrying the request id and a boolean.
// The approval decision is made by code and a human, never by the model.
const AUTO_APPROVE = new Set(["read_invoice"]);
const MAX_UNREVIEWED_IDR = 0; // every payment draft goes to a person
for (const item of response.output) {
if (item.type !== "mcp_approval_request") continue;
const args = JSON.parse(item.arguments);
const needsHuman =
!AUTO_APPROVE.has(item.name) ||
(item.name === "draft_payment" && args.amountIdr > MAX_UNREVIEWED_IDR);
const approve = needsHuman
? await approvals.ask({ user: userId, tool: item.name, args }) // UI, chat, email
: true;
await audit.log({ requestId: item.id, tool: item.name, args, approve });
pending.push({
type: "mcp_approval_response",
approval_request_id: item.id,
approve,
});
}Keep the decision outside the model. A rule like every payment draft needs a person, or anything above a rupiah threshold needs a second approver, is ordinary code with an audit row behind it. Asking the model whether a call looks safe hands the decision back to the component the attacker has just compromised.
Default to an allowlist of read tools that skip approval, not a blocklist of write tools that need it. A new tool added to the MCP server next month then starts out gated instead of silently auto-approved.
Run this before an agent gets production credentials, and again whenever it gains a tool:
Items six and eight are the ones teams most often skip, and they are the ones that matter on the worst day. If one leaked token exposes everything, or if revoking it takes the whole helpdesk offline, the design is still a shared admin key with extra steps.
Prompt defences decide how often an agent is fooled; credentials decide what a fooled agent can do. Keep secrets where the model cannot read them, bind every token to one user and one server, give each tool the narrowest identity that lets it do its job, and keep a person between the agent and any irreversible write.