AI
OpenAI Responses API Remote MCP: Approvals, Filters, Tunnels
October 202612 min read

Add a tool of type mcp to the tools array with a server_label and either server_url for a public server or tunnel_id for a private one behind Secure MCP Tunnel. OpenAI then lists the server's tools as an mcp_list_tools item and records each call as an mcp_call item. Add an OAuth token in the authorization field if the server needs one.
Yes. By default OpenAI asks for approval before any data is shared with a remote MCP server, and the call comes back as an mcp_approval_request item. You answer it with an mcp_approval_response item in a new request that uses previous_response_id. Set require_approval to never, or to a never filter with tool_names, only for tools you trust.
allowed_tools controls which of the server's tools the model can see at all, so a tool left off the list can never be called and costs no tokens. require_approval controls which visible tools can run without your code approving each call. Use both: a short allowed_tools list, and approval for every tool that writes data.
The Responses API deliberately does not store the authorization value and does not show it in the Response object, so a stored response can never leak the token. That means every create call, including each follow-up in an approval loop, must include the token again. Refresh short-lived tokens between rounds, because a human approval can outlast them.
Yes, through Secure MCP Tunnel. You create a tunnel in Platform tunnel settings, run tunnel-client inside your network with outbound HTTPS to api.openai.com on port 443, and pass the tunnel_id in the mcp tool instead of server_url. No inbound port is opened, but the OAuth authorization server is not tunnelled and must still be reachable.

Key Takeaway
The OpenAI Responses API remote MCP tool lets OpenAI call your MCP server for the model. Point it at a public server with server_url or a private one with tunnel_id, list only the tools you need in allowed_tools, require approval for every write, and resend the OAuth token on each request because the API never stores it.
The first time I pointed the Responses API at an internal ERP MCP server, the model could see every tool the server exposed, including the ones that write stock adjustments. Nothing had gone wrong yet, but nothing was stopping it either. The MCP server was built for a desktop client where a person clicks Allow on each call. In the Responses API, that person is your code, and if you configure the tool loosely, no one is clicking anything.
This post covers the client side of the OpenAI Responses API remote MCP tool: how OpenAI reaches your server (server_url, the Secure MCP Tunnel, and the deprecated connector_id), how allowed_tools and require_approval narrow what the model can do, how to run the mcp_approval_request loop in TypeScript, and why the authorization token has to travel with every call. Every field name and behaviour comes from OpenAI's MCP guide, the tunnel guide and the typed definitions in the official Python SDK. Building the server itself is a separate topic, covered in the earlier MCP server posts on this blog.
You add a tool of type mcp to the tools array. When the model decides it needs the server, the API calls the server's tools list and records the result as an mcp_list_tools output item. When the model calls a tool, OpenAI's infrastructure, not your process, sends the request to the MCP server and records the arguments and the output in an mcp_call item. If approval is required, the call is paused and an mcp_approval_request item appears instead. These are the fields that control that behaviour:
| Field | What it does | What to watch |
|---|---|---|
| server_label | Required name for the server, echoed in every MCP output item | Use it as the key for your audit log and approval rules |
| server_url / tunnel_id / connector_id | How OpenAI reaches the server. One of the three must be set | connector_id is deprecated for models released after 1 September 2026 |
| authorization | OAuth access token sent to the server | Not stored and not returned in the Response. Send it every time |
| headers | Optional map of extra HTTP headers for the server | The no-storage promise is documented for authorization, so keep the token there |
| allowed_tools | Array of tool names, or a filter object with tool_names and read_only | read_only trusts the server's own readOnlyHint annotation |
| require_approval | The string always or never, or a filter object with always and never keys | If you leave it out, every call needs approval |
| defer_loading | With tool search, loads tool definitions only when the model needs them | Useful for servers with dozens of tools |
The guide says the API works with servers that speak either Streamable HTTP or the older HTTP with SSE transport. There is no per-call fee. You pay for the tokens used to import tool definitions and for the tool calls themselves, and that is why the size of the tool list matters more than most people expect.
The key thing to understand about this tool is where the network call comes from. With function calling, your backend runs the tool. With the MCP tool, OpenAI connects to the server, so the server must be reachable from OpenAI, not just from your app. That leaves three options:
The connector_id deprecation will quietly break code. A request that works on gpt-5.2, the model OpenAI's legacy connector examples still use, is not guaranteed to work when you change the model string to a newer one. If you have connector_id anywhere in your codebase, the migration is to find an official remote MCP server for that service and switch to server_url, or to put your own server behind a tunnel. Do it before you upgrade the model, not as part of the same change.
Most ERP MCP servers I would connect should not be on the public internet. The Secure MCP Tunnel avoids that. You create a tunnel in the Platform tunnel settings, run tunnel-client on a host that can already reach the server, and the client long-polls OpenAI for queued MCP requests, forwards each JSON-RPC request locally and posts the response back. It needs outbound HTTPS to api.openai.com on port 443, or mtls.api.openai.com when control-plane mTLS is configured, and no inbound port at all.
# Run this inside the network that can already reach the ERP's MCP server.
# It only needs OUTBOUND HTTPS to api.openai.com:443. No inbound port opens.
export CONTROL_PLANE_API_KEY="sk-..." # runtime API key for tunnel-client
tunnel-client init \
--sample sample_mcp_stdio_local \
--profile erp-stdio \
--tunnel-id tunnel_0123456789abcdef0123456789abcdef \
--mcp-command "python /opt/erp-mcp/server.py"
# An HTTP server takes --mcp-server-url https://mcp.internal.example.com/mcp
# instead of --mcp-command.
tunnel-client doctor --profile erp-stdio --explain
tunnel-client run --profile erp-stdio
# Keep "run" alive under systemd or as a sidecar: while it is down, every
# tool call routed through the tunnel fails.
# Responses API side: tunnel_id REPLACES server_url.
# Never paste the OpenAI-hosted tunnel endpoint into server_url.
{
"type": "mcp",
"server_label": "erp",
"tunnel_id": "tunnel_0123456789abcdef0123456789abcdef"
}Before this works, two things need to be in place. The first is permissions: creating a tunnel needs Tunnels Read and Manage, while running tunnel-client or using the tunnel needs Tunnels Read and Use. These are set at the organisation level, not the project level, and the guide warns a new role can take up to 30 minutes to propagate. The second is OAuth: discovery metadata passes through the tunnel, but the authorization server itself is not tunnelled. If your identity provider is reachable only on the internal network, the OAuth flow can fail even when the MCP server is reachable. The guide suggests running the client as a Kubernetes sidecar next to the server, as a separate deployment, or as a systemd service on a VM.
tunnel-client exposes /healthz, /readyz and /metrics, plus a local admin UI at /ui that listens only on loopback by default. Point your existing monitoring at /readyz. While the client is disconnected, every tunnelled tool call fails, so treat it like any other production dependency, with alerts.
allowed_tools is the most effective control in this tool, and it costs nothing to set. A tool that is not on the list is never imported, so the model cannot call it, cannot be tricked into calling it, and you do not pay tokens for its definition. OpenAI's guide notes that exposing many tools adds cost and latency. My rule is simple: list the tools by name, read-only ones first, and add a write tool only when there is a real feature that needs it.
The filter form, an object with tool_names and read_only, is convenient, but read_only works by matching the server's readOnlyHint annotation. The MCP specification says clients must treat tool annotations as untrusted unless they come from a trusted server. A third-party server that labels a delete tool as read-only would get through that filter. For a server you did not write, use explicit names. For your own server, read_only is fine because you control both the code and the annotation.
For servers with a long tool list, defer_loading set to true is the other option. It works together with tool search: the model sees the server label and description, and loads individual function definitions only when it decides to search that server. This saves tokens, but it does not limit anything. A deferred tool can still be called, so defer_loading is a way to reduce cost, not a security control, and you still need allowed_tools.
By default, OpenAI asks for approval before any data is shared with a remote MCP server, so leaving require_approval out gives the safest behaviour. The string never skips approvals for every tool on that server. The object form takes a never key, an always key, or both, and each one holds tool_names and optionally read_only. OpenAI's own example skips approval only for named read tools. This is the builder I use, so the policy is computed in one place and cannot drift:
import type OpenAI from "openai";
type McpTool = OpenAI.Responses.Tool.Mcp;
// Tools I have read the server code for and know only run SELECTs.
const READ_ONLY = ["get_stock_level", "get_sales_order", "list_open_invoices"];
// The one write the model may propose. A person approves every call.
const WRITES = ["draft_stock_adjustment"];
// Wrong: if READ_ONLY ever ends up empty (a typo, a feature flag), a
// community bug report says an empty "never" list produced NO approval
// requests at all, the opposite of what the object appears to say.
const wrongPolicy: McpTool["require_approval"] = {
never: { tool_names: READ_ONLY },
};
// Right: fall back to the string form, which has only one meaning.
function approvalPolicy(skip: string[]): McpTool["require_approval"] {
return skip.length > 0 ? { never: { tool_names: skip } } : "always";
}
export function erpMcpTool(accessToken: string): McpTool {
return {
type: "mcp",
server_label: "erp",
server_description:
"Stock levels, sales orders and open invoices for the Jakarta warehouse.",
server_url: "https://mcp.erp.example.co.id/mcp",
// Not stored by the API and not echoed in the Response object,
// so it has to be sent again on every single create call.
authorization: accessToken,
// Anything else the server exposes is never shown to the model.
allowed_tools: [...READ_ONLY, ...WRITES],
require_approval: approvalPolicy(READ_ONLY),
};
}The model sees four tools and can run three of them without waiting. The fourth, the write, always comes back to my code as an approval request. Combined with allowed_tools, this gives two separate limits: what the model knows exists, and what it can run without a person. OpenAI's risk section recommends exactly this: use require_approval and allowed_tools together so that every sensitive action goes through an approval flow.
A December 2025 bug report on the OpenAI developer forum describes require_approval set to a never filter with an empty tool_names array, which produced no approval requests at all, and says combining never and always also behaved unexpectedly. There was no reply from OpenAI in the thread. I cannot confirm whether it has been fixed, so I never send an empty list: if nothing should skip approval, send the string always.
An approval request is not an error and not a callback. It is an output item with an id, the server_label, the tool name and the arguments as a JSON string. To answer it, you create a new response that references the old one with previous_response_id and sends an mcp_approval_response item with approve set to true or false. The SDK type also has an optional reason field. This is the loop, with the audit logging I would not ship without:
import OpenAI from "openai";
import { erpMcpTool } from "./erp-mcp-tool";
// tokens, audit and reviewer are your own modules: an OAuth token store,
// an append-only log, and whatever decides (a person, or a rule you wrote).
import { tokens, audit, reviewer } from "./erp-agent-deps";
const client = new OpenAI();
const MODEL = "gpt-6-astra";
const MAX_APPROVAL_ROUNDS = 3;
export async function askErp(userId: string, question: string) {
let response = await client.responses.create({
model: MODEL,
tools: [erpMcpTool(await tokens.forUser(userId))],
input: question,
});
for (let round = 0; round < MAX_APPROVAL_ROUNDS; round++) {
const decisions: OpenAI.Responses.ResponseInputItem.McpApprovalResponse[] = [];
for (const item of response.output) {
if (item.type !== "mcp_approval_request") continue;
// item.arguments is the exact JSON the MCP server will receive.
// Approve the payload, not the tool name.
const args = JSON.parse(item.arguments);
await audit.log({ userId, server: item.server_label, tool: item.name, args });
const approve = await reviewer.decide(item.name, args);
decisions.push({ type: "mcp_approval_response", approval_request_id: item.id, approve });
}
if (decisions.length === 0) break; // nothing pending: the answer is final
response = await client.responses.create({
model: MODEL,
previous_response_id: response.id,
// The token from the first call is gone. Resend it, refreshed if needed.
tools: [erpMcpTool(await tokens.forUser(userId))],
input: decisions,
});
}
// A failed call does not throw: it lands in the item's error field.
for (const item of response.output) {
if (item.type === "mcp_call" && item.error) {
await audit.log({ userId, tool: item.name, error: item.error });
}
}
return response.output_text;
}Each part of that loop is there for a reason:
The guide is explicit: the Responses API does not store the value of authorization, the value is not visible in the Response object, and you must send it with every create request. This is deliberate, because a token that is never stored cannot leak from a stored response. In practice, every follow-up call in the approval loop needs a valid token, so your token store has to refresh it between rounds. A person approving a write can easily take longer than a short-lived token lasts.
Use the end user's own OAuth token, not one service account for everyone. The MCP server then enforces that user's permissions, so even an approved call cannot read a branch or company the user could not open in the ERP. Approval is your control, and the token scope is the server's control. You want both.
The headers field, an optional map of HTTP headers that the SDK types describe as being for authentication or other purposes, is for servers that expect something other than a bearer token, such as a tenant id or an API key in a custom header. I use it for non-secret routing values only. The no-storage guarantee in OpenAI's guide is written for the authorization field, and I would rather not assume it also covers headers.
Keep the mcp_list_tools item in context. The guide says that as long as it is present, the API does not fetch the tool list from the server again on each turn, and OpenAI recommends keeping it for every conversation to reduce latency. previous_response_id handles this automatically. If you manage context yourself, for example with store set to false, you have to pass that item back as input, or you pay the import latency and tokens again on every turn.
On data: with store set to true, the data sent to MCP servers is already logged by the API for 30 days unless Zero Data Retention is enabled, and the guide still recommends keeping your own log. The MCP tool works with Zero Data Retention and data residency, but those guarantees end where the MCP server begins. Data sent to the server follows the server's own retention and residency policies. For an Indonesian company with data-location requirements, the tunnel setup helps here, because the server and its database stay on your own infrastructure. Also treat URLs returned in mcp_call output as untrusted. The guide warns against fetching or embedding them unless you trust the domain, because requesting an attacker-chosen URL is itself a way to leak data.
The rule I follow now: the MCP tool should be configured for the least trust that still works. Use explicit allowed_tools names, use the never filter only for tools whose code I have read, use the string always instead of an empty list, approve the arguments rather than the tool name, and send a fresh user-scoped token on every call. The tunnel keeps private servers private. Everything else is code that has to be written on purpose.