AI
OpenAI Agents API Tutorial: Hosted Agents and Sandboxes
October 202612 min read

It is a managed service, in public beta since 10 September 2026, that gives your application the Codex harness through an API. OpenAI manages sessions, orchestration, context compaction and recovery, while you provide tools and choose where the agent runs. You work with four objects: an agent, an optional environment, a durable session, and its events and items.
There is no separate platform fee. You pay the selected model's token rates, standard rates for OpenAI tools, and container rates for OpenAI-hosted sandboxes. The pricing page lists containers at 0.03 USD for 1 GB, 0.12 USD for 4 GB and 0.48 USD for 16 GB per 20-minute session.
Send an agent.session.input.message event to the same session. If a turn is active, the message steers that turn; if the session is idle, it starts a new turn with the existing conversation. To stop the work instead, send agent.session.input.cancel, which keeps the session and its earlier work.
No. Idle only means the session is ready for more input. Check for agent.session.turn.completed, turn.failed or turn.cancelled on the root turn, and inspect the agent's output, because a completed turn can still contain failed tool calls.
No. The Agents API currently supports data residency only in the United States and does not support Zero Data Retention. Choosing a self-hosted sandbox does not change that, because the harness and session state still run on OpenAI's side.

Key Takeaway
The OpenAI Agents API, in public beta since 10 September 2026, runs the Codex harness: you create a session with an agent and an environment, send a task, follow events or webhooks, steer or cancel mid-turn, then download artifacts and delete the session. There is no platform fee, but data residency is US-only and Zero Data Retention is unsupported.
The job I had in mind was dull and real. Every month-end, someone at an ERP client exports invoices per branch, totals them in a spreadsheet and chases the rows that do not reconcile. It is exactly the kind of task an agent with a Python sandbox should finish unattended, and until September 2026 doing that on OpenAI meant writing and hosting the agent loop myself.
This OpenAI Agents API tutorial walks the whole session lifecycle in TypeScript, using only fields documented in OpenAI's own guides: creating a session in an OpenAI-hosted sandbox, streaming a turn correctly, steering and cancelling, handling webhooks and function results, choosing between hosted and self-hosted sandboxes, and cleaning up. It ends with the two data-handling limits to check before anyone ships it.
The Agents API is a managed service around the open-source Codex harness. OpenAI's overview says it manages sessions, orchestration, context compaction and recovery, while your application provides tools and chooses the execution environment. The design rests on four objects:
What the harness adds over a hand-rolled loop is listed plainly: running commands in a sandbox, applying skills, connecting to tools or MCP, steering while the agent works, summarising earlier work to fit the context window, delegating to subagents and resuming a session where it left off. What stays yours is everything with business consequences: the function handlers that touch your systems, the credentials, the decision about where code runs, and the check that a finished turn actually did the job.
The beta lives under the beta.agents namespace of the official SDKs. Raw HTTP requests need the header OpenAI-Beta: agents=v1, which the SDKs add for you. Set the pieces up in this order:
import OpenAI from "openai";
// Application key: api.agents.read + api.agents.write + api.responses.write.
// It stays in your backend. Nothing in the sandbox should ever see it.
const client = new OpenAI();
const session = await client.beta.agents.sessions.create({
agent: {
model: "gpt-6-astra",
instructions:
"You reconcile ERP exports. Use Python, show the figures you used, " +
"and write every result file to /workspace/outputs.",
},
environment: {
type: "openai_hosted",
container_size: "small", // 1 vCPU / 1 GB. Default is "medium" (2 vCPU / 4 GB)
network: { access: "disabled" }, // default is "enabled" — opt out unless the task needs it
packages: { python: ["pandas==2.2.3"] },
files: [
// Inline files: 5 MiB each, 10 MiB per request, measured before base64.
{ type: "inline", path: "/workspace/invoices.csv", data: invoicesCsvBase64 },
],
},
// No input yet: an OpenAI-hosted session may start empty. The task is sent
// after the event stream is open, so no early event is missed.
});
// Store this next to your own job record — it is the handle for everything else.
console.log(session.id, session.environment.id);Two choices in that request are deliberate. Network access in an OpenAI-hosted sandbox defaults to enabled, so a job that only crunches a CSV should say disabled, or restricted with an allowed_domains list of 1 to 100 exact host names. And the session is created without input: the guides say to subscribe to the event stream before sending work, because a stream does not replay events you were not listening for. The documented sizes are small at 1 vCPU and 1 GB, medium at 2 vCPU and 4 GB, which is the default, and large at 4 vCPU and 16 GB.
Check setup before trusting the sandbox. A successful create call only means setup has started. Retrieving /v1/agents/environments/ followed by the session's environment.id returns provisioning while packages and setup_commands run, and connected once they succeed. A nonzero setup command exit status stops the agent from starting at all, which makes a setup command the cheapest place to assert that a required file or package is present.
A turn is one cycle of work inside a session. The helper below opens the stream, sends the task, prints text deltas and returns only when the root turn completes. It is adapted from the send-and-stream helper in OpenAI's Events and items guide, and every case in the switch is there for a reason the docs spell out.
async function runTask(client: OpenAI, sessionId: string, text: string) {
// Subscribe BEFORE sending input — streams do not replay missed events.
const events = await client.beta.agents.sessions.events.stream(sessionId);
try {
await client.beta.agents.sessions.events.create(sessionId, {
events: [
{
type: "agent.session.input.message",
input: [{ role: "user", content: [{ type: "input_text", text }] }],
},
],
});
for await (const event of events) {
switch (event.type) {
case "agent.session.turn.output_text.delta":
process.stdout.write(event.delta); // deltas may be absent; .done carries full text
break;
case "agent.session.idle":
continue; // Wrong to treat as success: idle only means "ready for input"
case "error":
throw new Error(event.error.message);
case "agent.session.failed":
case "agent.session.environment.failed":
throw new Error(`Agent lifecycle failure: ${event.type}`);
case "agent.session.turn.failed":
case "agent.session.turn.cancelled":
if (event.turn.subagent_id === null) throw new Error(event.type);
break; // a subagent's turn failing does not end the root turn
case "agent.session.turn.completed":
if (event.turn.subagent_id === null) return; // still inspect tool results
break;
}
}
throw new Error("Stream closed before the turn ended. Retrieve saved items.");
} finally {
events.controller.abort(); // closing the stream does NOT cancel the turn
}
}
await runTask(
client,
session.id,
"Sum the amount column per branch in /workspace/invoices.csv and " +
"write /workspace/outputs/totals.json. Read it back to verify it.",
);The trap I would have walked straight into is agent.session.idle. It fires when the session is ready for more input, not when the work succeeded, so code that treats idle as done will report a failed reconciliation as finished. The outcome events are turn.completed, turn.failed and turn.cancelled, and even a completed turn does not guarantee every tool call succeeded, so read the agent's output too. Two more details matter in production. Turn events from subagents carry a subagent_id and must not end the loop. And when a stream drops, missed events are gone: you recover by opening a new stream, retrieving the session and its saved items, and merging by item_id.
There is no separate steer endpoint. The same agent.session.input.message event starts a new turn when the session is idle and steers the active turn when the agent is working, so a correction such as excluding a test branch lands while the agent is still mid-task. Cancelling is a different input event on the same endpoint, and it stops the turn without discarding the session or its earlier work.
// One event type does two jobs. Sent to an idle session it starts a new turn;
// sent while a turn is running it steers that turn.
async function sendMessage(client: OpenAI, sessionId: string, text: string) {
await client.beta.agents.sessions.events.create(sessionId, {
events: [
{
type: "agent.session.input.message",
input: [{ role: "user", content: [{ type: "input_text", text }] }],
},
],
});
}
// Mid-turn correction: the agent is already working on the wrong branch set.
await sendMessage(client, session.id, "Exclude branch JKT-99, it is a test branch.");
// Stop the current turn. The session and its previous work survive.
await client.beta.agents.sessions.events.create(session.id, {
events: [{ type: "agent.session.input.cancel" }],
});Some settings can change on a live session and some cannot. A POST to the session itself can change the model, reasoning effort or service tier for turns started after the update, while an active turn keeps its settings even when you steer it. Instructions, tools and multi_agent are fixed for the life of the session, so changing what the agent may do means creating a new session. That is a useful property when an auditor asks what an agent was allowed to do on a given run.
# Change model, reasoning effort or service tier for LATER turns of one session.
# instructions, tools and multi_agent cannot change here: create a new session.
curl -sS -X POST "https://api.openai.com/v1/agents/sessions/$SESSION_ID" \
-H "OpenAI-Beta: agents=v1" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{ "agent": { "reasoning": { "effort": "low" }, "service_tier": null } }'A month-end job should not hold an HTTP stream open for as long as the agent needs. Session webhooks cover the state changes, namely agent.session.created, action_required, in_progress, idle and failed. They are signed, so the handler verifies the signature against your webhook secret before parsing anything. The interesting one is action_required, which fires when the session needs a function result, an environment connection or a computer-use approval.
import express from "express";
import OpenAI from "openai";
const app = express();
const client = new OpenAI({ webhookSecret: process.env.OPENAI_WEBHOOK_SECRET });
app.post("/webhooks/openai", express.raw({ type: "application/json" }), async (req, res) => {
const payload = req.body.toString("utf8");
try {
await client.webhooks.verifySignature(payload, req.headers);
} catch {
return res.status(400).send("Invalid signature");
}
const event = JSON.parse(payload);
res.sendStatus(200); // acknowledge fast; queue the slow part in production
// Webhook name: agent.session.action_required.
// The stream event for the same situation is agent.session.requires_action.
if (event.type !== "agent.session.action_required") return;
// The webhook omits call IDs and arguments — read them from the session.
const session = await client.beta.agents.sessions.retrieve(event.data.id);
for (const action of session.required_actions ?? []) {
if (action.type !== "function_call" || action.name !== "get_stock_level") continue;
const output = JSON.stringify(await getStockLevel(action.arguments));
await client.beta.agents.sessions.events.create(session.id, {
events: [{
type: "agent.session.input.tool_result",
turn_id: action.turn_id,
call_id: action.call_id,
success: true,
output, // a string; serialize objects yourself
}],
});
}
});
app.listen(Number(process.env.PORT ?? 8000));Two details in that handler come straight from the docs and are easy to miss. The webhook deliberately omits call IDs and arguments, so you retrieve the session and read required_actions rather than trusting the payload. And the webhook is named agent.session.action_required while the matching stream event is agent.session.requires_action, a one-word difference that silently breaks a handler written against the wrong page. For functions with side effects, such as posting a journal entry, store each result by session, turn and call ID, and resubmit the saved result after a disconnect instead of running the function twice.
The environment type decides who provisions compute and how you get files back. Only OpenAI-hosted sandboxes publish files written to /workspace/outputs as downloadable artifacts. A self-hosted environment returns files through your own provider, even from that same path.
| environment.type | Who provisions it | Getting files out | Use it when |
|---|---|---|---|
| none | Nobody, there is no sandbox | Read the output from session items | The agent only answers questions or calls your functions and external tools |
| openai_hosted | OpenAI, using the size, packages, files and network policy you set | Artifacts API for files under /workspace/outputs, still downloadable after the sandbox expires | You want a disposable Linux workspace without running any infrastructure |
| self_hosted | You, on a laptop, a container or a provider such as Modal, E2B or Daytona | Your provider's file API or a mounted filesystem | You need your own image, your own compute or a private network |
Self-hosting works by running the Codex executor, codex exec-server, inside your environment. It registers with an environment ID and a restricted environment key, then holds an outbound WebSocket to receive commands, so no inbound port is opened. The environment key can only connect environments and cannot authorise any other API action, which is why it is the only OpenAI credential that should exist inside the sandbox.
# Inside YOUR sandbox (a container, a Modal/E2B/Daytona box, a laptop).
# CODEX_API_KEY is the restricted environment key: it can only connect environments.
# The application OPENAI_API_KEY never enters this machine.
export CODEX_API_KEY="$OPENAI_EXECUTOR_API_KEY"
# Outbound only: api.openai.com (register) + wss://codex-cloud-environments.chatgpt.com
codex exec-server \
--remote "<session.environment.remote_url>" \
--environment-id "<session.environment.id>"Sandbox options at launch: besides OpenAI-hosted sandboxes, the announcement lists first-class integrations with Blaxel, Cloudflare, Daytona, DigitalOcean, E2B, Modal, Oracle Cloud, Runloop and Vercel, and the self-hosted guide also documents AWS Lambda MicroVMs. OpenAI-hosted sandboxes offer configurable CPU, GPU and memory, either fully managed or deployed into your own VPC.
The Agents API itself carries no additional fee; the announcement says you pay for the tokens and tools you consume. On OpenAI's pricing page, gpt-6-astra, the model in every Agents API example, lists standard short-context rates of 10 USD per million input tokens, 1 USD per million cached input tokens and 50 USD per million output tokens. The containers line lists 0.03 USD for 1 GB, 0.12 USD for 4 GB and 0.48 USD for 16 GB per 20-minute session, with eligible sessions billed by the minute and a five-minute minimum. Those are the same memory tiers as the small, medium and large sandbox sizes.
Cleanup has an order. Artifacts are immutable copies published when a turn completes, and the API downloads one per request with no batch endpoint, so either loop over them or ask the agent to zip its outputs into one file. Download what you need first, then delete the session. A 409 on deletion means setup or execution is still finishing, so retry with a bounded number of attempts.
# Python — copy outputs out BEFORE deleting the session.
def download_artifact(client, session_id, turn_id, path, destination):
for artifact in client.beta.agents.sessions.artifacts.list(session_id):
if artifact.turn_id != turn_id or artifact.path != path:
continue
with client.beta.agents.sessions.artifacts.with_streaming_response.content(
artifact.id, session_id=session_id
) as response:
response.stream_to_file(destination)
return
raise FileNotFoundError(f"No artifact for {path!r} in turn {turn_id}")
download_artifact(client, session_id, turn_id, "/workspace/outputs/totals.json", "totals.json")
# Then delete. A 409 means setup or execution is still finishing:
# wait and retry a bounded number of times.
client.beta.agents.sessions.delete(session_id)Do not lean on expiry as your cleanup mechanism. A connected sandbox receives keep-alives, including between turns, and only becomes eligible for deletion after an hour without activity or keep-alives, a timeout you cannot configure. Closing your event stream does not cancel the task either. For self-hosted sessions the warning is sharper: deleting a session emits no webhook and does not stop provider compute, so your provisioning code owns the shutdown.
The overview states both limits plainly. The Agents API currently supports data residency only in the United States, and it does not support Zero Data Retention. Sessions retain state by design, which is what makes follow-ups work, and you can delete sessions and published artifacts when you are done, but deletion is a cleanup step rather than a retention guarantee.
Choosing a self-hosted sandbox does not make the Agents API ZDR-eligible. Running the executor on your own hardware keeps files and commands on your machine, but the harness, the conversation and the session state still live with OpenAI in the US. If a client contract requires in-region processing or zero retention, the Agents API in its current beta is the wrong runtime, whichever sandbox you pick.
For my month-end example, that splits the work cleanly. Aggregated, already-anonymised exports are a reasonable pilot for a hosted session with networking disabled. Raw ledgers containing customer names, or anything an Indonesian client has contractually kept in-region, stay on a runtime where I control the loop and the storage. Secrets follow the same logic: the hosted-sandbox guide warns that agent-generated code can read env values, and points to vault credentials for anything sensitive.
The Agents API removes the part of agent engineering nobody wanted to own, namely the loop, compaction, recovery and subagent plumbing, and leaves you the parts that carry risk. Treat it as a session protocol: open the stream before sending work, judge success by turn outcomes and output rather than idle, keep the application key out of the sandbox, download artifacts before deleting, and check residency and retention before the first real record goes in.