AI
Claude Managed Agents vs OpenAI Agents API: Full Comparison
October 202612 min read

Both are hosted agent harnesses built on an agent, an environment, a session and events, and both can run tools in a vendor-managed or self-hosted sandbox. Claude treats agents and environments as reusable, versioned resources and offers scheduled deployments on a cron. OpenAI lets one request define the whole session, supports an environment type of none, and runs the managed Codex harness.
No. Anthropic states that Claude Managed Agents is not eligible for Zero Data Retention or HIPAA BAA coverage because sessions store history, sandbox state and outputs server-side. OpenAI states that the Agents API does not support Zero Data Retention and that a self-hosted sandbox does not make it eligible. Both let you delete sessions when the work is done.
Claude Managed Agents bills tokens at model rates plus $0.08 per session-hour, counted only while the session is running. OpenAI lists no separate agent fee: tokens at model rates, its tools at standard rates and hosted sandboxes at container rates, which are $0.03, $0.12 and $0.48 per 20-minute session for 1 GB, 4 GB and 16 GB.
Claude Managed Agents has scheduled deployments that start a session from a POSIX cron expression in an IANA timezone, with up to 1,000 per organisation. Runs get up to 15 percent jitter, capped at 9 minutes. The OpenAI Agents API documents no scheduler, so you trigger sessions from your own cron or queue and follow them with webhooks.
Only partly. A self-hosted sandbox keeps the files, processes and network egress of tool execution on your infrastructure, which helps when the agent must reach private systems. The agent loop still runs on the vendor's harness, so prompts, commands and tool results still reach Anthropic or OpenAI. If data must never leave your servers, run an agent library in your own process instead.

Key Takeaway
Claude Managed Agents and the OpenAI Agents API are hosted agent harnesses built on the same four objects: agent, environment, session and events, with managed or self-hosted sandboxes. Neither is eligible for Zero Data Retention. Claude adds cron scheduling and bills $0.08 per running session-hour; OpenAI adds a no-sandbox mode, US-only residency and container-rate sandboxes.
Consider a team running a home-grown agent loop for an ERP back office in Jakarta. A Python worker calls the model, executes tools in a Docker container, trims the history when it grows too long, and restarts from a checkpoint when the VPS reboots. Every one of those pieces is code the team maintains. In 2026 both Anthropic and OpenAI began offering to run that loop for them, Anthropic with Claude Managed Agents in April and OpenAI with the Agents API in September, and the question on the table became whether to delete the worker.
This Claude Managed Agents vs OpenAI Agents API comparison is written for exactly that decision. It covers the shared object model, a full feature table, the same task written against both SDKs, custom tools, sandboxes, scheduling, data controls and cost. Every claim comes from the two vendors' own documentation as read on 1 October 2026. Both products are in beta, so check the linked pages before committing.
The two products converged on almost the same nouns. Both describe an agent, an environment, a session and the events that flow between your application and the running agent. The table maps them one to one; the differences sit in the details, not in the concepts.
| Concept | Claude Managed Agents | OpenAI Agents API |
|---|---|---|
| Agent | A versioned resource holding the model, system prompt, tools, MCP servers and skills. Sessions reference it by ID or pin a version. | Model, instructions, tools and MCP servers. Pass it inline when creating a session, or save it once and pass agent_id. |
| Environment | Created separately and required by every session: an Anthropic-managed cloud sandbox or a self-hosted one. | Optional and declared per session: openai_hosted, self_hosted, or none for an agent that only calls tools. |
| Session | A running agent instance inside an environment, with a persistent filesystem and history. Created first, then started by a user event or initial_events. | A durable instance that works on tasks. The create request can carry the first input and stream the first turn. |
| Events | In: user.message, user.interrupt, user.tool_confirmation, user.custom_tool_result. Out: agent.message, agent.tool_use, session.status_idle. | In: agent.session.input.message, input.cancel, input.tool_result. Out: turn outcomes such as agent.session.turn.completed. |
| Transport | Server-sent events, with the full event history persisted server-side and fetchable later. | Server-sent events on the session stream, or signed webhooks for state changes without an open stream. |
Two differences matter in practice. Claude treats the agent and the environment as long-lived resources that you create once and version, which suits a fleet of sessions sharing one reviewed configuration. OpenAI lets a single request carry everything and adds an environment type of none, which suits an agent that only talks to remote MCP servers and never needs a shell.
This is the table to put in front of whoever signs off on the architecture. Where a cell says not documented, the vendor's pages say nothing either way as of October 2026, which is not the same as saying no.
| Dimension | Claude Managed Agents | OpenAI Agents API |
|---|---|---|
| Launch and status | Public beta since 8 April 2026 | Public beta since 10 September 2026 |
| Beta flag | anthropic-beta: managed-agents-2026-04-01 on every endpoint; the SDK sets it for you | OpenAI-Beta: agents=v1 header; SDK methods live under client.beta.agents |
| Harness | Anthropic's agent harness with built-in prompt caching and compaction | The managed Codex harness with compaction, recovery and resume |
| Built-in tools | Bash, file operations (read, write, edit, glob, grep), web search and web fetch | Sandbox commands and file editing, web search, programmatic tool calling, and computer use since 29 September 2026 |
| Your own tools | Custom tools with an input_schema; results return as user.custom_tool_result | Function tools with JSON Schema parameters; results return as agent.session.input.tool_result |
| MCP | MCP servers on the agent; MCP tunnels for private servers are a limited research preview | MCP servers over HTTP transport, called by the harness directly |
| Sandbox options | Anthropic-managed cloud sandbox or self-hosted sandbox; an environment is always required | OpenAI-hosted sandbox, self-hosted sandbox, or no environment at all |
| Hosted sandbox | Secure sandboxes with pre-installed packages and network access; sizing not documented on the overview | Linux workspace in three sizes up to 4 vCPU and 16 GB; network enabled, disabled or allowlisted |
| Multi-agent | Multi-agent coordination is a research preview, available on request | Subagents through multi_agent, with a max_concurrent_subagents limit |
| Scheduling | Scheduled deployments: POSIX cron plus an IANA timezone, up to 1,000 per organisation | No scheduler documented; trigger sessions from your own cron or queue |
| Retention | History, sandbox state and outputs stored server-side until you delete the session | Session state retained across turns until you delete the session |
| ZDR and HIPAA | Not eligible for Zero Data Retention or HIPAA BAA coverage | No Zero Data Retention, and a self-hosted sandbox does not change that; BAA not documented on the overview |
| Data residency | Inference can be pinned to the US with inference_geo, billed at 1.1x token rates | United States only |
| Pricing | Tokens at model rates plus $0.08 per session-hour while running; no separate container-hours | Tokens at model rates, OpenAI tools at standard rates, hosted sandboxes at container rates; no separate agent fee listed |
Read the ZDR and residency rows before the feature rows. Both vendors made the same trade: a stateful hosted harness keeps your transcripts on their servers, so neither offers Zero Data Retention for it. The feature rows are close enough that most teams will decide on data controls, on scheduling, and on which model family already passes their evaluations.
The quickest way to feel the difference is to write one job twice. The block below asks each platform to reconcile a stock export against a general-ledger extract for one warehouse, in Python, using only the methods shown in each vendor's documentation. The model names are the ones those docs use in October 2026.
# --- Claude Managed Agents -------------------------------------------
from anthropic import Anthropic
claude = Anthropic() # the SDK adds the managed-agents-2026-04-01 beta header
# Agent and environment are separate, reusable resources. Create them once,
# keep the IDs in config, and reference them from every session.
agent = claude.beta.agents.create(
name="Stock reconciler",
model="claude-opus-5-5",
system="Reconcile the stock export against the GL extract. Report, never post.",
tools=[{"type": "agent_toolset_20260401"}], # bash, files, web search/fetch
)
environment = claude.beta.environments.create(
name="recon-env",
config={"type": "cloud",
"networking": {"type": "limited", "allow_package_managers": True}},
)
session = claude.beta.sessions.create(agent=agent.id, environment_id=environment.id)
# Open the stream BEFORE sending the event, so nothing is emitted unseen.
with claude.beta.sessions.events.stream(session.id) as stream:
claude.beta.sessions.events.send(session.id, events=[{
"type": "user.message",
"content": [{"type": "text", "text": "Run the reconciliation for JKT-01."}],
}])
for event in stream:
if event.type == "session.status_idle":
break # check event.stop_reason: end_turn OR requires_action
# --- OpenAI Agents API -----------------------------------------------
from openai import OpenAI
oai = OpenAI()
# Agent and environment can be inline; OpenAI provisions the sandbox.
# To reuse settings, save one with oai.beta.agents.create and pass agent_id.
with oai.beta.agents.sessions.create(
agent={"model": "gpt-6-astra",
"instructions": "Reconcile the stock export against the GL extract. "
"Report, never post."},
environment={"type": "openai_hosted", "network": {"access": "disabled"}},
input="Run the reconciliation for JKT-01.",
stream=True,
) as events:
for event in events:
# The turn outcome is the signal. agent.session.idle is not success.
if event.type in ("agent.session.turn.completed",
"agent.session.turn.failed",
"agent.session.turn.cancelled"):
breakClaude needs three calls before any work starts, because agent, environment and session are separate resources. In exchange the agent ID is reusable and versioned, so a code review can approve one configuration for every session that follows. OpenAI fits in a single call. The stop conditions also differ: on Claude, session.status_idle with a stop_reason of end_turn means the agent has nothing left to do, while OpenAI reports an explicit turn outcome, completed, failed or cancelled, and that outcome is what a job runner should wait for.
Built-in tools run in the sandbox, but anything that touches the ERP should be your own code: a read-only stock lookup, a draft journal, a vendor search. Both platforms implement this the same way. The agent asks for a call, the session pauses, your application runs the function and posts the result back as an event.
# Claude: tool declared on the agent as {"type": "custom", "input_schema": ...}.
# The session emits agent.custom_tool_use, then idles with requires_action.
if event.type == "session.status_idle" and event.stop_reason.type == "requires_action":
for event_id in event.stop_reason.event_ids:
call = seen_tool_uses[event_id] # the agent.custom_tool_use event
claude.beta.sessions.events.send(session.id, events=[{
"type": "user.custom_tool_result",
"custom_tool_use_id": event_id,
"content": [{"type": "text", "text": run_erp_lookup(call.name, call.input)}],
}])
# OpenAI: tool declared as {"type": "function", "parameters": ...}.
# The stream emits agent.session.requires_action; the session holds the details.
pending = oai.beta.agents.sessions.retrieve(session_id).required_actions
for action in pending:
if action.type != "function_call":
continue # environment_connection, etc.
oai.beta.agents.sessions.events.create(session_id, events=[{
"type": "agent.session.input.tool_result",
"turn_id": action.turn_id,
"call_id": action.call_id,
"success": True,
"output": json.dumps(run_erp_lookup(action.name, action.arguments)),
}])
# Both: persist (session, call id) -> result BEFORE you send it. After a crash
# you resend the saved result instead of running a side effect twice.The shapes differ in one useful way. Claude puts the blocking event IDs directly on the idle event, so a streaming handler already holds everything it needs. OpenAI's requires_action event only tells you to look, and the retrieved session's required_actions list tells you what to do. That second design is easier to recover after a restart, because the source of truth is one API read rather than a stream you may have missed.
A paused session is a session waiting on your infrastructure. If the worker that answers tool calls is down, the agent simply waits, and on Claude the runtime meter stops while the session is idle, so the bill will not warn you either. Track pending tool calls as a metric, and make every tool that writes idempotent by saving its result against the session and call ID before sending it back.
Both platforms let the sandbox run on your own machines, and both are explicit that this moves execution, not the agent loop. The harness and the model stay with the vendor.
The consequence is about what you can promise a client. A self-hosted sandbox keeps the files the agent works on, the processes it spawns and its network egress on your side, which is real value when the agent needs a database that is not publicly routable. It does not keep the conversation on your side: commands and tool results still travel to the vendor's harness so the model can decide what to do next. If the requirement is that data never leaves your servers, neither product meets it, and an agent library running in your own process is the honest answer.
Back-office agents are usually scheduled, not chatted with: a nightly reconciliation, a Monday aged-debt summary, a month-end accrual check. Claude Managed Agents answers this with scheduled deployments, which start a session from a cron expression in a named timezone. The OpenAI Agents API documentation covers sessions, streaming and webhooks but no scheduler, so the trigger stays in your own infrastructure.
# Claude scheduled deployment: the platform starts a session on a cron.
deployment = claude.beta.deployments.create(
name="Nightly stock reconciliation",
agent=agent.id,
environment_id=environment.id,
initial_events=[{ # at least one event is required
"type": "user.message",
"content": [{"type": "text",
"text": "Reconcile yesterday's movements for every warehouse."}],
}],
# Asia/Jakarta has no DST, so the skipped/doubled 2 AM trap never applies.
schedule={"type": "cron", "expression": "30 4 * * *", "timezone": "Asia/Jakarta"},
)
print(deployment.schedule.upcoming_runs_at) # confirm the fire times
# A trigger that fails creates a run record, not a session. Alert on these.
for run in claude.beta.deployment_runs.list(deployment_id=deployment.id, has_error=True):
print(run.created_at, run.error.type) # e.g. session_rate_limited_error
# OpenAI Agents API: no scheduler is documented, so the cron is yours,
# e.g. a worker that calls oai.beta.agents.sessions.create(...) at 04:30 WIB
# and listens for agent.session.idle / agent.session.failed webhooks.Three details from the scheduled deployments page change how you operate it. Execution adds jitter of up to 15 percent of the interval, at least 5 seconds and at most 9 minutes, so a 04:30 job can start at 04:39. A failed trigger, such as a rate-limited session creation, is recorded as a deployment run with an error type and is not retried until the next occurrence. And an optional budget is copied onto every session the deployment starts, capping each run on its own rather than the total across runs.
Both vendors say nearly the same thing. Claude Managed Agents is stateful by design, storing conversation history, sandbox state and outputs server-side, and is therefore not eligible for Zero Data Retention or HIPAA Business Associate Agreement coverage. The OpenAI Agents API retains session state to continue work across turns, does not support Zero Data Retention, and states that choosing a self-hosted sandbox does not make it eligible. Both let you delete sessions, and Claude also lets you delete uploaded files separately.
Residency is where they differ, and only slightly. The OpenAI Agents API supports data residency only in the United States. On Claude, an agent's model configuration can pin inference_geo to us, which bills that agent's tokens at 1.1 times the standard rates. Neither product's documentation lists a Southeast Asian region.
For an Indonesian client, settle the data question before the demo. If the agent will see personal data or financial records covered by a residency or retention clause, both hosted harnesses keep transcripts outside Indonesia until you delete them. Design for that on purpose: give the agent IDs and aggregates instead of raw rows, keep sensitive lookups behind your own custom tools, and delete sessions on a schedule once the output is saved.
Claude's pricing is the easier one to forecast. Tokens are billed at model rates, and session runtime costs $0.08 per session-hour, metered to the millisecond only while the session is running and never while it sits idle waiting for input or a tool confirmation. It replaces container-hour billing rather than adding to it; the pricing page's own example, a one-hour Claude Opus 5 session with 50,000 input and 15,000 output tokens, totals $0.705. OpenAI lists no separate agent fee: tokens at model rates, tools at standard rates, and hosted sandboxes at container rates, which its pricing page sets at $0.03, $0.12 and $0.48 per 20-minute session for 1 GB, 4 GB and 16 GB, billed by the minute with a 5-minute minimum. With the cost shape clear, four questions settle the choice.
Whichever you pick, keep your tools as plain functions behind a thin adapter. The custom-tool contracts above are close enough that porting a tool takes an afternoon, while a session history cannot be moved between the two vendors at all.
The rule worth carrying is that a hosted agent harness is a data decision first and a feature decision second. Claude Managed Agents and the OpenAI Agents API remove the same work, the loop, the sandbox, compaction and recovery, and both keep your transcripts until you delete them. If that is acceptable, choose on model quality, scheduling and pricing shape. If it is not, keep the loop in your own process.
Sources