AI
OpenAI Agents API vs Agents SDK vs Responses API: Which?
October 202611 min read

The Agents API is a hosted service: OpenAI runs a managed Codex harness and keeps the session, turns and items on its side. The Agents SDK is a Python or TypeScript library whose runner executes the agent loop inside your own application, so you control deployment, storage and approvals.
No. OpenAI documents that the Agents API does not support Zero Data Retention, and that choosing a self-hosted sandbox does not make it eligible. If you need ZDR, use the Agents SDK or the Responses API and keep conversation history in your own storage.
Not with the Agents API, which currently supports data residency only in the United States. The Responses API is listed for regional storage in ten regions, including Singapore, the nearest one to Indonesia. Residency is configured per project and carries a 10% uplift for eligible models released on or after 5 March 2026.
Use it when the job is one model call with a schema, such as extracting a purchase order or classifying a ticket. An agent runtime adds sessions, turn limits and state to clean up without adding value there. For genuine multi-step agents with custom tools, the Agents SDK saves you writing the loop yourself.
No. ChatKit is an embeddable chat interface, not an agent runtime. OpenAI recommends connecting it to your own agentic service through the ChatKit Python SDK, which pairs naturally with an Agents SDK backend. The Agent Builder-hosted ChatKit path is legacy because Agent Builder shuts down on 30 November 2026.

Key Takeaway
OpenAI documents four ways to build an agent. The Agents API runs a managed Codex harness for you but supports only US data residency and no Zero Data Retention. The Agents SDK runs the loop inside your own process with your own storage. The Responses API is raw model calls. ChatKit is only an embeddable chat UI.
Picture an accounts-payable assistant for an ERP team: it reads open invoices, explains mismatches to a clerk and never posts anything on its own. OpenAI now offers four starting points for building it, and the first question a finance client asks is rarely which model. It is where the invoice data will sit. That one question can eliminate an option before any code is written.
This guide compares the Agents API, the Agents SDK, the raw Responses API and ChatKit on the axes that actually decide the choice: who owns the agent loop, where state lives, Zero Data Retention eligibility, data residency, cost and lock-in. Every capability claim comes from OpenAI's own documentation as of October 2026, linked at the end, and each option is mapped to a concrete job such as an ERP back-office run, a customer chat or a CI task.
Every agent is a loop: call the model, run the tools it asks for, feed the results back, decide whether to stop. The four options OpenAI documents differ mainly in who writes and hosts that loop, and the rest of the comparison follows from that one fact.
OpenAI's own comparison rates the integration effort as low for the Agents API, medium for the SDK and high for the Responses API. That rating is honest, and it is also the trade: the less you integrate, the less you control where data goes and how the loop behaves when something fails.
The table below combines OpenAI's runtime comparison with the data-controls page, which is where the ZDR and residency rows come from. Those two rows are the ones most comparison posts leave out, and for regulated ERP data they are usually the ones that decide the case.
| Dimension | Agents API | Agents SDK | Responses API | ChatKit |
|---|---|---|---|---|
| Who owns the loop | OpenAI, through the managed Codex harness | The SDK runner, inside your process | You write it from scratch | Nobody: it is a UI over a backend you choose |
| State between tasks | Saved session configuration, turns and items on OpenAI | Your storage via SDK sessions, or Responses conversation state | Manual history, response chaining or Conversations | Whatever the backend behind it keeps |
| Execution environment | OpenAI-hosted sandbox, self-hosted sandbox or none | Your runtime plus sandbox provider integrations | Your own execution environment | Browser front end plus your server |
| Integration effort, as OpenAI rates it | Low | Medium | High | Not rated: it is a UI layer |
| Zero Data Retention | Not supported, even with a self-hosted sandbox | Follows the endpoints it calls: Responses is eligible with limitations, Conversations-backed state is not | /v1/responses eligible with limitations, store forced to false; /v1/conversations not eligible | Follows the backend it talks to |
| Data residency | United States only | Follows the endpoints it calls, so regional storage is possible | /v1/responses listed for regional storage in ten regions, including Singapore | Follows the backend it talks to |
| Cost beyond model tokens | No platform fee; OpenAI tools at standard rates, hosted sandboxes at container rates | Your compute, your storage, your sandbox provider's bill | Your compute, plus every hosted tool you enable | Hosting the server it connects to |
| Lock-in | Highest: session objects and harness behaviour live on OpenAI | Moderate: you own the code, and manual history works with any provider | Lowest at the API level, but you rebuild everything the others give you | Low: the UI can point at another backend |
| Use it for | Long-running tasks where OpenAI manages the agent and saves its progress | Custom tools, approvals and workflows inside your application | Direct model control or an agent built from scratch | Adding an embedded chat experience |
Two rows carry most of the weight. The ZDR row is binary for the Agents API: OpenAI states plainly that it does not support Zero Data Retention and that choosing a self-hosted sandbox does not make it eligible. The residency row is the same story. The other three options inherit the controls of whichever endpoints they call, which is why their cells say follows rather than yes or no.
The Agents API is built on four objects: the agent (model, instructions, tools, MCP servers), an optional environment, a durable session, and the events and items that flow in and out. Turns run asynchronously. A message sent to an idle session starts a new turn, and the same message type sent during an active turn steers it instead. The whole lifecycle fits in one block, written against the beta SDK surface the docs use.
import OpenAI from "openai";
const client = new OpenAI();
// 1. Create the session AND start the first turn in one request.
// OpenAI provisions the sandbox; you never see the agent loop.
const events = await client.beta.agents.sessions.create({
agent: {
model: "gpt-6-astra",
instructions:
"Reconcile yesterday's goods receipts against open purchase orders. " +
"Report mismatches; never post anything.",
},
environment: { type: "openai_hosted" },
input: "Run the reconciliation for warehouse JKT-01.",
stream: true,
});
for await (const event of events) {
// An idle session is NOT success. Wait for the turn outcome:
// agent.session.turn.completed | .failed | .cancelled
console.log(event.type);
}
// 2. The same call steers a running turn or starts a new one on an idle
// session. Persist the session id next to your own job record.
const sessionId = job.agentSessionId; // saved from the session's events
await client.beta.agents.sessions.events.create(sessionId, {
events: [
{
type: "agent.session.input.message",
input: [
{
role: "user",
content: [{ type: "input_text", text: "Skip JKT-02, it is mid-stocktake." }],
},
],
},
],
});
// 3. Stop the turn; the session and its earlier work survive.
await client.beta.agents.sessions.events.create(sessionId, {
events: [{ type: "agent.session.input.cancel" }],
});Three details in that code are easy to miss. First, an idle session does not mean the turn succeeded; the docs say to check for the completed, failed or cancelled turn event and to inspect the output, because a completed turn does not guarantee every tool call worked. Second, function tools still run in your code: when the session needs a function result it surfaces required_actions, and nothing continues until you answer. Third, the runtime accepts requests up to 4 MiB, input plus output schema, so you upload large documents to the environment rather than inlining them.
Read the data-controls line before the feature list. The Agents API currently supports data residency only in the United States and does not support Zero Data Retention, and a self-hosted sandbox does not change either fact. Session state is retained so work can continue across turns until you delete the session. If a client contract says invoice or payroll data must not leave a region, this option is out regardless of how convenient the harness is.
The Agents SDK is a library, available in Python and TypeScript, whose runner handles the loop and handoffs inside your application. The part that matters for this comparison is state. The Python docs list four strategies: rebuild the input yourself with to_input_list, use an SDK session backed by your own store, point at a server-side conversation_id, or chain with previous_response_id. Only the first two keep the conversation off OpenAI's servers.
from agents import Agent, Runner, SQLiteSession, function_tool
@function_tool
def open_invoices(vendor_code: str) -> list[dict]:
"""Read-only lookup against the ERP database you already run."""
return erp.query_open_invoices(vendor_code)
agent = Agent(
name="AP clerk",
instructions="Answer accounts-payable questions. Never approve payments.",
tools=[open_invoices],
)
# Option A: history lives in YOUR storage (SQLite here; swap the class for
# your own store). Works with any provider the SDK can call.
session = SQLiteSession("vendor-V0042")
result = await Runner.run(agent, "What is open for V0042?", session=session, max_turns=8)
result = await Runner.run(agent, "Which of those are overdue?", session=session, max_turns=8)
# Option B: history lives on OpenAI as a Conversation object.
conversation = await client.conversations.create()
result = await Runner.run(agent, "What is open for V0042?", conversation_id=conversation.id)
# Wrong: session= together with conversation_id / previous_response_id.
# The SDK does not allow both state strategies in the same run.Two rules from the running-agents docs saved me a debugging session. A session cannot be combined with conversation_id, previous_response_id or auto_previous_response_id in the same run, so pick one strategy per agent and keep it. And max_turns is a real safety net: a run that exceeds it raises MaxTurnsExceeded rather than looping until the bill notices. Set it explicitly on every back-office job; passing None disables the limit, which is exactly what an unattended ERP job should never do.
OpenAI rates the Responses API as the high-effort option, and for a genuine multi-step agent it is. But a lot of work called agentic is one model call with a schema: extract a purchase order, classify a ticket, summarise a document. For those jobs an agent runtime adds sessions, turn limits and state you then have to clean up, and buys nothing.
from openai import OpenAI
client = OpenAI()
# You own everything: the loop, the tool dispatch, the history list.
history = [{"role": "user", "content": "Summarise PO-2026-0917 for the buyer."}]
response = client.responses.create(
model="gpt-6-astra",
input=history,
tools=TOOLS, # your function definitions
store=False, # forced to False anyway under ZDR
)
# Not stored server-side, so the next turn cannot lean on OpenAI state:
# append response.output plus your tool results to history and call again,
# and decide yourself when to stop.The data controls are also the clearest here. By default, Responses keeps application state for 30 days; under Zero Data Retention the store parameter is always treated as false, even if a request sets it to true. Conversations, by contrast, are kept until you delete them and are not ZDR-eligible. If you need ZDR and server-side history at the same time, you cannot have both from OpenAI, so the history has to live in your own database.
There is no Indonesia region in OpenAI's data-residency table. Singapore is the nearest, and /v1/responses is listed for regional storage there, but not regional processing, and the region requires Modified Abuse Monitoring or ZDR. Residency is configured per project through OpenAI's sales team, and it carries a 10% price uplift for eligible models released on or after 5 March 2026. Budget for that before promising it to a client.
ChatKit sits in the same starting-point table, which invites the wrong comparison. It is an embeddable chat interface with tool invocation, file attachments and visualisations, available as React bindings, a JS SDK and a Python server SDK. It answers how users talk to an agent, not how the agent runs. In practice it pairs with one of the other three, most naturally the Agents SDK behind your own server.
The deployment advice changed this year. The ChatKit docs now recommend connecting it to your own agentic service through the Python SDK, and treat the Agent Builder-hosted integration as a legacy path, because Agent Builder is scheduled to shut down on 30 November 2026. If a tutorial tells you to publish a workflow in Agent Builder and embed it with ChatKit, it is describing the path that is ending.
Abstract comparisons only go so far. These are the four jobs I would actually put in front of a team, with the option I would pick and the deciding reason.
| Job | Pick | Deciding reason |
|---|---|---|
| Nightly ERP back-office run: reconcile goods receipts against purchase orders and flag mismatches | Agents SDK on your own worker | Financial data, read-only tools against your own database, approvals in your code, and history in your storage; the Agents API's US-only, no-ZDR terms rule it out for many clients |
| Customer-facing order-status chat on a company website | ChatKit front end with an Agents SDK backend | ChatKit gives you the UI, the SDK keeps tools and auth on your server, and nothing depends on Agent Builder |
| CI task: reproduce a failing test, investigate, draft a fix | Agents API with an OpenAI-hosted sandbox | Long-running, needs a shell and files, benefits from managed compaction, recovery and subagents, and the repository is usually not regulated data |
| High-volume extraction: invoice PDF to structured JSON | Responses API | One call with a schema, no loop, nothing to clean up afterwards, and ZDR-eligible if the organisation has it |
The CI row is the one where the Agents API earns its place. Running a sandboxed shell, compacting context on a long investigation and recovering from a crashed worker are exactly the parts I would rather not build, and source code for an internal tool is a far easier data conversation than a vendor ledger. The ERP row goes the other way for the same reason.
OpenAI's docs end the comparison with a warning worth taking literally: an Agents API session, an SDK session, a Responses conversation and a sandbox are different resources, each with its own state and cleanup rules. Moving between them later is a migration, not a config change. This is the order I make the decision in.
Lock-in is not automatically bad. The Agents API's lock-in buys a harness that OpenAI maintains; the SDK's freedom costs you the operations work. What hurts is choosing one by accident and discovering the data-controls row after the client asks.
The rule I now carry is short: decide where the data may live, then decide who runs the loop, and only then look at features. For most ERP back-office work that lands on the Agents SDK with state in your own storage; for sandboxed engineering tasks the Agents API is the least code; for single-shot extraction plain Responses is enough; and ChatKit is the face you put on whichever one you chose.
Sources