AI
OpenAI Agents SDK vs Google ADK vs Claude Agent SDK (2026)
October 202611 min read

The OpenAI Agents SDK is a small set of primitives, agents, a Runner, handoffs, guardrails, sessions and sandbox agents, in Python and TypeScript. Google ADK 2.0 adds graph workflows where agents and functions are nodes joined by edges, built-in evaluation with adk eval, and SDKs for Python, Go and TypeScript. ADK is also more model-agnostic out of the box.
Yes, through its LiteLLM or Any-LLM adapters, which the SDK documentation marks as beta. You install the extra, for example openai-agents with the litellm extra, and pass a model name with the litellm prefix. Structured outputs, tool calling and usage metrics then depend on the upstream provider, so test them before relying on them.
No, but it is built on the Claude Code harness, so it is strongest where an agent reads files, edits them or runs shell commands. You get Claude Code's built-in tools, permissions, hooks, subagents and sessions without writing them. For a narrow agent that only calls a few APIs, it is a heavier runtime than the job needs.
Google ADK, by a clear margin. It ships the adk eval command, an adk web interface and a pytest AgentEvaluator, with metrics for the exact tool-call trajectory, ROUGE-1 response match and LLM-judged quality. Neither the OpenAI Agents SDK nor the Claude Agent SDK lists an eval runner among its primitives.
With the SDKs, yes: the code runs in your own process or container. Each vendor also has a hosted option: the OpenAI Agents API, Claude Managed Agents, and Google Cloud's Agent Runtime for ADK. If you do not want to operate agent servers, compare those hosted runtimes before choosing the SDK.

Key Takeaway
The OpenAI Agents SDK is the lightest framework, with handoffs, guardrails and sandbox agents in Python and TypeScript. Google ADK 2.0 adds graph workflows, built-in evaluation and a Go SDK. The Claude Agent SDK embeds the Claude Code harness with its file and shell tools. Choose by loop shape and production hosting target, not by model vendor alone.
The week OpenAI's Agents API went into beta, I had the documentation of three vendor agent SDKs open side by side, trying to decide which one my next agent should be built on. Each vendor's docs answered a slightly different question, because each SDK has a different idea of what an agent is.
This is a buying guide, not a tutorial. It compares the OpenAI Agents SDK, Google's Agent Development Kit 2.0 and Anthropic's Claude Agent SDK as they stand in October 2026, across languages, multi-agent design, sandboxing, memory, tracing, evaluation, model portability and the hosted runtime each one leads to. Every capability claim comes from the vendors' own documentation, linked at the end.
If you want the smallest set of primitives and plan to run on OpenAI models, the OpenAI Agents SDK is the shortest path: an Agent, a Runner, handoffs between agents, guardrails, sessions and sandbox agents with their own workspace. If you want deterministic orchestration, built-in evaluation or a Go SDK, Google ADK 2.0 is the stronger framework. If the agent's job is to work on files and run commands, the Claude Agent SDK hands you Claude Code's tools, permissions and hooks without writing any of them.
The less obvious split is where each SDK expects to run in production. OpenAI and Anthropic now both offer a hosted harness next to the library, the Agents API and Claude Managed Agents, while ADK deploys to Google Cloud's Agent Runtime, Cloud Run or GKE. Choosing an SDK quietly chooses a migration path as well.
The table compresses each vendor's documentation into the nine questions that decided my shortlist. Where a cell says a capability is not listed, the vendor docs did not offer it as an SDK primitive, which is not the same as impossible.
| Dimension | OpenAI Agents SDK | Google ADK 2.0 | Claude Agent SDK |
|---|---|---|---|
| Languages | Python and TypeScript | Python, Go and TypeScript at 2.0 GA; Java and Kotlin also documented | Python and TypeScript; other languages by running the CLI headless |
| Core abstraction | Agent objects plus a Runner that owns the loop | Agent and Workflow, a graph of agent nodes and function nodes | query() or ClaudeSDKClient driving a Claude Code CLI subprocess |
| Multi-agent model | Handoffs, or agents exposed as tools | Graph edges and routes, plus sequential, parallel and loop patterns | Subagents declared with AgentDefinition |
| Sandboxing | SandboxAgent with a Manifest; Unix-local, Docker and hosted providers | Not an SDK primitive; isolation comes from the deployment target | Runs in a container you provide; permissions and hooks gate each tool call |
| Sessions and memory | Session backends for SQLite, SQLAlchemy, Redis, MongoDB, Dapr and encrypted storage | Sessions, memory and artifacts assembled into context; 2.0 changed the session schema | JSONL transcripts on local disk with resume and fork; a SessionStore adapter mirrors them |
| Tracing | Tracing is a core primitive, with observability integrations | Logging, metrics and traces documented | OpenTelemetry traces, metrics and logs configured by environment variables |
| Evaluation | No eval runner listed among the SDK primitives | adk eval CLI, adk web UI and pytest AgentEvaluator with trajectory metrics | No eval runner listed |
| Model portability | OpenAI by default; LiteLLM and Any-LLM adapters, both beta | Gemini and Gemma plus Claude, OpenAI, Ollama, vLLM and LiteLLM | Claude only, via the Anthropic API, Amazon Bedrock or Google Cloud |
| Hosted runtime | Agents API, a managed harness in OpenAI's service | Agent Runtime, Cloud Run or GKE on Google Cloud | Claude Managed Agents, beta header managed-agents-2026-04-01 |
Two rows matter more than they look. Model portability decides whether you can try a cheaper model later without a rewrite, and the hosted-runtime row decides what moving off your own servers would cost once the agent is in production.
The same small task, a billing question routed to a specialist, looks different in each SDK, and the difference is the whole story. OpenAI gives you objects and a Runner that loops over them. ADK 2.0 gives you a graph in which an LLM agent and a plain Python function are both nodes, joined by edges. The Claude Agent SDK gives you a process: query() starts the Claude Code CLI, which already knows how to read, edit, search and run shell commands.
# 1) OpenAI Agents SDK (pip install openai-agents)
# The Runner owns the loop; agents are objects you wire together.
from agents import Agent, Runner, SQLiteSession
billing = Agent(name="Billing", instructions="Answer invoice questions.")
triage = Agent(
name="Triage",
instructions="Route the user to the right specialist.",
handoffs=[billing], # Billing takes over the conversation, not just one call
)
session = SQLiteSession("user-42") # history survives between runs
result = Runner.run_sync(triage, "Why was INV-118 charged twice?", session=session)
print(result.final_output)
# 2) Google ADK 2.0 (pip install google-adk)
# A Workflow is a graph: agents and plain functions are both nodes.
from google.adk import Agent, Workflow
extract = Agent(name="extract", model="gemini-flash-latest",
instruction="Return only the invoice number in the message.")
def lookup(node_input: str):
return {"invoice": node_input, "status": "paid twice"} # deterministic code node
root_agent = Workflow(name="root_agent",
edges=[("START", extract, lookup)])
# 3) Claude Agent SDK (pip install claude-agent-sdk)
# query() spawns the Claude Code CLI; its built-in tools come with it.
import asyncio
from claude_agent_sdk import query, ClaudeAgentOptions, AgentDefinition
options = ClaudeAgentOptions(
allowed_tools=["Read", "Grep", "Bash"], # Claude Code's own tools, no wiring
max_turns=20, # there is no session timeout, so bound the loop yourself
agents={"reviewer": AgentDefinition(
description="Reviews a diff for bugs",
prompt="You are a strict code reviewer.",
tools=["Read", "Grep"])},
)
async def main():
async for message in query(prompt="Find why ./billing tests fail", options=options):
print(message)
asyncio.run(main())That last point is easy to underrate. With OpenAI and ADK, every tool the agent can touch is a function you wrote or a hosted tool you enabled. With the Claude Agent SDK, Read, Edit, Bash and the rest arrive pre-built, so the work shifts from writing tools to restricting them with allowed_tools, permission modes and hooks. For a document or code agent that is a large head start. For a narrow API agent it is a far bigger runtime than the job needs.
All three can run more than one agent, but they disagree about who holds control while the work is happening.
OpenAI made the sandbox a first-class SDK object. A SandboxAgent gets a workspace built from a Manifest of files, repositories and mounts, runs through a sandbox client such as UnixLocalSandboxClient or the Docker client, and can resume from saved sandbox state or a snapshot. Conversation memory is separate, a Session backed by SQLite, SQLAlchemy, Redis, MongoDB or Dapr. ADK treats sessions, memory and artifacts as its own context layer, and the 2.0 release changed the schema for custom session storage, so a 1.x store needs migrating.
The Claude Agent SDK starts from the opposite end: everything is local by default. Each session is a CLI subprocess with its own shell and working directory, writing a JSONL transcript under the projects folder of ~/.claude. Anthropic's hosting guide suggests 1 GiB of RAM, 5 GiB of disk and 1 CPU per agent as a starting point, and concurrency on a host is bounded by how many of those subprocesses its RAM can hold.
A Claude Agent SDK container that restarts loses its transcripts, CLAUDE.md memory files and working-directory artifacts. A SessionStore adapter mirrors transcripts only, and mirror writes are best effort: a batch that fails is dropped with a mirror_error system message while the query carries on. Alert on that message if users expect to resume their sessions.
Evaluation is where ADK pulls clearly ahead. It ships an adk eval command, an adk web UI with metric thresholds and an AgentEvaluator for pytest, and it scores the trajectory as well as the answer. tool_trajectory_avg_score demands the exact tool-call sequence, response_match_score uses ROUGE-1 against a reference, and LLM-judged and rubric-based metrics cover answers with no single correct wording. Neither the OpenAI nor the Claude SDK lists an eval runner among its primitives.
# Google ADK: evaluation is a CLI verb, not a separate product
adk eval ./billing_agent ./billing_agent/refunds.evalset.json \
--config_file_path=./billing_agent/test_config.json \
--print_detailed_results
# test_config.json -- thresholds per metric
# tool_trajectory_avg_score 1.0 = the exact tool-call sequence must match
# response_match_score 0.8 = ROUGE-1 against the reference answer
{
"criteria": {
"tool_trajectory_avg_score": 1.0,
"response_match_score": 0.8
}
}
# Claude Agent SDK: traces come from env vars, exported over OTLP
CLAUDE_CODE_ENABLE_TELEMETRY=1
CLAUDE_CODE_ENHANCED_TELEMETRY_BETA=1 # required for traces, not for metrics/logs
OTEL_TRACES_EXPORTER=otlp
OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf
OTEL_EXPORTER_OTLP_ENDPOINT=http://collector.example.com:4318Tracing is closer to a draw. Tracing is a core primitive of the OpenAI SDK, with integrations for outside observability tools. The Claude Agent SDK inherits OpenTelemetry settings from the environment, as above, and leaves prompt text and tool inputs out of exports unless you opt in. ADK documents logging, metrics and traces. In practice I would send all three to the same OpenTelemetry collector and compare them on one dashboard rather than in each vendor's viewer.
ADK is the most model-agnostic of the three: its docs list Gemini and Gemma beside Claude, OpenAI, Ollama, vLLM and LiteLLM. The OpenAI Agents SDK reaches other providers through two adapters, LiteLLM and Any-LLM, which its own documentation marks as beta.
# pip install "openai-agents[litellm]" (adapter is marked beta)
from agents import Agent, ModelSettings
agent = Agent(
name="Portable",
instructions="Summarise the purchase order.",
model="litellm/gemini/gemini-2.5-flash", # litellm/<provider>/<model>
# Some upstream providers return no usage numbers through the adapter;
# without this, your cost dashboard silently reads zero.
model_settings=ModelSettings(include_usage=True),
)The Claude Agent SDK runs Claude only. What varies is the route: the Anthropic API, Amazon Bedrock or Google Cloud, each needing outbound HTTPS to its own endpoint. That is real lock-in at the model layer, traded for a built-in tool set that the other two make you write yourself.
Portability through an adapter is not free. For LiteLLM, the OpenAI SDK docs warn that structured outputs, tool calling and usage metrics depend on the upstream provider. Run your eval set against the second model before you promise anyone a cheaper fallback.
This is the shortlist I would hand someone today, with the one reason that decides each case.
The rule I took away is to choose an agent SDK by the shape of its loop and by where it will run in production, then check model portability. Model choice is a constraint worth confirming early, but the orchestration model is the hardest thing to change once real traffic depends on it.