AI
OpenAI Agents SDK Handoffs vs Agents as Tools Explained
October 202611 min read

A handoff transfers ownership of the conversation to another agent, which then sees the history and answers the user directly. Agent.as_tool() runs the other agent as a nested run on generated input and returns its output to the calling agent, which keeps control and writes the reply.
Input guardrails run only for the first agent in the chain, because they are meant to check user input. A guardrail attached to a specialist that is reached only through a handoff never executes. Attach it to the entry agent instead, or move the check onto the function tools as a tool guardrail.
Yes, by default the new agent sees the entire previous conversation. You can change that with an input_filter on the handoff, a global handoff_input_filter on RunConfig, or the opt-in nest_handoff_history beta. None of these redacts content on its own, so sensitive tool output needs a filter that removes it explicitly.
Use agents as tools when one agent should own the final answer, combine results from several specialists in one turn, or enforce one set of output guardrails. Use handoffs when the specialist should talk to the user directly and keep the conversation over several turns with its own instructions.
Yes. A common shape is a manager agent that calls specialists as tools for ordinary questions, plus a handoff to an escalation agent that takes over when the conversation needs a different owner. The two primitives answer different questions: who borrows expertise, and who owns the reply.

Key Takeaway
In the OpenAI Agents SDK, a handoff transfers the conversation to a specialist, which then sees the history and answers the user. Agent.as_tool() runs the specialist as a nested run on generated input and returns its output to the orchestrator, which keeps control. Guardrails differ: input guardrails run only for the first agent, output guardrails only for the last.
Consider a typical first version of an ERP helpdesk agent: a triage agent and two specialists, sales and finance, wired together with handoffs. It routes well. Then an input guardrail is added to the finance agent so it refuses to discuss any account other than the one the customer is signed in with, someone asks about another customer's invoice, and the finance agent answers anyway. The guardrail never ran. Nothing is broken; the composition primitive was chosen without reading its guardrail rules.
The OpenAI Agents SDK gives you two ways to make agents work together. A handoff moves ownership of the conversation to another agent. Agent.as_tool() turns an agent into a callable tool, so the calling agent keeps control and decides what to do with the result. This post compares them on the things that change in production: who answers the user, what history each agent sees, where guardrails fire and how the run shows up in a trace. It uses one sales versus finance triage example throughout, the Python SDK, and the SDK's own documentation and source as the authority.
A handoff is exposed to the model as a tool, named transfer_to_ followed by the agent name by default. When the model calls it, the runner does not return anything to the triage agent. It swaps the current agent, keeps the input, and re-runs the loop, so the specialist becomes the active agent for the rest of the turn and writes the final answer itself. You can pass an Agent straight into handoffs, or wrap it in handoff() to add an on_handoff callback, an input_type schema for model-generated metadata, an input_filter, or an is_enabled switch.
from pydantic import BaseModel
from agents import Agent, Runner, RunContextWrapper, function_tool, handoff
from agents.extensions.handoff_prompt import RECOMMENDED_PROMPT_PREFIX
@function_tool
async def get_invoice(invoice_no: str) -> str:
"""Return status, due date and open amount for one invoice."""
return await erp.invoices.summary(invoice_no)
@function_tool
async def get_price_list(customer_code: str) -> str:
"""Return the price list and discount tier assigned to a customer."""
return await erp.pricing.for_customer(customer_code)
class RouteReason(BaseModel):
reason: str # model-generated metadata, e.g. "payment_not_matched"
customer_code: str
async def log_route(ctx: RunContextWrapper[None], data: RouteReason) -> None:
# Runs BEFORE the specialist takes over. Raise here to stop the transfer;
# returning normally lets it continue.
await audit.write("helpdesk.route", reason=data.reason, customer=data.customer_code)
sales_agent = Agent(
name="Sales agent",
instructions=f"{RECOMMENDED_PROMPT_PREFIX}\nYou answer price lists, quotations and discounts.",
tools=[get_price_list],
)
finance_agent = Agent(
name="Finance agent",
instructions=f"{RECOMMENDED_PROMPT_PREFIX}\nYou answer invoice status, due dates and payments.",
tools=[get_invoice],
)
triage_agent = Agent(
name="Helpdesk triage",
instructions=f"{RECOMMENDED_PROMPT_PREFIX}\nRoute pricing to sales, invoices and payments to finance.",
handoffs=[
sales_agent, # exposed as transfer_to_sales_agent
handoff(finance_agent, on_handoff=log_route, input_type=RouteReason),
],
)
result = await Runner.run(triage_agent, "INV-2026-0912 still shows unpaid, we transferred last week")
print(result.last_agent.name) # "Finance agent" - it now owns the conversationTwo details matter here. First, result.last_agent is the agent that should usually handle the next user turn, so a chat loop that passes it back into Runner.run keeps the customer talking to finance instead of re-triaging every message. Second, input_type is metadata about the transfer, such as a reason or a priority, and not a way to pick a destination: handoff() always transfers to the agent you wrapped, so you register one handoff per specialist and let the model choose. The receiving agent still sees the conversation history unless you change it.
Agent.as_tool() returns a FunctionTool. When the orchestrator calls it, the SDK starts a fresh Runner.run with the specialist as the starting agent, feeds it the tool arguments, and hands the final output back to the orchestrator as the tool result. The orchestrator then carries on: it can call another specialist, combine both answers, or ask a follow-up. The SDK docstring states the two differences plainly. In a handoff the new agent receives the conversation history and takes over; as a tool, the new agent receives generated input and the original agent continues.
helpdesk_manager = Agent(
name="Helpdesk manager",
instructions=(
"You own the reply to the customer. Ask the specialists for facts, "
"then answer in one message. Never forward a specialist's text verbatim."
),
tools=[
sales_agent.as_tool(
tool_name="ask_sales",
tool_description="Prices, quotations and discount rules for one customer.",
),
finance_agent.as_tool(
tool_name="ask_finance",
tool_description="Invoice status, due dates and open balance for one customer.",
max_turns=4, # the nested run gets its own budget (default is 10)
failure_error_function=None, # raise, instead of handing the model an error string
),
],
)
# The specialist does NOT see this message. It sees whatever the manager
# writes into the tool call, by default {"input": "..."}.
result = await Runner.run(
helpdesk_manager,
"Can CUST-0418 get the Q4 distributor discount, and is INV-2026-0912 still open?",
)
print(result.last_agent.name) # "Helpdesk manager" - control never movedBecause it is a separate run, the nested agent has its own settings. It gets its own max_turns, which falls back to the SDK default of 10 when you do not set one, its own run_config and hooks, and it does not inherit the parent's history unless you pass a session, conversation_id or previous_response_id explicitly. Failure handling also differs from a handoff: by default a failure inside the nested run is turned into an error message for the orchestrator's model through failure_error_function, and passing None makes it raise instead. For a finance lookup I would rather see an exception than a model improvising around one.
The SDK's multi-agent guide frames the choice in one line each. Agents as tools suit a manager that should own the final answer, combine specialists or enforce shared guardrails. Handoffs suit a triage step whose chosen specialist should respond directly with focused instructions. The table below adds the mechanical consequences.
| Concern | Handoff | Agent as tool |
|---|---|---|
| Who answers the user | The specialist, directly | The orchestrator, after reading the specialist's output |
| Active agent for the next turn | The specialist, via result.last_agent | Still the orchestrator |
| What the specialist sees | Full conversation history, unless filtered | Only the tool arguments, by default a single input string, or a typed model via parameters |
| Several specialists in one turn | No, one transfer ends the triage agent's part | Yes, the orchestrator can call several and merge the results |
| Input guardrails | Only the first agent's, usually the triage agent | The orchestrator's on the user message; the specialist's on its generated input, since it starts its own run |
| Output guardrails | Only those of whichever specialist ends the turn | The orchestrator's always cover the final answer |
| Tool guardrails on the call itself | Not applied, handoffs use their own pipeline | Not exposed directly by as_tool() today |
| Trace shape | A handoff_span, then the specialist's agent span | A function tool call that wraps a nested agent run |
| Approval before it runs | Authorise inside on_handoff and raise to stop | needs_approval pauses the run with an interruption |
The guardrail rows deserve their own section, because they are where the two patterns stop being interchangeable. The trace row is worth a sentence too: with handoffs a trace reads as a relay, one agent after another, while with agents as tools it reads as one agent making calls, which is easier to scan when a manager consults three specialists in a single turn.
The guardrails page states both rules in one place. Input guardrails run only for the first agent in the chain, because they are meant to check user input. Output guardrails run only for the agent that produces the final output. In a handoff chain the first agent is the triage agent and the last is whichever specialist answered, so a guardrail attached to the wrong agent is valid code that never executes. That was my bug.
from agents import (
Agent, GuardrailFunctionOutput, InputGuardrailTripwireTriggered,
RunContextWrapper, Runner, input_guardrail,
)
@input_guardrail
async def one_customer_only(ctx: RunContextWrapper[None], agent: Agent, input) -> GuardrailFunctionOutput:
codes = extract_customer_codes(input) # regex over CUST-xxxx
allowed = ctx.context.session_customer_code # who is actually logged in
leaked = [c for c in codes if c != allowed]
return GuardrailFunctionOutput(output_info=leaked, tripwire_triggered=bool(leaked))
# Wrong: the guardrail sits on a specialist that is only reached by a handoff.
finance_agent = Agent(name="Finance agent", tools=[get_invoice],
input_guardrails=[one_customer_only])
triage_agent = Agent(name="Helpdesk triage", handoffs=[sales_agent, finance_agent])
await Runner.run(triage_agent, text) # one_customer_only never runs
# Right for handoffs: guard the entry point, and put the output check on
# EVERY agent that can end the turn - any specialist may be the last one.
triage_agent = Agent(name="Helpdesk triage", handoffs=[sales_agent, finance_agent],
input_guardrails=[one_customer_only])
finance_agent.output_guardrails = [no_internal_margin]
sales_agent.output_guardrails = [no_internal_margin]
try:
result = await Runner.run(triage_agent, text, context=ctx)
except InputGuardrailTripwireTriggered:
reply = "I can only discuss the account you are signed in with."With agents as tools the coverage looks different. The manager is both the first and the last agent of the outer run, so its input guardrail sees every user message and its output guardrail sees every reply, whichever specialists it consulted. Each specialist also starts its own nested run, so its own guardrails fire against the input the manager generated for it. That is a different string from what the user typed, which matters for a check like one_customer_only: the manager may have paraphrased the customer code away. Put user-facing checks on the manager and data checks on the function tools.
A guardrail on a handoff target is not an error and produces no warning. If a check must hold no matter which agent ends up answering, attach it as a tool guardrail to the function tools that touch the data, such as get_invoice. Tool guardrails run on every guarded call, in either pattern, because the function tool executes inside whichever agent is active.
A handoff forwards the whole conversation by default, which is convenient and is also how a sales agent ends up reading finance tool output it has no business with. The SDK gives you four levers, and they interact.
from agents import RunConfig, handoff
from agents.extensions import handoff_filters
# Handoff: drop structured tool items the triage agent produced. This does NOT
# redact tool text that was already copied into ordinary messages.
faq_handoff = handoff(faq_agent, input_filter=handoff_filters.remove_all_tools)
# Opt-in beta: summarise earlier history into <CONVERSATION HISTORY> segments.
# Ignored when an input_filter is set on the handoff or on the RunConfig.
config = RunConfig(nest_handoff_history=True)
# Agent as tool: the specialist only ever sees the arguments, so make them typed.
class InvoiceQuery(BaseModel):
invoice_no: str
customer_code: str
ask_finance = finance_agent.as_tool(
tool_name="ask_finance",
tool_description="Look up one invoice for one customer.",
parameters=InvoiceQuery,
include_input_schema=True,
)Agents as tools sidestep most of this, because the specialist sees only what the orchestrator writes into the call. The cost is that the orchestrator's model decides what to write. Giving the tool a typed schema with parameters and include_input_schema turns a free-text request into fields the specialist can rely on, and the nested agent can read the parsed value from RunContextWrapper.tool_input.
Nested handoff history changes how the transcript is represented; the docs say explicitly that it does not redact anything. Tool arguments and outputs can survive inside the generated summary. If finance data must not reach the sales agent, write an input filter that sanitises input_history, pre_handoff_items and new_items, not just the items you forward.
The two patterns put the authorisation hook in different places. For a handoff, is_enabled is evaluated while the SDK prepares the available handoffs, before the model has produced any arguments, so it can hide the finance route for a user without finance access but it cannot inspect the customer code in the request. The docs say to perform that check at the start of on_handoff and to raise rather than return, because the transfer continues once on_handoff returns successfully.
An agent tool takes needs_approval, a boolean or a callable policy. When it triggers, the run pauses and the pending call appears in result.interruptions, and you resume it with the run state's approve or reject. That is a better fit for an ERP action with consequences, such as a specialist that drafts a credit note, because a human can see the exact arguments before the nested run starts. Both patterns accept is_enabled, so you can show a specialist only to the roles that should reach it.
Prefix every agent's instructions in a handoff graph with RECOMMENDED_PROMPT_PREFIX from agents.extensions.handoff_prompt, or call prompt_with_handoff_instructions. Without it, specialists tend not to understand that they were handed a conversation midway and may greet the user again or re-ask what triage already established.
After rebuilding the helpdesk both ways, I settled on a short decision procedure. It assumes you already decided the task needs more than one agent, which is not a given: one agent with good tools is often enough.
The two are not exclusive. The version I kept is a manager that calls sales and finance as tools for ordinary questions, plus one handoff to a human-escalation agent that takes over the conversation when a dispute needs a person. Routing that is part of the workflow is a handoff; borrowing expertise for a sub-question is a tool call.
Choose the primitive by who should own the reply. If the specialist should speak to the user and keep the conversation, hand off, then guard the first agent and every agent that can end the turn. If the specialist only supplies facts, call it as a tool and let one manager own the answer and its guardrails. Either way, put the checks that must never be skipped on the function tools.