AI
OpenAI Agents SDK Guardrails: Input, Output and Tool Checks
October 202611 min read

Guardrails are checks the SDK runs on an agent's input, its final output, or individual function-tool calls. Each guardrail returns a verdict, and if its tripwire is triggered the runner raises an exception and halts the run. Input and output guardrails are attached to an agent, while tool guardrails are attached to a specific tool.
Yes, by default. With run_in_parallel=True the guardrail and the agent start together, which keeps latency low, but the agent may already have used tokens and called tools before a tripwire cancels it. Setting run_in_parallel=False makes the guardrail finish first, so a blocked request never reaches the model or any tool.
Input guardrails only run for the first agent in a run, and output guardrails only for the agent that produces the final output. A specialist that receives a handoff is not the first agent, so its input guardrails are skipped. For checks that must apply whichever agent acts, put a tool guardrail on the function tool itself.
reject_content skips the tool call, or replaces its output, with a message the model reads instead, so the run continues and the agent can correct itself. raise_exception halts the whole run with ToolInputGuardrailTripwireTriggered or ToolOutputGuardrailTripwireTriggered. Use the first for recoverable mistakes and the second for states the agent should never reach.
No. Tool guardrails use the function-tool pipeline, so hosted tools such as WebSearchTool, FileSearchTool and HostedMCPTool, and built-in tools such as ComputerTool and ShellTool, bypass them. Handoff calls and Agent.as_tool() also lack tool guardrail options. Local MCP servers can attach guardrails to every tool they expose.

Key Takeaway
OpenAI Agents SDK guardrails come in three layers: input guardrails run only on the first agent, output guardrails only on the last, and tool guardrails on every guarded function-tool call. Input checks run in parallel by default, so the agent may spend tokens and call tools before a tripwire fires. Use blocking mode or a tool guardrail for side effects.
The first version of a finance assistant I sketched for an ERP back-office had one guardrail: an input check that refused anything that was not about the ledger. It looked complete until I read the execution-mode section of the docs. By default that check runs at the same time as the agent, and the documentation says plainly that if the tripwire fires, the agent may already have consumed tokens and executed tools. For an agent whose main tool posts journal entries, may already have executed tools is the whole problem.
This post walks through the three guardrail layers of the OpenAI Agents SDK as documented for openai-agents 0.22.3, the PyPI release of 17 September 2026: input guardrails, output guardrails and the per-call tool guardrails. Each one is applied to a real risk in an accounting workflow, an off-topic request, personal data leaking from a reply, and a tool call that would post to the wrong branch, with the exceptions you need to catch. The TypeScript SDK documents the same three layers with camelCase names, such as runInParallel and toolExecution, so the design carries over unchanged.
The single most useful fact about guardrails is where they run, because it is not where most people assume. Guardrails are configured on the agent or the tool, but the runner only fires them at specific workflow boundaries.
| Layer | When it runs | What it sees and can do | Exception on tripwire |
|---|---|---|---|
| Input guardrail, on Agent.input_guardrails | Only when the agent is the first agent in the run. Parallel with the agent by default, or before it with run_in_parallel=False | The same input passed to the agent. Returns GuardrailFunctionOutput with tripwire_triggered and output_info | InputGuardrailTripwireTriggered |
| Output guardrail, on Agent.output_guardrails | Only on the agent that produces the final output, always after it completes. No parallel option | The final output, typed as the agent's output_type. Same GuardrailFunctionOutput shape | OutputGuardrailTripwireTriggered |
| Tool guardrail, on a FunctionTool | On every invocation of that tool, input check before execution and output check after, whichever agent calls it | The tool name and raw argument string, or the tool's result. Returns allow, reject_content or raise_exception | ToolInputGuardrailTripwireTriggered or ToolOutputGuardrailTripwireTriggered |
Read the second column twice. In a triage setup where a front-desk agent hands off to a finance specialist, the specialist's input guardrails never run, because it is not the first agent, and the triage agent's output guardrails never run, because it does not produce the final answer. The docs say it directly: if you need checks around each custom function-tool call in a workflow with managers, handoffs or delegated specialists, use tool guardrails instead of relying only on agent-level ones.
An input guardrail is any function that takes the run context, the agent and the input, and returns a GuardrailFunctionOutput. The common pattern is a small classifier agent with a typed output_type, run from inside the guardrail. The only decision that really matters is the execution mode, which you set on the decorator.
from pydantic import BaseModel
from agents import (
Agent,
GuardrailFunctionOutput,
RunContextWrapper,
Runner,
TResponseInputItem,
)
from agents.decorators import input_guardrail
class ScopeCheck(BaseModel):
in_scope: bool
reason: str
# A small, cheap agent whose only job is a yes/no verdict.
# Point it at the cheapest model you trust for classification.
scope_checker = Agent(
name="Finance scope check",
instructions=(
"Answer in_scope=true only if the request is about this company's "
"ledger, journals, invoices or branch accounts. Anything else is false."
),
output_type=ScopeCheck,
)
# Blocking: the finance agent does not start until this returns.
# Costs one extra round-trip of latency; buys zero tokens and zero
# tool calls on a request that was never going to be served.
@input_guardrail(name="finance_scope", run_in_parallel=False)
async def finance_scope(
ctx: RunContextWrapper, agent: Agent, input: str | list[TResponseInputItem]
) -> GuardrailFunctionOutput:
verdict = await Runner.run(scope_checker, input, context=ctx.context)
return GuardrailFunctionOutput(
output_info=verdict.final_output, # kept for logging
tripwire_triggered=not verdict.final_output.in_scope,
)Parallel mode, the default, starts the guardrail and the agent together, so a clean request pays no extra latency. Blocking mode makes the agent wait for the verdict, which costs one classifier round-trip on every request, but when the tripwire fires the agent never executes: no tokens on the expensive model and no tool calls. I chose blocking for any agent that owns a side-effecting tool, and kept parallel for read-only assistants where a cancelled run only wastes a few tokens.
Parallel mode is a cost optimisation, not a safety boundary. The SDK documents that when a parallel input guardrail trips, the agent may already have consumed tokens and executed tools before being cancelled. If the first turn can call a tool that writes to the ledger, sends an email or calls a payment API, either set run_in_parallel=False or put the real check on the tool itself.
Output guardrails receive the final output after the last agent finishes, which is why there is no parallel mode to choose. A model-based check works here too, but for personal data a deterministic pattern is faster and easier to reason about. The example below trips on a 16-digit NIK or an email address in the reply.
import re
from agents import Agent, GuardrailFunctionOutput, RunContextWrapper
from agents.decorators import output_guardrail
# NIK (Indonesian national ID) is 16 digits. NPWP and bank account numbers
# deserve their own patterns; start with the ones your ledger actually stores.
NIK = re.compile(r"\b\d{16}\b")
EMAIL = re.compile(r"[\w.+-]+@[\w-]+\.[\w.]+")
@output_guardrail(name="no_pii_in_reply")
async def no_pii_in_reply(
ctx: RunContextWrapper, agent: Agent, output: str
) -> GuardrailFunctionOutput:
hits = {
"nik": len(NIK.findall(output)),
"email": len(EMAIL.findall(output)),
}
return GuardrailFunctionOutput(
# Log counts, never the matched values: output_info ends up in logs.
output_info=hits,
tripwire_triggered=any(hits.values()),
)
finance_agent = Agent(
name="Finance assistant",
instructions="Answer questions about journals and branch balances.",
input_guardrails=[finance_scope],
output_guardrails=[no_pii_in_reply],
)Two behaviours from the docs shape how you use this. First, a tripwire and a crash are treated differently: when the tripwire fires, the session keeps the completed tool calls and outputs but drops the rejected final answer, while an exception raised inside the guardrail is treated as an unknown verdict and the final-turn items are persisted before the error surfaces. So return a verdict, do not raise. Second, when tool_use_behavior makes a tool's result the final output and the guardrail rejects it, the SDK replaces the stored payload with the placeholder Output withheld by an output guardrail. You can change that text with RunConfig.output_guardrail_blocked_message, but keep it free of data, because it is persisted and replayed.
Tool guardrails are the layer built for the ERP case. They attach to the tool, so they run on every call regardless of which agent made it, and they see the arguments the model actually produced. The input guardrail below reads the raw tool_arguments string and the caller's own context object, then decides between three outcomes.
import json
from dataclasses import dataclass
from pydantic import BaseModel
from agents import ToolGuardrailFunctionOutput, function_tool
from agents.decorators import tool_input_guardrail
@dataclass
class ErpUser: # passed as Runner.run(..., context=ErpUser(...))
user_id: str
allowed_branches: set[str]
class JournalLine(BaseModel):
account: str
debit: float
credit: float
@tool_input_guardrail
def branch_and_balance(data):
# data.context is a ToolContext: it carries tool_name and the RAW
# tool_arguments string, and .context is your own ErpUser object.
args = json.loads(data.context.tool_arguments or "{}")
user: ErpUser = data.context.context
branch = args.get("branch_code")
if branch not in user.allowed_branches:
# Recoverable: the call is skipped and the model reads this text
# instead of a tool result, so it can ask the user which branch.
return ToolGuardrailFunctionOutput.reject_content(
f"Branch {branch} is not one this user may post to. "
f"Allowed: {', '.join(sorted(user.allowed_branches))}."
)
lines = args.get("lines", [])
debit = round(sum(l["debit"] for l in lines), 2)
credit = round(sum(l["credit"] for l in lines), 2)
if debit != credit:
# Not recoverable by rephrasing: an unbalanced entry means the model
# has lost the plot. Halt the run with ToolInputGuardrailTripwireTriggered.
return ToolGuardrailFunctionOutput.raise_exception(
output_info={"debit": debit, "credit": credit}
)
return ToolGuardrailFunctionOutput.allow()
@function_tool(tool_input_guardrails=[branch_and_balance])
def post_journal_entry(branch_code: str, memo: str, lines: list[JournalLine]) -> str:
"""Post a balanced journal entry to the given branch ledger."""
... # call the ERP here; the guardrail has already runThe two rejection paths do different jobs. reject_content skips the call and hands the model a message in place of a tool result, so the run continues and the agent can ask the user which branch they meant. raise_exception halts the whole run with ToolInputGuardrailTripwireTriggered, which is right for a state the model should never have reached, such as debits that do not equal credits. Note that the branch list comes from your context object, never from the prompt, so the model cannot talk its way into another branch.
If the tool also has needs_approval, input tool guardrails normally run after a human approves and just before execution. Set RunConfig.tool_execution to ToolExecutionConfig with pre_approval_tool_input_guardrails=True and the same check also runs before the approval request is raised, so nobody is asked to approve an entry that was going to be rejected anyway. It still runs again after approval.
The gaps are documented, and each one is a place where an agent can act without the check you think is in front of it.
The practical rule I took from this list: anything that writes must be a function tool you own, with a tool guardrail on it. If a write goes through a hosted tool or a remote MCP server, the guardrail has to live on the server side, because the SDK cannot see inside that call.
Each layer raises its own exception, and they carry their data in different places. Agent-level tripwires expose guardrail_result, which names the guardrail and holds its output. Tool tripwires expose the guardrail and its ToolGuardrailFunctionOutput directly on the exception. All four are importable from the agents package.
from agents import (
InputGuardrailTripwireTriggered,
OutputGuardrailTripwireTriggered,
Runner,
ToolInputGuardrailTripwireTriggered,
ToolOutputGuardrailTripwireTriggered,
)
async def ask_finance(message: str, user: ErpUser) -> dict:
try:
result = await Runner.run(finance_agent, message, context=user)
return {"status": 200, "reply": result.final_output}
except InputGuardrailTripwireTriggered as exc:
# Agent-level tripwires carry guardrail_result; it names the guardrail.
name = exc.guardrail_result.guardrail.get_name()
log.info("input tripwire", guardrail=name, user=user.user_id)
return {"status": 422, "reply": "I can only help with ledger questions."}
except OutputGuardrailTripwireTriggered:
# The draft answer was rejected. Do not retry blindly with the same
# prompt; it will usually produce the same identifiers again.
return {"status": 502, "reply": "The answer contained personal data and was withheld."}
except ToolInputGuardrailTripwireTriggered as exc:
# Tool tripwires expose the verdict directly on exc.output.
log.warning("journal blocked", info=exc.output.output_info, user=user.user_id)
return {"status": 409, "reply": "That entry does not balance, so nothing was posted."}
except ToolOutputGuardrailTripwireTriggered as exc:
log.warning("tool output blocked", info=exc.output.output_info)
return {"status": 502, "reply": "A lookup returned data I am not allowed to show."}Map each one to a distinct response, because they mean different things to the caller: an out-of-scope request is a client error, a withheld answer is a server-side refusal, and a blocked journal entry is a conflict nothing was written for. The docs also note that exception.run_data keeps the guardrail results gathered before the run stopped, including tool guardrail results from completed turns, which is the right thing to send to your audit log. run_data can be None when the exception is raised outside a runner-managed path, so guard for that before reading it.
OpenAI's own guide separates the two cleanly: guardrails validate input, output or tool behaviour automatically, while human review pauses the run so a person or policy can approve a sensitive action. In practice I layer them in this order for any agent that touches money.
None of these replaces permissions in the ERP itself. A guardrail is a check inside your process; the ledger's own role and branch permissions are the boundary that still holds when the agent, the SDK or your guardrail code has a bug.
The rule I carry from this: put a check on the boundary where the action happens, not where the conversation starts. Input guardrails keep cost and scope under control, output guardrails protect what leaves, but only a tool guardrail is guaranteed to stand in front of every write, whichever agent makes it.