AI
LangChain create_agent Middleware Guide: PII, Limits, Fallback
October 202613 min read

LangChain 1.0 recommends langchain.agents.create_agent instead of langgraph.prebuilt.create_react_agent. It still runs on the LangGraph runtime, but the prompt argument is renamed system_prompt, and pre_model_hook, post_model_hook and ToolNode error handling are replaced by middleware passed in a middleware list.
before_agent and before_model hooks run from the first middleware in the list to the last. after_model and after_agent hooks run in reverse, last to first. wrap_model_call and wrap_tool_call nest like function calls, so the first middleware listed is the outermost wrapper.
No. PIIMiddleware checks user input by default, while apply_to_output and apply_to_tool_results both default to False. If a tool returns customer records, set apply_to_tool_results=True for every PII type that can appear in that data, or the values reach the model unchanged.
ModelRetryMiddleware defaults to on_failure='continue', which turns the final error into an AIMessage rather than raising it. The fallback middleware only reacts to exceptions, so it never sees the failure. Set on_failure='error' on the retry and list the fallback before it so the fallback is the outer wrapper.
Use HumanInTheLoopMiddleware with an interrupt_on entry for the tool and add a when predicate that returns True only when the arguments exceed your threshold. The when option requires langchain 1.3.3 or later. The agent also needs a checkpointer and a thread_id so the paused run can be resumed with a Command carrying the decision.

Key Takeaway
In LangChain 1.x, create_agent replaces create_react_agent, and the old pre_model_hook, post_model_hook and ToolNode error handling become middleware. Built-in classes cover PII redaction, model and tool-call limits, retries, model fallback and human approval. Order matters: before hooks run first to last, after hooks last to first, and wrap hooks nest.
Search for how to add a guardrail to a LangChain agent and many of the top answers still import create_react_agent from langgraph.prebuilt and pass it a pre_model_hook. That code describes the API before LangChain 1.0. The 1.0.0 release, published on PyPI on 17 October 2025, moved the agent to langchain.agents.create_agent and turned the scattered hook arguments into a single middleware list. Copying an older answer into a current project now fails at the first ToolNode argument.
This guide covers langchain create_agent middleware as documented for langchain 1.4.3, the PyPI release of 28 September 2026, which requires Python 3.10 or later. It migrates a small accounts-payable agent from the old API, explains the six hooks and the order they run in, and then builds the guardrails an ERP agent actually needs: PII redaction that also covers tool results, call budgets, a custom closed-period guard, retries that hand over to a fallback model, and human approval above an amount threshold. Every class and parameter name comes from the LangChain documentation and source.
The rename is the smallest part of the migration. create_agent still runs on the LangGraph runtime, so persistence, checkpointing and interrupts behave as before, but most customisation arguments moved. The migration guide lists the changes; these are the ones that break existing code.
| Before 1.0 | LangChain 1.x | What to watch for |
|---|---|---|
| from langgraph.prebuilt import create_react_agent | from langchain.agents import create_agent | Same LangGraph runtime underneath, so checkpointers carry over |
| prompt= | system_prompt= | A string; dynamic prompts move to the @dynamic_prompt middleware |
| pre_model_hook= | Middleware with before_model | Several can stack; summarisation ships as SummarizationMiddleware |
| post_model_hook= | Middleware with after_model | Tool approval ships as HumanInTheLoopMiddleware |
| model= a callable that picks a model | wrap_model_call with request.override(model=...) | Models pre-bound with bind_tools are rejected |
| tools=ToolNode(..., handle_tool_errors=...) | tools=[...] plus a wrap_tool_call middleware | ToolNode instances are no longer accepted |
| Custom state as Pydantic model or dataclass | TypedDict extending AgentState | Set via state_schema on the agent or on a middleware |
| Stream node named agent | Stream node named model | Any filter matching on the node name silently stops matching |
| Dependencies in config configurable | context= at invoke, context_schema= on create_agent | Typed, and reachable from middleware through runtime.context |
Here is a typical pre-1.0 agent and its 1.x equivalent. The trimming hook becomes the built-in summarisation middleware, and the ToolNode error handler becomes a four-line wrap_tool_call function.
# Before: LangGraph prebuilt agent (pre-1.0 style)
from langgraph.prebuilt import create_react_agent, ToolNode
def trim_history(state): # pre_model_hook
...
agent = create_react_agent(
model="openai:gpt-5.4-mini",
tools=ToolNode(
[lookup_invoice, post_journal_entry],
handle_tool_errors=lambda e: f"Tool error: {e}",
),
prompt="You are an accounts-payable assistant.",
pre_model_hook=trim_history,
)
# After: langchain 1.x
from langchain.agents import create_agent
from langchain.agents.middleware import SummarizationMiddleware, wrap_tool_call
from langchain.messages import ToolMessage
@wrap_tool_call
def tool_errors_to_model(request, handler):
try:
return handler(request)
except ValueError as exc:
# Only bad-but-schema-valid input belongs here. Network failures go to
# ToolRetryMiddleware; bugs in the tool should still bubble up.
return ToolMessage(
content=f"Tool error: {exc}",
tool_call_id=request.tool_call["id"],
)
agent = create_agent(
model="openai:gpt-5.4-mini",
tools=[lookup_invoice, post_journal_entry], # plain tools, no ToolNode
system_prompt="You are an accounts-payable assistant.",
middleware=[
SummarizationMiddleware(model="openai:gpt-5.4-mini", trigger={"tokens": 4000}),
tool_errors_to_model,
],
)
# Streaming filters change too: the node is now called "model", not "agent".Two changes are easy to miss in review. Legacy chains, retrievers, the indexing API and the hub module moved to the separate langchain-classic package, so an import from langchain.chains needs that package installed. And the streaming node rename is silent: code that filters events where the node equals agent keeps running and simply receives nothing.
Middleware exposes two kinds of hook. Node-style hooks, before_agent, before_model, after_model and after_agent, run at fixed points and return a dict that is merged into state. Wrap-style hooks, wrap_model_call and wrap_tool_call, receive the request and a handler, and decide whether to call it zero times, once or several times. With three middleware in the list, one model turn runs like this:
A node-style hook can also end the run early by returning jump_to with end, tools or model, provided the hook declares the targets with can_jump_to. That is how the built-in limit middleware stops a loop without raising. A custom middleware can also declare state_schema to add its own fields to agent state, such as a counter that before_model reads and after_model increments.
Use the decorators, such as @before_model and @wrap_tool_call, for a single hook with no configuration, and subclass AgentMiddleware when one concern needs several hooks, init-time settings, or both sync and async versions. Every class hook has an async twin with an a prefix, for example abefore_model, and an agent invoked with ainvoke needs it.
PIIMiddleware takes a PII type, a strategy and three scope flags. The built-in types are email, credit_card, ip, mac_address and url, and any other name works when you supply a detector as a regex string, a compiled pattern or a function. The strategies are block, which raises PIIDetectionError, redact, which substitutes a REDACTED placeholder named after the type, mask, which keeps the last characters, and hash, which replaces the value with a deterministic hash.
from langchain.agents.middleware import PIIMiddleware, PIIDetectionError
pii = [
# Built-in detector. Also scrub tool results: a vendor lookup returns emails.
PIIMiddleware(
"email",
strategy="redact",
apply_to_input=True,
apply_to_tool_results=True,
),
# Custom type: NIK, the 16-digit Indonesian national ID. The regex string
# is the detector; mask keeps the last digits so a human can still match it.
PIIMiddleware(
"nik",
detector=r"\b\d{16}\b",
strategy="mask",
apply_to_input=True,
apply_to_output=True,
apply_to_tool_results=True,
),
# Fail closed: a card number in a chat about invoices is never legitimate.
PIIMiddleware("credit_card", strategy="block", apply_to_input=True),
]
try:
result = agent.invoke({"messages": [{"role": "user", "content": text}]})
except PIIDetectionError:
# Raised by strategy="block". Tell the user, do not echo the input back.
reply = "Please remove the card number and send the request again."Consider an ERP helpdesk agent for an Indonesian distributor. Customers paste their NIK, the 16-digit national ID, into chat, and the vendor lookup tool returns contact emails straight from the master data. A custom nik type with mask strategy covers the first case; apply_to_tool_results covers the second, which is the one most often missed. Hashing is the better choice when the model must still tell two people apart without seeing either number.
The defaults only scan user input. apply_to_output and apply_to_tool_results are both False, so out of the box a tool that returns a customer record sends every identifier in it to the model provider unchanged. Set the flags on every type that can appear in your tool data, and remember that a regex detector is a filter, not a compliance programme.
Two built-in classes stop an agent that loops. ModelCallLimitMiddleware caps model round-trips, and ToolCallLimitMiddleware caps tool calls either globally or for one named tool. Both take run_limit, which resets on every user message, and thread_limit, which counts across the whole conversation and therefore needs a checkpointer to persist the count.
from langchain.agents.middleware import (
ModelCallLimitMiddleware,
ToolCallLimitMiddleware,
)
from langchain.agents.middleware.tool_call_limit import ToolCallLimitExceededError
from langgraph.checkpoint.memory import InMemorySaver
budgets = [
# Hard ceiling on model round-trips in one user turn: stops a tool loop.
ModelCallLimitMiddleware(run_limit=12, exit_behavior="end"),
# Soft global cap: over-limit calls get an error ToolMessage and the
# model decides how to wrap up (exit_behavior="continue" is the default).
ToolCallLimitMiddleware(run_limit=20),
# A chatty lookup the model likes to call in a loop.
ToolCallLimitMiddleware(tool_name="search_vendor", run_limit=5),
# The side-effecting tool: one posting per turn, three per conversation,
# and exceeding either is an exception, not a polite message.
ToolCallLimitMiddleware(
tool_name="post_journal_entry",
run_limit=1,
thread_limit=3,
exit_behavior="error",
),
]
agent = create_agent(
model="openai:gpt-5.4-mini",
tools=[lookup_invoice, search_vendor, post_journal_entry],
middleware=budgets,
checkpointer=InMemorySaver(), # required for any thread_limit
)
try:
agent.invoke(payload, config={"configurable": {"thread_id": "ap-0042"}})
except ToolCallLimitExceededError:
alert_finance_team("journal posting budget exceeded", thread="ap-0042")The limits are independent, so stacking a global cap with per-tool caps is normal. ModelCallLimitMiddleware has only end and error as exit behaviours, and defaults to end, which ends the turn gracefully instead of raising.
Built-in classes cover generic risks; business rules need your own middleware. A wrap_tool_call function sees the tool name and arguments before execution and can return a ToolMessage without calling the handler, which means the tool never runs. Here, an agent is refused any journal posting dated inside a closed fiscal period, with the closing date passed in through the typed runtime context.
from dataclasses import dataclass
from datetime import date
from langchain.agents import create_agent
from langchain.agents.middleware import wrap_tool_call
from langchain.messages import ToolMessage
@dataclass
class ErpContext:
user_id: str
company: str
closed_until: date # end of the last closed fiscal period
@wrap_tool_call
def closed_period_guard(request, handler):
call = request.tool_call
if call["name"] != "post_journal_entry":
return handler(request)
ctx: ErpContext = request.runtime.context
posting_date = date.fromisoformat(call["args"]["posting_date"])
if posting_date <= ctx.closed_until:
# Short-circuit: handler() is never called, so nothing reaches the ERP.
# status="error" tells the model this was a refusal, not a result.
return ToolMessage(
content=(
f"Posting date {posting_date} falls in a closed period "
f"(closed until {ctx.closed_until}). Ask the user for a date "
"in an open period. Do not retry with the same date."
),
tool_call_id=call["id"],
status="error",
)
return handler(request)
agent = create_agent(
model="openai:gpt-5.4-mini",
tools=[lookup_invoice, post_journal_entry],
context_schema=ErpContext,
middleware=[closed_period_guard],
)
agent.invoke(
{"messages": [{"role": "user", "content": "Post the September freight accrual"}]},
context=ErpContext(user_id="u-17", company="PT Contoh", closed_until=date(2026, 8, 31)),
)Write the refusal for the model, not for a log. A message that states the rule and the next step, ask for a date in an open period, lets the model recover in the same turn, while a bare error string tends to produce the same call again. Because this check runs in Python with no model call, it adds no latency and cannot be talked out of its decision.
ModelRetryMiddleware retries a failed model call with exponential backoff; by default it makes two retries after the first attempt, waiting from one second and doubling, with jitter. ModelFallbackMiddleware catches an exception from the primary model and calls the inner chain again with each fallback model in turn, re-raising the last error if all of them fail. ToolRetryMiddleware does the same for tools and can be scoped to named tools and exception types.
from langchain.agents.middleware import (
ModelFallbackMiddleware,
ModelRetryMiddleware,
ToolRetryMiddleware,
)
# Wrong: ModelRetryMiddleware defaults to on_failure="continue", which turns
# the final error into an AIMessage. The outer fallback only acts on an
# exception, so it never sees one and the user gets an error message instead.
middleware = [
ModelFallbackMiddleware("anthropic:claude-sonnet-4-6"),
ModelRetryMiddleware(max_retries=2),
]
# Right: the exhausted retry raises, the fallback catches it and calls the
# inner chain again with the fallback model, which is retried the same way.
middleware = [
ModelFallbackMiddleware("anthropic:claude-sonnet-4-6"), # outermost
ModelRetryMiddleware(
max_retries=2, # 3 attempts in total
initial_delay=1.0,
backoff_factor=2.0, # 1 s, then 2 s, with jitter
on_failure="error",
),
# Tools get their own retry, scoped to the one that calls a flaky API.
ToolRetryMiddleware(
tools=["lookup_exchange_rate"],
max_retries=3,
retry_on=(ConnectionError, TimeoutError),
on_failure="continue", # the model reads the error and can carry on
),
]Because wrap hooks nest, the middleware listed first is the outermost. With fallback listed before retry, each model, primary or fallback, gets its full retry budget before the fallback moves on, which is usually the behaviour you want for transient rate-limit errors.
ModelRetryMiddleware defaults to on_failure continue, which converts the final error into an AIMessage instead of raising it. ModelFallbackMiddleware only reacts to an exception, so with the default setting the fallback never fires and the user reads an error message. Set on_failure to error on any retry that sits inside a fallback.
HumanInTheLoopMiddleware pauses the run after the model proposes a tool call and before the tool executes. Each entry in interrupt_on maps a tool name to True, False or a config with allowed_decisions, chosen from approve, edit, reject and respond. Since langchain 1.3.3 a config can also take a when predicate, so only calls that match a condition pause. Approval needs a checkpointer and a thread_id, because the paused state must survive until someone answers.
from langchain.agents.middleware import HumanInTheLoopMiddleware, ToolCallRequest
from langgraph.checkpoint.memory import InMemorySaver
from langgraph.types import Command
APPROVAL_THRESHOLD_IDR = 50_000_000
def needs_approval(request: ToolCallRequest) -> bool:
lines = request.tool_call["args"].get("lines", [])
return sum(line["debit"] for line in lines) >= APPROVAL_THRESHOLD_IDR
hitl = HumanInTheLoopMiddleware(
interrupt_on={
"post_journal_entry": {
"allowed_decisions": ["approve", "edit", "reject"],
"when": needs_approval, # below the threshold: no pause
},
"lookup_invoice": False, # read-only, never pauses
},
description_prefix="Journal entry awaiting finance approval",
)
agent = create_agent(
model="openai:gpt-5.4-mini",
tools=[lookup_invoice, post_journal_entry],
context_schema=ErpContext,
middleware=[hitl],
checkpointer=InMemorySaver(), # use a persistent saver, e.g. AsyncPostgresSaver
)
config = {"configurable": {"thread_id": "ap-2026-10-0042"}}
result = agent.invoke(payload, config=config, context=ctx, version="v2")
for interrupt in result.interrupts:
for action in interrupt.value["action_requests"]:
queue_for_review(action["name"], action["arguments"], thread="ap-2026-10-0042")
# Later, when the reviewer clicks reject on the approval screen.
# One decision per paused action, in the same order as action_requests.
agent.invoke(
Command(resume={"decisions": [{
"type": "reject",
"message": (
"Rejected by finance: wrong cost centre. Ask the user which cost "
"centre to use. Do not post this entry again unchanged."
),
}]}),
config=config,
context=ctx,
version="v2",
)The reject decision accepts a message that is returned to the model as the tool result. Make it explicit about what happens next, abandon, ask the user, or try another route, because the default text only says the tool was not run. Do not use respond to deny a side-effecting tool: the documentation warns that its text is treated as a successful tool result.
Decisions are a list with one entry per paused action, in the same order as action_requests. A batch of three proposed postings therefore needs three decisions, and a mismatch raises a ValueError rather than guessing.
Put together, the middleware list is not a set; its order decides who sees what first. The table summarises where each piece sits in the example below and why.
| Middleware | Hook | Why it sits there |
|---|---|---|
| PIIMiddleware entries | before_model, after_model | First in the list, so input is scrubbed before any other hook or model sees it |
| ModelFallbackMiddleware | wrap_model_call | Listed before retry, so it is the outer wrapper and only acts once retries are exhausted |
| ModelRetryMiddleware | wrap_model_call | Inside the fallback, with on_failure error so the fallback can see the failure |
| closed_period_guard | wrap_tool_call | Runs at execution, after approval, as the last deterministic check before the ERP |
| ToolCallLimitMiddleware for postings | after_model | Listed after the approval middleware, so its after_model runs first and an over-budget run stops before anyone is asked to approve |
agent = create_agent(
model="openai:gpt-5.4-mini",
tools=[lookup_invoice, search_vendor, lookup_exchange_rate, post_journal_entry],
system_prompt="You are an accounts-payable assistant for PT Contoh.",
context_schema=ErpContext,
checkpointer=InMemorySaver(),
middleware=[
*pii, # before_model, first
ModelCallLimitMiddleware(run_limit=12),
ModelFallbackMiddleware("anthropic:claude-sonnet-4-6"), # outermost wrap
ModelRetryMiddleware(max_retries=2, on_failure="error"), # inside the fallback
ToolRetryMiddleware(tools=["lookup_exchange_rate"], retry_on=(ConnectionError, TimeoutError)),
closed_period_guard, # at execution time
hitl, # after_model
ToolCallLimitMiddleware( # after_model, runs
tool_name="post_journal_entry", # before hitl because
run_limit=1, # after hooks run
exit_behavior="error", # last to first
),
],
)One trade-off remains visible here. The closed-period guard runs only when the tool executes, which is after a reviewer has approved the call, so a reviewer can approve an entry the guard then refuses. That is acceptable for a rare case, and the refusal still reaches the model; if it becomes common, run the same date check inside the when predicate as well, so a closed-period posting never reaches the approval queue.
The rule to carry from LangChain 1.x is that an agent is now a model, a tool list and an ordered middleware list, and the order is part of the design. Use the built-in classes for PII, limits, retries, fallback and approval, write wrap_tool_call guards for business rules, and check every on_failure and exit_behavior default before trusting the stack.