Backend
Durable AI Agents With Temporal: Survive Crashes and Waits
October 202612 min read

A durable AI agent is one whose progress is stored outside the process running it, so a crash, restart or long wait resumes on the same step instead of starting over. With Temporal, the agent loop runs in a workflow and every model call and tool call runs as an activity whose result is recorded in event history. On recovery, completed steps return their recorded results rather than running again.
Install the temporalio-openai-agents package, register OpenAIAgentsPlugin on the Temporal Client, and call Runner.run inside a workflow as usual. The plugin turns each model call into an activity. Tools that do I/O must be wrapped with activity_as_tool and their activity functions registered on the worker.
In Temporal, the workflow waits with workflow.wait_condition on a value set by a signal or update handler, with an optional timeout such as three days. The wait is a server-side timer, so no worker holds it in memory and it survives restarts and deploys. When the approval signal arrives, the workflow continues from that point.
Temporal replays in-flight workflows with the new code, so a change to the sequence of steps can cause a non-determinism error. Gate new steps behind workflow.patched, or use Worker Versioning to pin executions to the code that started them. Replay tests against recorded histories catch the changes that need a patch before they ship.
Temporal suits agents that are one step in a larger business process needing timers, retries and audit history, because it records every activity and signal. LangGraph checkpointers save graph state per super-step and are lighter to operate if the agent already runs as a graph. For ERP writes on LangGraph, use the sync durability mode so each checkpoint is written before the next step.

Key Takeaway
Durable AI agents on Temporal keep the agent loop in a deterministic workflow and push every model call and tool call into activities, whose results are recorded in event history. A crashed worker replays that history instead of repeating work, an approval can wait for days on a durable timer, and patch markers let in-flight agents survive a deploy.
Consider a purchase request agent in an ERP. It reads the request, checks the cost centre's remaining budget, writes a summary for the department manager and then waits for a decision. The manager is at a supplier site until Thursday. On Tuesday night the team deploys a new release, the worker pod restarts, and the in-memory await that was holding the agent's place is gone. Nobody notices until the requester asks why the PO was never created.
That is the problem durable AI agents with Temporal solve: the agent's position is stored by the Temporal server, not in a process, so a crash, a restart or a three-day wait resumes on the same step. This post explains the workflow and activity split, implements the purchase approval with the temporalio-openai-agents plugin for the OpenAI Agents SDK, shows how to ship a new release while approvals are still open, and compares the same job in Pydantic AI and LangGraph. Every API name below comes from the official Temporal, Pydantic, OpenAI and LangChain documentation listed at the end.
An agent run is a loop of model calls and tool calls, and each step depends on the results of the previous ones. Held in a Python coroutine, that loop has the lifetime of the process. If the worker crashes after the third tool call, a naive retry starts from the top: it calls the model again, gets a slightly different plan, reads the ERP again and may write to it twice. If the loop has to wait for a person, the process has to stay up for as long as the person takes, which nobody can promise.
Temporal's guarantee is narrower and more useful than retry. As the workflow makes progress, the server records each scheduled activity and its result in the workflow's event history. When a worker dies, another worker replays the workflow code against that history: completed model calls and tool calls return their recorded results instead of running again, and execution continues from the first step that has no result yet. A wait is a timer or a condition on the server, so a workflow idle for three days consumes no worker at all.
Everything rests on one rule from the Temporal documentation: workflow code must be deterministic, because it is re-executed on replay, while activities may do any I/O but restart from the beginning if they fail part-way. For an agent, that sorts every line of code into one of two places.
| Piece of the agent | Where it runs | Why |
|---|---|---|
| Model call | Activity, automatically | A network call with a non-deterministic answer. The plugin routes every Runner.run model request through an activity, so a replay reads the recorded response instead of asking the model again. |
| ERP read tool | Activity, via activity_as_tool | Database or HTTP I/O. A plain function tool would run inside the workflow, where I/O is not allowed. |
| ERP write, such as converting a PR to a PO | Activity, called by the workflow, not by the model | The side effect that must happen exactly once, so it needs an idempotency key and should be a decision of the workflow, not of the model. |
| Pure computation, such as a tax total | Workflow, as a function tool | Deterministic and free of I/O, so it can safely run again on every replay. |
| Waiting for a manager | Workflow, through a signal and wait_condition | The wait and its timeout live on the Temporal server, so the worker can restart or be redeployed during it. |
| Clock, random numbers, environment | Workflow APIs or an activity | Reading datetime.now() or an environment variable in workflow code gives a different answer on replay, which is exactly the non-determinism replay cannot tolerate. |
The third row is the design decision that matters most in an ERP. The model is allowed to read and to summarise, but the step that creates a purchase order is called by the workflow after a person has approved it. That keeps the irreversible action out of the model's hands and gives it one obvious place for an idempotency key.
The integration now ships as its own package. Temporal's documentation says version 1.0.0 moved it out of the Python SDK: the temporalio[openai-agents] extra became the temporalio-openai-agents distribution, and imports moved from temporalio.contrib.openai_agents to temporalio.openai_agents. It needs Python 3.10 or later and Temporal Python SDK 1.33.0 or later. Start with the activities, which are ordinary Temporal activities with no agent code in them.
# uv add temporalio-openai-agents
# (1.0.0 moved the plugin out of temporalio.contrib.openai_agents)
# activities.py: everything that touches the network lives here
from dataclasses import dataclass
from temporalio import activity
@dataclass
class PurchaseRequest:
pr_number: str
requester: str
total_idr: int
cost_center: str
@activity.defn
async def get_purchase_request(pr_number: str) -> PurchaseRequest:
"""Read one purchase request from the ERP. Read-only."""
return await erp.fetch_pr(pr_number)
@activity.defn
async def get_budget_remaining(cost_center: str) -> int:
"""Remaining budget for a cost centre this period, in rupiah."""
return await erp.budget_remaining(cost_center)
@activity.defn
async def notify_approver(pr_number: str, summary: str) -> None:
await chat.send_to_manager(pr_number, summary)
@activity.defn
async def convert_pr_to_po(pr_number: str, approver: str) -> str:
# A retried activity runs again from the top, so the ERP call takes
# the PR number as its idempotency key and returns the existing PO.
return await erp.create_po_from_pr(
pr_number, approved_by=approver, idempotency_key=pr_number
)The workflow holds the agent. Inside it you write plain OpenAI Agents SDK code, because the plugin redirects Runner.run so each model call becomes an activity; there is no Temporal-specific runner. Tools that perform I/O are wrapped with activity_as_tool, which exposes the activity to the model as a tool and schedules the activity whenever the model calls it.
# workflows.py: deterministic orchestration, no I/O
import asyncio
from dataclasses import dataclass
from datetime import timedelta
from temporalio import workflow
from temporalio.openai_agents.workflow import activity_as_tool
with workflow.unsafe.imports_passed_through():
from agents import Agent, Runner
from activities import (
convert_pr_to_po,
get_budget_remaining,
get_purchase_request,
notify_approver,
)
@dataclass
class Decision:
approver: str
approved: bool
note: str = ""
APPROVAL_WINDOW = timedelta(days=3)
@workflow.defn
class PurchaseApprovalWorkflow:
def __init__(self) -> None:
self.decision: Decision | None = None
@workflow.signal
def decide(self, decision: Decision) -> None:
self.decision = decision
@workflow.query
def status(self) -> str:
return "decided" if self.decision else "waiting for approver"
@workflow.run
async def run(self, pr_number: str) -> str:
reviewer = Agent(
name="pr-reviewer",
instructions=(
"Summarise the purchase request for a manager in Bahasa Indonesia: "
"total, cost centre, and whether budget remains. "
"Do not recommend approval or rejection."
),
tools=[
activity_as_tool(get_purchase_request,
start_to_close_timeout=timedelta(seconds=15)),
activity_as_tool(get_budget_remaining,
start_to_close_timeout=timedelta(seconds=15)),
],
)
# Ordinary Agents SDK code. The plugin turns every model call into an
# activity, and activity_as_tool does the same for each tool call, so
# a replay after a crash reads their results from history.
review = await Runner.run(reviewer, input=f"Review {pr_number}")
await workflow.execute_activity(
notify_approver,
args=[pr_number, review.final_output],
start_to_close_timeout=timedelta(seconds=30),
)
try:
# A durable timer on the Temporal server. No worker holds this
# wait in memory, so it survives restarts and deploys.
await workflow.wait_condition(
lambda: self.decision is not None, timeout=APPROVAL_WINDOW
)
except asyncio.TimeoutError:
return f"{pr_number}: no decision in {APPROVAL_WINDOW.days} days, escalated"
if not self.decision.approved:
return f"{pr_number}: rejected by {self.decision.approver}"
po_number = await workflow.execute_activity(
convert_pr_to_po,
args=[pr_number, self.decision.approver],
start_to_close_timeout=timedelta(seconds=60),
)
return f"{pr_number}: approved, {po_number} created"Three details carry the durability. The agent's reads go through activities, so a crash after get_budget_remaining replays the recorded budget rather than reading it again. The wait is wait_condition with a three-day timeout, which becomes a server-side timer. And the PO is created by the workflow only after a signal carries a decision, with the PR number as the ERP idempotency key so that a retried activity returns the PO it already created.
The worker and every client that starts or signals the workflow must use the same OpenAIAgentsPlugin, because the plugin also configures the Pydantic data converter that serialises payloads on both sides. ModelActivityParameters controls the model activity: start_to_close_timeout defaults to 60 seconds, and it also accepts a retry policy, heartbeat timeout and task queue. activity_as_tool only describes how the agent invokes an activity, so the activity functions must still be listed on the worker.
# worker.py: the same plugin goes on every Client, worker and caller alike
from datetime import timedelta
from temporalio.client import Client
from temporalio.openai_agents import ModelActivityParameters, OpenAIAgentsPlugin
from temporalio.worker import Worker
PLUGIN = OpenAIAgentsPlugin(
model_params=ModelActivityParameters(start_to_close_timeout=timedelta(seconds=60))
)
client = await Client.connect("localhost:7233", plugins=[PLUGIN])
worker = Worker(
client,
task_queue="erp-approvals",
workflows=[PurchaseApprovalWorkflow],
# activity_as_tool does not register activities; list them here too.
activities=[get_purchase_request, get_budget_remaining,
notify_approver, convert_pr_to_po],
)
await worker.run()
# api.py: one workflow per PR, and the PR number is the workflow id
await client.start_workflow(
PurchaseApprovalWorkflow.run,
"PR-2026-10-0042",
id="pr-approval-PR-2026-10-0042",
task_queue="erp-approvals",
)
# Two days later, from the manager's Approve button, in any process:
handle = client.get_workflow_handle("pr-approval-PR-2026-10-0042")
await handle.signal(
PurchaseApprovalWorkflow.decide,
Decision(approver="budi.santoso", approved=True),
)Using the PR number as the workflow id gives the approval a business key. The approve button needs nothing but that id to find the waiting execution, and the status query lets the ERP screen show waiting for approver without touching the workflow's state. If you need the caller to get a result back from the same call, Temporal also offers Updates, which the integration docs use for chat turns that return the agent reply; a signal is enough for a fire-and-forget approval.
The Agents SDK has its own approval flow, needs_approval on a tool, and inside Temporal the documented hook for hosted MCP tools is an on_approval_request callback that runs in workflow context and can wait on a signal or update. For an ERP approval, prefer the explicit pattern above: the model finishes its summary, the workflow waits, and the write happens outside the model loop. It is easier to audit and to test.
Durability creates an obligation. A workflow started last week will be replayed by this week's code, so any change that alters the sequence of steps an in-flight execution would take breaks its replay. Temporal's versioning guide offers two tools: patched(), which records a marker in history so old executions keep the old path while new ones take the new path, and Worker Versioning, which pins executions to the deployment version of the worker code that started them.
# The new release adds a vendor blacklist check before the approval wait.
# Wrong: insert the step directly. A workflow that started last week and
# is parked in wait_condition replays a history with no such activity,
# and the worker fails that workflow task with a non-determinism error.
await workflow.execute_activity(
check_vendor_blacklist, pr_number, start_to_close_timeout=timedelta(seconds=15)
)
# Right: gate the new step behind a patch marker. Old histories take the
# old path; new executions record the marker and take the new one.
if workflow.patched("vendor-blacklist-check"):
await workflow.execute_activity(
check_vendor_blacklist, pr_number, start_to_close_timeout=timedelta(seconds=15)
)
# Once no pre-patch execution is still open, swap patched() for
# workflow.deprecate_patch("vendor-blacklist-check"), then remove it.The patch lifecycle has three steps, and the guide is explicit about the order: deploy with patched(), switch to deprecate_patch() once no pre-patch execution is still open, and remove both calls once those histories are out of retention. For approval workflows with a timeout of days, that is a few deploys, not a few hours. Temporal also recommends replay testing, running recorded histories against new code before it ships, to find the changes that need a patch.
Names are part of the contract too. Activity names come from the function names, and in Pydantic AI the documentation says the agent name and toolset ids are required under Temporal and must not change once deployed, because renaming them breaks active workflows. Rename a tool in the same release that drains its last open approval, never before.
Pydantic AI supports Temporal natively and follows the same split: model requests, tool calls that may need I/O and MCP server communication become activities, while the agent run lives in the workflow. The current API is a capability. You attach TemporalDurability to an ordinary Agent, list the agent on a PydanticAIWorkflow subclass and connect with PydanticAIPlugin. The older TemporalAgent wrapper is deprecated and will be removed in v3; the docs say workflows started under it replay correctly after switching, as long as the agent name, toolset ids and model registry keys stay the same.
# uv add "pydantic-ai[temporal]"
from datetime import timedelta
from pydantic_ai import Agent
from pydantic_ai.durable_exec.temporal import (
PydanticAIPlugin,
PydanticAIWorkflow,
TemporalDurability,
)
from temporalio import workflow
from temporalio.workflow import ActivityConfig
reviewer = Agent(
"openai:gpt-5.6-sol",
# Required under Temporal: activity names derive from the agent name.
# Renaming it in production breaks every workflow still in flight.
name="pr-reviewer",
instructions="Summarise the purchase request for a manager. Do not recommend.",
tools=[get_purchase_request, get_budget_remaining], # plain async functions
capabilities=[
TemporalDurability(
activity_config=ActivityConfig(start_to_close_timeout=timedelta(seconds=60))
)
],
)
@workflow.defn
class PurchaseApprovalWorkflow(PydanticAIWorkflow):
__pydantic_ai_agents__ = [reviewer] # the plugin registers their activities
@workflow.run
async def run(self, pr_number: str) -> str:
# Must be the async API: run_sync() inside a workflow raises UserError.
review = await reviewer.run(f"Review {pr_number}")
... # same notify / wait_condition / convert_pr_to_po steps as before
client = await Client.connect("localhost:7233", plugins=[PydanticAIPlugin()])Two differences from the OpenAI plugin decide how the code looks. Plain async function tools are routed through activities automatically, so there is no activity_as_tool step. In exchange, everything crossing into an activity must be serialisable by Pydantic, including deps, model_settings and tool metadata, and Temporal caps each payload at 2 MB by default. The default activity timeout is 60 seconds unless you pass activity_config, and the docs recommend turning off the provider client's own retries, for example max_retries set to 0, so Temporal's retry policy is not multiplied by two more layers.
Temporal is not the only way to make an agent outlive its process. LangGraph persists graph state per thread through its checkpointers, which its docs list for human-in-the-loop and fault tolerance, and the Agents SDK can serialise a paused run on its own. They differ in what is saved, how often, and what you operate.
| Approach | What is persisted | After a crash mid-run | How it waits for a person |
|---|---|---|---|
| OpenAI Agents SDK on Temporal | Every model call, tool activity, signal and timer in the event history | Replays to the first step without a recorded result | Signal or update plus wait_condition with a server-side timeout |
| Pydantic AI with TemporalDurability | The same event history, with Pydantic-serialised payloads up to 2 MB each by default | Same replay behaviour; failed activities retry under the configured policy | Workflow signal or update around agent.run |
| LangGraph with a Postgres checkpointer | Graph state at each super-step, plus per-node pending writes | Resumes from the last saved checkpoint; how recent depends on the durability mode | interrupt in a node, resumed with Command and the same thread_id |
| Agents SDK RunState alone | A JSON snapshot of a paused run: pending tool calls, arguments and approval decisions | Only what you saved at the last interruption; a crash between interruptions loses the run | needs_approval on the tool, then approve or reject on the restored state |
In practice the choice follows the shape of the job. If the agent is one step in a business process that already needs timers, retries and audit history, Temporal gives you all of it from one system. If the agent is the process and runs in a graph you already maintain, a LangGraph checkpointer in sync mode is less to operate. A bare RunState snapshot suits a single approval pause, but the OpenAI docs warn that it is not authenticated: load it only from storage you trust, because it carries tool arguments.
# LangGraph: durability is a per-call setting on a checkpointed graph.
# "exit" saves only when the run exits or interrupts: a crash mid-run loses it
# "async" saves while the next step runs: a small window can be lost
# "sync" saves before the next step starts: the one to use for ERP writes
graph.invoke(inputs, {"configurable": {"thread_id": "PR-2026-10-0042"}}, durability="sync")
# OpenAI Agents SDK without Temporal: a paused run is a JSON snapshot you own.
state = result.to_state()
db.save("PR-2026-10-0042", state.to_json()) # pending tool call + arguments
# ...days later, in another process...
state = await RunState.from_json(agent, db.load("PR-2026-10-0042"))
state.approve(state.get_interruptions()[0])
result = await Runner.run(agent, state)Durable execution removes one class of failure and makes a few others easier to commit. Before an agent like this goes live in an ERP:
Durability is not authorisation. A replayed workflow will happily finish a PO conversion that was approved by the wrong person if the signal handler does not check. Validate who sent the decision, in the authenticated API in front of the signal, and record the approver in the decision itself, before it reaches workflow state.
The rule to carry is simple: whatever the agent must not forget goes into history, and whatever the agent must not repeat goes into an activity with an idempotency key. With the model calls and tools in activities, the wait on a server-side timer and new steps behind patch markers, a purchase approval can sit for three days across two deploys and resume on exactly the step where it stopped.
Sources and further reading