AI
OpenAI Agents SDK Sessions: SQLite, Redis and Postgres Memory
October 202612 min read

A session is an object that stores conversation history for the Agents SDK runner. Before each run the runner reads the stored items and prepends them to the input, and after the run it appends every new message, tool call and tool output. Any class with a session_id and the async methods get_items, add_items, pop_item and clear_session can act as a session.
Use a backend every worker can reach: RedisSession, SQLAlchemySession on Postgres or MySQL, MongoDBSession, DaprSession or the server-side OpenAIConversationsSession. SQLiteSession and AsyncSQLiteSession store history in a local file, so each container or host ends up with its own copy. Pick the shared backend you already operate and back up.
No. The Agents SDK does not allow a session in the same run as the server-managed options conversation_id, previous_response_id or auto_previous_response_id. Sessions are client-managed memory, and mixing the two layers can duplicate context. Choose one persistence strategy per conversation.
Conversation objects and their items are not subject to the 30-day TTL that applies to stored Response objects, so they persist until you delete them. Standalone responses are kept for 30 days by default unless you set store to false. If your policy requires deleting chat logs after a fixed period, you have to run that deletion against the OpenAI API yourself.
Wrap any session in EncryptedSession, installed with the openai-agents encrypt extra. It derives a per-session Fernet key from your master key with HKDF, using the session id as salt, and can ignore items older than an optional ttl. Keep the master key in a secret manager, because changing it without re-encrypting makes existing history unreadable.

Key Takeaway
OpenAI Agents SDK sessions store conversation history so the runner can prepend it to every turn. SQLiteSession suits one process; RedisSession and SQLAlchemySession on Postgres suit several workers; OpenAIConversationsSession keeps history on OpenAI with no 30-day expiry. EncryptedSession and OpenAIResponsesCompactionSession wrap any backend to add encryption at rest and compaction.
The bug report for a session problem almost never mentions sessions. It says the assistant forgot the purchase order the user mentioned two messages ago, and only sometimes. The agent worked perfectly on a laptop with SQLiteSession and a local file, then went behind a load balancer with three replicas, and each replica kept its own conversations.db. A user whose second request landed on a different container was talking to an agent with no memory of them.
This guide is about choosing the memory backend for OpenAI Agents SDK sessions in a real deployment rather than a notebook. It covers what the Session protocol actually does, a comparison table of every built-in and extension backend, production configuration for RedisSession and SQLAlchemySession on Postgres, the retention rules behind the server-side OpenAIConversationsSession, and the compaction and encryption wrappers that sit on top of any of them. Every class name, argument and default comes from the Python SDK documentation, version 0.22.x at the time of writing.
A session is a small contract, not a framework. Any object with a session_id and four async methods, get_items, add_items, pop_item and clear_session, is a session. When you pass one to Runner.run, the runner reads the stored items before the call and prepends them to the input, and after the call it appends every new item the run produced: user messages, assistant messages, tool calls and tool outputs. You stop maintaining result.to_input_list() by hand, and different agents can share one session, which is what makes handoffs feel continuous.
from agents import Agent, Runner, SQLiteSession
agent = Agent(name="ERP helpdesk", instructions="Answer questions about purchase orders.")
# Wrong for anything long-lived: no db_path means an in-memory database,
# so every conversation disappears when the process restarts.
session = SQLiteSession("user_123")
# Right for a single-process tool, a CLI or a notebook: a file on disk.
session = SQLiteSession("user_123", "conversations.db")
await Runner.run(agent, "What is the status of PO-2026-0412?", session=session)
# Before this second call the runner reads the history and prepends it;
# after it, every new item (messages, tool calls, tool outputs) is appended.
await Runner.run(agent, "And who approved it?", session=session)
# Wrong: a session plus server-managed continuation in the same run.
# The SDK does not allow it; pick one persistence strategy per conversation.
await Runner.run(agent, "...", session=session, previous_response_id="resp_...")Two rules follow from that design. First, the default SQLiteSession with no file path is an in-memory database, so it is lost when the process ends; it is a test fixture, not storage. Second, a session is client-managed memory, and the SDK does not let you combine it in the same run with the server-managed options conversation_id, previous_response_id or auto_previous_response_id. The docs are blunt about why: mixing client-managed history with OpenAI-managed state can duplicate context. Decide where a conversation lives once, per conversation, and the rest of this post is about making that decision well.
The SDK ships three sessions in its core package and a longer list under agents.extensions.memory. They differ less in API, since all of them satisfy the same four methods, than in where the bytes live and therefore what happens when you run more than one worker.
| Session class | Install | Where history lives | Safe with several workers | Expiry and extras |
|---|---|---|---|---|
| SQLiteSession | Core package | In memory, or a local SQLite file when you pass db_path | No. Each container or host gets its own file | None. Trusts the file completely and does not detect external edits |
| AsyncSQLiteSession | aiosqlite | A local SQLite file, accessed asynchronously | No, for the same reason | None. Avoids blocking the event loop in async apps |
| RedisSession | openai-agents[redis] | A Redis list of items plus a metadata hash, under key_prefix | Yes. Every worker talks to the same server | Optional ttl in seconds. Durability depends on your Redis persistence settings |
| SQLAlchemySession | openai-agents[sqlalchemy] | agent_sessions and agent_messages tables in Postgres, MySQL or SQLite | Yes, on a networked database. Writes run in a transaction | No built-in expiry. You own retention, backups and migrations |
| MongoDBSession | openai-agents[mongodb] | agent_sessions and agent_messages collections, one batch document per add_items call | Yes | Ordering by a monotonically increasing seq field. ping() checks connectivity |
| DaprSession | openai-agents[dapr] | Whichever Dapr state store you configure | Yes | Optional ttl and a strong-consistency option |
| OpenAIConversationsSession | Core package | The OpenAI Conversations API, on OpenAI's servers | Yes. Workers only need the conversation id | Not subject to the 30-day response TTL. Cannot be wrapped in the compaction session |
| AdvancedSQLiteSession | agents.extensions.memory | A local SQLite file with turns, branches and usage data | No | Branching from any turn and per-turn token usage |
Read the fourth column first. If the answer is no, the backend is fine for a CLI, a single long-running worker or a test suite, and wrong for anything that scales horizontally. Among the ones that say yes, the choice is mostly about what you already operate. A team with a Postgres database and a migration tool gets the most from SQLAlchemySession; a team already running Redis for caching or queues gets the lowest latency from RedisSession; a team that would rather operate nothing reaches for OpenAIConversationsSession and accepts that history lives on OpenAI.
SQLiteSession is not a toy, it is a single-process store, and there are plenty of single-process agents. It is the right backend in three situations:
Outside those cases, the failure is the one in the introduction: the history is real, it is just on a different disk. Sticky sessions at the load balancer hide it until a deploy or a crash moves a user to a fresh container. Note also the trust boundary the SDK documents: SQLiteSession assumes the application trusts the file and its storage, and does not authenticate rows or detect edits, deletions, reordering or replay. Anything that can write to that file can rewrite what the agent believes was said. If you are on SQLite and async, use AsyncSQLiteSession so history reads do not block the event loop.
RedisSession stores each conversation as a Redis list of serialised items, with a hash beside it for session_id, created_at and updated_at. Keys are namespaced by key_prefix, which defaults to agents:session, and an optional ttl in seconds lets Redis remove abandoned conversations without a cleanup job. The important decision is who owns the client. In a web app, create one async client per worker at startup and inject it, rather than calling from_url on every request.
# pip install "openai-agents[redis]"
from contextlib import asynccontextmanager
import redis.asyncio as redis
from fastapi import FastAPI, Request
from agents import Agent, Runner
from agents.extensions.memory import RedisSession
@asynccontextmanager
async def lifespan(app: FastAPI):
# One client and connection pool per worker, owned by the app, not the session.
app.state.redis = redis.from_url("redis://redis:6379/0")
yield
await app.state.redis.aclose()
app = FastAPI(lifespan=lifespan)
agent = Agent(name="ERP helpdesk")
@app.post("/chat/{conversation_id}")
async def chat(conversation_id: str, body: dict, request: Request):
user = request.state.user # set by your auth middleware
# Wrong: RedisSession(conversation_id, ...) lets any caller who guesses
# an id read someone else's history. Scope the key to the identity.
session = RedisSession(
f"{user.tenant}:{user.id}:{conversation_id}",
redis_client=app.state.redis, # injected, so close() is a no-op
key_prefix="erp:helpdesk", # default is "agents:session"
ttl=60 * 60 * 24 * 7, # time-to-live, in seconds
)
result = await Runner.run(agent, body["message"], session=session)
return {"reply": result.final_output}The session id is the other decision, and it is a security one. A conversation id that arrives in a URL is user input. If it becomes the session id directly, anyone who can guess or replay an id reads another user's history, including every tool output the agent fetched on their behalf, which in an ERP context means invoice amounts and supplier names. Build the id from the authenticated tenant and user plus the conversation, as above. Before relying on ttl as a sliding expiry for idle chats, check in your own Redis whether the key's TTL is refreshed on each write, because the docs only specify ttl as a time-to-live in seconds.
RedisSession.from_url creates and owns its Redis client. After close(), that session is terminal and every later operation raises RuntimeError. A session constructed with redis_client=... behaves the opposite way: close() does nothing and the caller keeps ownership. Mixing the two in one codebase is how a shutdown hook ends up breaking live requests.
SQLAlchemySession is the backend I would pick for an ERP back office, because the conversations end up next to the data they are about, under the same backups, access controls and retention policy. It works with any database SQLAlchemy supports through an async driver; the sqlalchemy extra already includes asyncpg for URLs starting with postgresql+asyncpg, MySQL needs aiomysql, and SQLite needs aiosqlite.
# pip install "openai-agents[sqlalchemy]" (asyncpg comes with the extra)
from sqlalchemy.ext.asyncio import create_async_engine
from agents.extensions.memory import SQLAlchemySession
# One engine per worker process. Building it per request opens a new pool
# per request, and Postgres runs out of connections long before you run out of users.
engine = create_async_engine(
"postgresql+asyncpg://agents:***@db:5432/erp",
pool_size=10,
pool_pre_ping=True,
)
def session_for(tenant: str, user_id: str, conversation_id: str) -> SQLAlchemySession:
return SQLAlchemySession(
f"{tenant}:{user_id}:{conversation_id}",
engine=engine,
create_tables=False, # the constructor default: tables come from migrations
sessions_table="agent_sessions",
messages_table="agent_messages",
ensure_ascii=False, # keep "Faktur sudah disetujui" readable in the JSON column
)
# Note: SQLAlchemySession.from_url(...) builds its own engine. Fine for a script,
# wasteful in a web app. On shutdown: await engine.dispose()Three arguments in that block matter more than they look. create_tables defaults to False in the constructor, deliberately, because in production the agent_sessions and agent_messages tables should come from your migrations, not from whichever pod starts first; the docs' quick-start examples pass True for convenience. ensure_ascii defaults to True to preserve the historical storage format, which stores Indonesian or any non-ASCII text as escape sequences, so set it to False if anyone will ever query the JSON. And reuse one engine: from_url builds a new engine each time, which is fine in a script and a connection leak in a request handler. On the write path, add_items inserts the batch inside a transaction that locks the session, and pop_item removes the newest item with DELETE ... RETURNING where the database supports it, so two concurrent writes to one conversation do not interleave in the middle of a batch.
For multi-tenant systems, a custom session can accept a keyword-only wrapper argument on all four methods and receive the run's RunContextWrapper. That lets one session class route each tenant to its own schema or database, or refuse access, using the same context object your tools already read.
OpenAIConversationsSession stores history in the OpenAI Conversations API instead of your infrastructure. Construct it with no argument to create a conversation, or pass a conversation_id to resume one; every worker then needs only that id. Conversations store messages, tool calls, tool outputs and other items as a durable object.
Retention is where it differs from everything else in the table, and it is easy to get backwards. Response objects are saved for 30 days by default, and you can disable that with store set to false. Conversation objects and their items are not subject to the 30-day TTL, and any response attached to a conversation has its items persisted with no 30-day TTL. So the server-side option is not the short-lived one; it keeps history until you delete it. If your data-protection policy says chat logs are deleted after 90 days, that deletion is now a job you run against the OpenAI API, not a cron against your database.
Cost is the other consideration. With server-managed state you send only the new turn, but the docs note that even with previous_response_id, all previous input tokens in the chain are billed as input tokens. Server-side storage saves bandwidth and code, not tokens. What controls the token bill is how much history reaches the model, which is the job of the wrappers and limits in the next two sections.
Two wrappers take any session as underlying_session and are themselves sessions, so they compose. OpenAIResponsesCompactionSession calls the Responses API compaction endpoint to replace long history with a compacted window, automatically after a turn according to should_trigger_compaction, or on demand with run_compaction. EncryptedSession encrypts each item before it reaches the underlying store, with an optional ttl.
# pip install "openai-agents[encrypt,sqlalchemy]"
import os
from agents import Runner
from agents.extensions.memory import EncryptedSession, SQLAlchemySession
from agents.memory import OpenAIResponsesCompactionSession
base = SQLAlchemySession(session_id, engine=engine)
encrypted = EncryptedSession(
session_id=session_id,
underlying_session=base,
# Master key. HKDF derives a separate Fernet key per session, salted
# with the session id, so one leaked row decrypts nothing else.
encryption_key=os.environ["AGENT_SESSION_KEY"],
ttl=60 * 60 * 24 * 30, # items older than 30 days are silently ignored on read
)
session = OpenAIResponsesCompactionSession(
session_id=session_id,
underlying_session=encrypted,
# Auto-compaction waits before the run completes and can stall a stream.
# Disable it on the hot path and compact from a background job instead.
should_trigger_compaction=lambda _: False,
)
result = await Runner.run(agent, message, session=session)
# Later, from an idle hook or a worker, through the same wrapper:
await session.run_compaction({"force": True})
# Wrong: base.add_items(...) while compaction runs. Direct writes to the
# underlying session bypass the wrapper's serialisation and recovery.The details that decide whether this works in production are all in the docs and easy to miss. Auto-compaction can block streaming, because it runs before the run completes; on a chat endpoint, disable it and compact on idle or every N turns. compaction_mode defaults to auto, which falls back to rebuilding the request from session items when the agent runs with store set to false, and input mode makes the session contents the source of truth. The wrapper serialises add_items, pop_item and clear_session against a running compaction and skips a stale compaction if another run changed the history first, but only for writes that go through the wrapper. Do not wrap OpenAIConversationsSession in it, since the two manage history in different ways. For encryption, EncryptedSession derives a 32-byte Fernet key per session with HKDF, using your master key and the session id as salt, and items older than ttl are silently ignored on read rather than raising.
Install the wrappers with the encrypt extra, for example pip install openai-agents[encrypt,sqlalchemy] for encrypted Postgres sessions. Keep the master key in a secret manager, not in the repo: rotating it without re-encrypting makes every existing conversation unreadable.
Storing everything and sending everything are separate decisions. SessionSettings with a limit caps how many recent items the session returns, and RunConfig accepts a session_input_callback that receives the retrieved history and the new input and returns exactly what goes to the model. AdvancedSQLiteSession adds what the others do not have: turn listing, content search, branching from a past turn and per-turn token usage.
from agents import RunConfig, Runner, SessionSettings
from agents.extensions.memory import AdvancedSQLiteSession, SQLAlchemySession
# Cap what is read back per run, not what is stored.
session = SQLAlchemySession(sid, engine=engine, session_settings=SessionSettings(limit=40))
def keep_recent_history(history, new_input):
# history = what the session returned; new_input = this turn.
# The return value is exactly what the model sees.
return history[-10:] + new_input
result = await Runner.run(
agent, message, session=session,
run_config=RunConfig(session_input_callback=keep_recent_history),
)
# Branching and per-turn token accounting: AdvancedSQLiteSession only.
review = AdvancedSQLiteSession(
session_id="po-review-77", db_path="review.db", create_tables=True
)
result = await Runner.run(agent, "Draft the approval note for PO-77", session=review)
await review.store_run_usage(result)
turns = await review.get_conversation_turns()
await review.create_branch_from_turn(2) # try another answer from turn 2
result = await Runner.run(agent, "Make it stricter on budget", session=review)
await review.switch_to_branch("main")
print(await review.get_session_usage())Trim on turn boundaries, not raw item counts, when you can. One user turn can produce a message, several tool calls, their outputs and reasoning items, and the SDK documents errors in multi-turn runs when a reasoning item is separated from the item it must stay paired with. Test any callback against a conversation with tool calls before shipping it. Branching is narrower than it sounds: it is an AdvancedSQLiteSession feature, so it fits review tools and evaluation harnesses, where you want to replay turn two with a different instruction and compare token cost with get_session_usage, more than it fits a horizontally scaled chat backend.
The rule I carry is to choose a session by deployment shape first and features second. One process means SQLiteSession; several workers mean RedisSession or SQLAlchemySession on a database you already back up; no infrastructure means OpenAIConversationsSession with retention you delete yourself. Then build the session id from the authenticated user, add EncryptedSession when history holds business data, and compact off the hot path.