Backend
OpenAI Assistants API Shutdown Migration to the Responses API
October 202611 min read

OpenAI sunset the Assistants API on 26 August 2026, and it is no longer available. The deprecations page lists the Responses API and the Conversations API as the recommended replacement. Calls to the assistants, threads and runs endpoints no longer work.
OpenAI's migration guide maps Assistants to prompts, Threads to Conversations, Runs to Responses and Run steps to Items. Because reusable prompt objects are themselves scheduled to shut down on 30 November 2026, the safer replacement for an assistant is its model, instructions and tools kept as versioned configuration in your own code.
No. OpenAI provides no automatic converter, and the call that listed a thread's messages stopped working at the sunset. You have to rebuild history from messages your application stored itself, converting user text to input_text and assistant text to output_text, and adding at most 20 items per Conversations API call.
Vector stores are separate resources from the Assistants API, and the Responses API's file_search tool takes vector_store_ids directly. Instead of attaching them through tool_resources on an assistant or thread, you reference them in the tool definition on each request. List your stores first and confirm every id your code uses still resolves.
Response objects are kept for 30 days by default, and setting store to false avoids that. Conversation objects, their items and any response attached to a conversation are not subject to the 30-day TTL, so they persist until you delete them. Plan a deletion policy before you migrate customer data.

Key Takeaway
The OpenAI Assistants API was shut down on 26 August 2026 and replaced by the Responses and Conversations APIs. Migrate assistants into versioned config in code rather than prompt objects, which shut down on 30 November 2026, rebuild thread history from your own stored messages in batches of 20 items, and move the tool loop into application code.
The failure mode after a sunset is quiet. A support widget that has worked for a year starts returning a generic error, the logs fill with failed calls to /v1/threads and /v1/threads/runs, and the person who wrote the integration left the company. That is the position many teams are in this month: the Assistants API stopped answering on 26 August 2026, and the code that called it is still deployed.
This is the playbook I would follow to repair a live Assistants integration in October 2026. It is built from OpenAI's own migration guide, its deprecations page and the Conversations API reference, and it covers the four things that actually need rewriting: where the assistant's configuration lives, how thread history is rebuilt, how the run loop is replaced, and how file_search and code_interpreter move. It also flags two traps in the official guide that will cost you a second migration if you follow it literally.
OpenAI's migration guide states it plainly: the Assistants API was officially sunset on 26 August 2026 and is no longer available. The deprecations page lists the Responses API and the Conversations API as the recommended replacement. Everything under the beta assistants, threads, runs and run steps surface went with it, including the call that lists a thread's messages. That last point matters most, because it means the history stored inside OpenAI's threads can no longer be read back through the API.
What did not go is everything that was never part of the Assistants surface. Files and vector stores are their own resources, and the Responses API's file_search tool takes vector store ids directly, so the indexed knowledge base you built is the one asset that carries over by reference. Your migration is therefore lopsided: configuration and conversation state must be rebuilt, retrieval data mostly has to be re-pointed, and the orchestration that used to happen inside a run now has to be written by you.
OpenAI's guide reduces the change to four renames. The table adds what each one means in practice, and a column for what I would actually use today, because in two rows the official answer has already moved on.
| Assistants concept | Official replacement | What I would use in October 2026 | The trap |
|---|---|---|---|
| Assistant | Prompt object, created in the dashboard | Model, instructions and tools as a versioned module in your repo | Reusable prompt objects shut down on 30 November 2026 |
| Thread | Conversation | Conversation, with the legacy thread id kept in metadata | No automatic converter; history comes from your own store |
| Run | Response | One responses.create call per model turn | Instructions and tools are sent with each request |
| Run step | Item | Typed output items: message, function_call, file_search_call | Code that parsed run steps for citations must read annotations instead |
| requires_action and submit_tool_outputs | function_call and function_call_output items | An explicit tool loop with a turn cap | The loop is now yours, including the infinite-loop guard |
| file_search via tool_resources | file_search tool with vector_store_ids | The same vector stores, referenced on every request | Thread-level vector stores have no thread to hang on |
| code_interpreter via tool_resources | code_interpreter tool with a container | container type auto with file_ids | Generated files come back as container_file_citation annotations |
The direction of every row is the same: state and orchestration that OpenAI used to hold for you inside an assistant and a run now sit in your application. The guide frames this as separation of concerns, with your code handling history pruning, the tool loop and retries. For a backend team that is a fair trade, because each of those things becomes testable, but it is also why this migration is a rewrite of the call site and not a find-and-replace of endpoint names.
The first step of the official guide is to open each assistant in the dashboard and click Create prompt, which turns it into a reusable prompt object you reference by id. The same page now carries a note that reusable prompt objects are also being deprecated, and the deprecations page gives the dates: prompt creation was de-emphasised on 3 June 2026, and the v1/prompts API and reusable prompt objects are scheduled to shut down on 30 November 2026. Following step one literally buys you about two months.
If you already moved assistants to prompt objects in the dashboard, you have a second deadline on 30 November 2026. OpenAI's own guidance for that migration is to move the prompt content out of the managed object and into your application code, which is where it should go in the first place.
// Wrong: the dashboard "Create prompt" path. Each assistant becomes a
// reusable prompt object (pmpt_...), and v1/prompts is scheduled to shut
// down on 30 November 2026 — you would migrate the same config twice.
await client.responses.create({
prompt: { id: "pmpt_123", version: "3" },
conversation: conversationId,
input,
});
// Right: the assistant's instructions + tools live in a module you review,
// diff and roll back like any other code. src/agents/invoice-helper.ts
export const INVOICE_HELPER = {
version: "2026-10-01",
model: "gpt-6-astra",
instructions: [
"You answer questions about purchase invoices in the ERP.",
"Quote invoice numbers exactly. Never invent a due date.",
].join("\n"),
tools: [
{
type: "file_search",
vector_store_ids: [process.env.POLICY_VECTOR_STORE_ID!],
},
{
type: "function",
name: "get_invoice",
description: "Fetch one purchase invoice by its number.",
parameters: {
type: "object",
properties: { number: { type: "string" } },
required: ["number"],
additionalProperties: false,
},
strict: true,
},
],
};The module above is the whole replacement for an assistant object. It gives you what the prompt object promised, a reviewable, diffable and versioned definition, through the tool you already use for that: git. Put a version field in it and log that version with every response, so that when a customer reports a bad answer you can tell which instructions produced it. If you need per-tenant variants, compose them in code from a base definition rather than keeping copies.
OpenAI does not convert threads into conversations. The guide's example shows how history could have been copied before the sunset by paging through threads.messages.list, then says that call no longer works and that you should use your stored messages instead. If your application kept its own copy of each message, you can rebuild every active thread. If it only kept a thread id and relied on OpenAI to hold the content, that history is gone, and the honest fix is to start those users on a fresh conversation.
The conversion itself is small, with one detail the guide's example skips. Its sample passes every converted message to a single conversations.create call, but the Conversations API reference says you may add up to 20 items at a time, both on create and on items.create. A demo thread of five messages passes; a customer thread of sixty does not. Batch it:
# Rebuild one legacy thread as a Conversation from messages YOU stored.
# OpenAI's example reads threads.messages.list — that call stopped working
# at the sunset, so the source is your own table, oldest message first.
from openai import OpenAI
client = OpenAI()
BATCH = 20 # conversations.create and items.create take up to 20 items per call
def to_item(row):
# User text is input_text; assistant text must be output_text, or the
# rebuilt history reads as if the user said the assistant's lines.
part = "input_text" if row["role"] == "user" else "output_text"
return {"role": row["role"], "content": [{"type": part, "text": row["text"]}]}
def migrate_thread(legacy_thread_id, rows):
items = [to_item(r) for r in rows if r["text"]]
conversation = client.conversations.create(
items=items[:BATCH],
# Keep the old id, so support can trace a conversation to its thread.
metadata={"legacy_thread_id": legacy_thread_id},
)
# Wrong: create(items=items) for a 60-message thread. It works in a demo
# with five messages and fails on the first real customer history.
for start in range(BATCH, len(items), BATCH):
client.conversations.items.create(
conversation.id, items=items[start:start + BATCH]
)
return conversation.idTwo smaller rules come from the same example. User text becomes an input_text part and assistant text an output_text part, and images become input_image parts with their URL and detail. Map each legacy thread id to the new conversation id in your own database during the migration, so that the next request from an existing user lands on the rebuilt history rather than creating a new one.
Only migrate threads that someone will come back to. A thread untouched for six months is cheaper to archive in your own database than to rebuild, and the conversation you would create for it will sit forever, because conversations are not covered by the 30-day response retention.
A run was an asynchronous job: create it, poll its status, handle requires_action by submitting tool outputs, then poll again. A response is a request that returns output items. The tool loop that requires_action hid is now explicit: you read function_call items from the output, run your function, and send back function_call_output items that reference each call by call_id.
// Before (Assistants, openai v4 SDK): start a run, then poll it.
let run = await openai.beta.threads.runs.create(threadId, {
assistant_id: assistantId,
});
while (["queued", "in_progress"].includes(run.status)) {
await sleep(1000);
run = await openai.beta.threads.runs.retrieve(threadId, run.id);
}
// ...then requires_action -> submit_tool_outputs -> poll again.
// After (Responses + Conversations): one request per model turn.
import OpenAI from "openai";
import { INVOICE_HELPER } from "./agents/invoice-helper";
const client = new OpenAI();
const MAX_TOOL_TURNS = 5;
export async function ask(conversationId: string, text: string) {
const base = {
model: INVOICE_HELPER.model,
// Sent on every call: the conversation stores items, not your config.
instructions: INVOICE_HELPER.instructions,
tools: INVOICE_HELPER.tools,
conversation: conversationId,
};
let response = await client.responses.create({
...base,
input: [{ role: "user", content: text }],
});
// requires_action used to hide this loop. Now it is yours: run each
// function_call, answer it by call_id, and send the outputs back.
for (let turn = 0; turn < MAX_TOOL_TURNS; turn++) {
const calls = response.output.filter((item) => item.type === "function_call");
if (calls.length === 0) return response.output_text;
const outputs = await Promise.all(
calls.map(async (call) => ({
type: "function_call_output" as const,
call_id: call.call_id,
output: JSON.stringify(await runTool(call.name, JSON.parse(call.arguments))),
})),
);
response = await client.responses.create({ ...base, input: outputs });
}
// A cap the old run never needed you to write: fail loudly, don't spin.
throw new Error(`Tool loop did not settle in ${MAX_TOOL_TURNS} turns`);
}Two choices in that code are deliberate. Instructions and tools are sent on every call instead of being trusted to persist, because the conversation stores items, not your configuration, and sending them each time also means a config change takes effect on the next turn. And the loop has a hard cap. A run would eventually fail or expire on its own; your loop will happily call a misbehaving tool forever, so give it a ceiling and an error you will notice in your logs.
In Assistants, hosted tools were declared on the assistant and their data was attached through tool_resources, either on the assistant or on an individual thread. In Responses there is no assistant object and no thread to attach anything to, so each tool carries its resources in its own definition, on each request.
// Before: tools declared on the assistant, files bound via tool_resources
// (on the assistant or on the thread).
await openai.beta.assistants.create({
model: "gpt-4o",
tools: [{ type: "file_search" }, { type: "code_interpreter" }],
tool_resources: {
file_search: { vector_store_ids: ["vs_policies"] },
code_interpreter: { file_ids: ["file-ledger-q3"] },
},
});
// After: each tool brings its own resources, on every request.
tools: [
{
type: "file_search",
vector_store_ids: ["vs_policies"], // the same vector store, referenced by id
max_num_results: 8,
},
{
type: "code_interpreter",
// "auto" creates a container, or reuses one already in the context.
container: { type: "auto", file_ids: ["file-ledger-q3"] },
},
]
// Read citations from the output message, not from run steps:
// annotations of type file_citation (file_search) and
// container_file_citation (files code_interpreter wrote).For file_search the change is mostly mechanical, because the vector stores are the same objects and are referenced by id. The case that needs thought is a vector store that was attached to a single thread, typically files a user uploaded during one chat. Record that store id next to the conversation in your database and add it to the tool definition for requests on that conversation. For code_interpreter, an auto container is created for you or reused from earlier in the context, and files the model generates come back as container_file_citation annotations on the output message, which is where any download link in your UI should now read from.
The state model changes in a way that matters for anyone holding customer or ERP data. OpenAI's conversation-state guide sets out three rules worth writing into your migration notes:
So a conversation is durable until you delete it. That is what you want for a support chat that a user returns to, and it is a liability for an invoice assistant that sees supplier names, amounts and bank details. Decide the retention period before you migrate, store the conversation id against the user or tenant that owns it, and delete conversations when that relationship ends. Under Indonesia's personal data protection law, keeping chat history indefinitely without a stated purpose is not a position you want to defend.
Do not create one shared conversation per tenant to save effort. Every user in that tenant would then read the same history, including what other users asked. One conversation per end user per chat session is the safe default.
This is the order I would work in, so that each step is verifiable before the next one depends on it:
The Assistants API shutdown is less a rename than a transfer of ownership. Configuration moves into your repository, conversation history into your database first and a Conversation second, and the tool loop into your code. Move each of those to the place you control, not to the next managed object on the deprecations list, and this is the last migration this integration needs.