AI
OpenAI ChatKit Tutorial: Self-Hosted Agent Chat in React
October 202612 min read

ChatKit is OpenAI's embeddable chat interface for agent experiences. It ships widgets, tool-invocation display, file attachments and chain-of-thought visualisations, so you configure a chat UI instead of building one. In React you render it with the ChatKit component and the useChatKit hook from @openai/chatkit-react.
Install the Python server with pip install openai-chatkit, subclass ChatKitServer and override respond to run your Agents SDK agent through stream_agent_response. Expose one POST endpoint that passes the request body to server.process, and point the useChatKit api.url option at it. You also implement a Store for threads and items.
Yes. OpenAI states that Agent Builder is scheduled to shut down on 30 November 2026 but ChatKit remains available. The hosted path that backed ChatKit with an Agent Builder workflow is for existing users during the transition, and new work should use the advanced integration with your own server-side agent.
The UI is loaded from OpenAI's CDN and the iframe verifies your domain key against api.openai.com on load. The ChatKit production guide also says the iframe sends its own telemetry to OpenAI-controlled endpoints, which it states contains no PII or message content. Your threads, files and agent run on your own server.
Attachments are disabled by default, so set composer.attachments.enabled to true and choose an uploadStrategy of two_phase or direct under the api option. On the server, pass an AttachmentStore to ChatKitServer; with two_phase its create_attachment returns an upload URL the browser uploads to. Converting a stored file into model input is done in a ThreadItemConverter.

Key Takeaway
OpenAI ChatKit is an embeddable React chat UI that, when self-hosted, talks to your own Python ChatKitServer endpoint, which runs an Agents SDK agent and streams replies, widgets and progress events. You still own authentication, the thread Store and attachment storage, and the UI itself loads from OpenAI's CDN behind a registered domain key.
The chat widget on this site is 826 lines of React in Chatbot.tsx and another 416 in its API route. A good share of that is plumbing nobody sees: reading the response body with getReader and a TextDecoder, cancelling the request with an AbortController when the panel closes, and keeping the last 20 messages in localStorage because there is no server-side thread store. So when OpenAI positioned ChatKit as the way to build agent chat without reinventing the chat UI, the question I cared about was which of those lines it would actually delete, and which new ones it would add.
This OpenAI ChatKit tutorial builds the self-hosted version end to end: @openai/chatkit-react on the front end, the open-source ChatKit Python server on the back end, and an Agents SDK agent behind it, using an ERP order-status desk as the example. It is the path OpenAI now recommends, because Agent Builder, which backed the hosted ChatKit option, is scheduled to shut down on 30 November 2026. Every API name below comes from OpenAI's ChatKit guides, the chatkit-python repository and its production guide, or the official advanced samples.
ChatKit has two integration paths. The hosted one creates a session against an Agent Builder workflow and hands the browser a client secret; OpenAI's guide now limits it to teams who already have such a workflow, during the transition window. The other, which the docs call the advanced or custom server integration, is what this post builds. Self-hosted is slightly misleading as a name, because it describes the back end only. There are three pieces, and you run two of them:
The upside of that split is that your agent, your data and your model calls stay on your infrastructure, and the guide lists custom authentication, data residency and on-premises deployment as the reasons to choose it. The cost is that everything a hosted product would have done for you, namely login, tenancy, retention and file storage, is now code you write. The rest of this post is that code, in the order I would write it.
Install the server with pip install openai-chatkit, which pulls in the openai-agents package as a dependency. ChatKitServer is generic over a context type of your choosing, and the one method you must override is respond, an async generator that receives the thread metadata and the new user message and yields thread stream events. The chatkit.agents module provides the bridge: simple_to_agent_input converts stored thread items into Agents SDK input, and stream_agent_response turns a streamed run into ChatKit events, including tool progress and any widgets your tools emit.
# server.py
import os
from typing import AsyncIterator
from agents import Agent, RunContextWrapper, Runner, function_tool
from chatkit.agents import AgentContext, simple_to_agent_input, stream_agent_response
from chatkit.server import ChatKitServer
from chatkit.types import ThreadMetadata, ThreadStreamEvent, UserMessageItem
from .context import RequestContext # your own dataclass: user_id, tenant_id, locale
from .erp import fetch_order # your own read-only service call
@function_tool(description_override="Look up one sales order by its number.")
async def get_order_status(ctx: RunContextWrapper[AgentContext], order_no: str) -> dict:
# request_context is whatever the endpoint passed to server.process(), so the
# authenticated user travels with the run and the tool can scope its query.
user: RequestContext = ctx.context.request_context
await ctx.context.stream_progress(icon="document", text=f"Looking up {order_no}")
return await fetch_order(order_no, tenant_id=user.tenant_id)
order_desk = Agent[AgentContext](
name="Order desk",
instructions="Answer questions about the user's own sales orders. "
"Always call the tool; never guess a status.",
model=os.environ["CHATKIT_MODEL"],
tools=[get_order_status],
)
class OrderDeskServer(ChatKitServer[RequestContext]):
async def respond(
self,
thread: ThreadMetadata,
input_user_message: UserMessageItem | None,
context: RequestContext,
) -> AsyncIterator[ThreadStreamEvent]:
# Wrong: order="asc", limit=20 returns the FIRST 20 items. Once a thread
# passes 20 items, the message the user just sent is not in the page.
# Right: take the newest 20, then flip them back into chronological order.
page = await self.store.load_thread_items(
thread.id, after=None, limit=20, order="desc", context=context
)
input_items = await simple_to_agent_input(list(reversed(page.data)))
agent_context = AgentContext(thread=thread, store=self.store, request_context=context)
result = Runner.run_streamed(order_desk, input_items, context=agent_context)
async for event in stream_agent_response(agent_context, result):
yield eventTwo details in that block matter more than they look. The model name comes from configuration because nothing in ChatKit ties you to one; the guide says the server can connect to any agentic service, and respond is just a generator, so it could call another provider entirely. And the tool reads the authenticated user from ctx.context.request_context, which is the same object your HTTP layer passed in. That is how a tool scopes an ERP query to the caller's tenant without the model ever being told the tenant id.
OpenAI's two official snippets disagree on how to load history. The custom-integration guide loads the newest 20 items in descending order and reverses them; the chatkit-python quickstart loads 20 items in ascending order with no cursor. The second returns the oldest 20 items, so once a thread grows past 20 the user's new message silently drops out of the model's input. Use descending order and reverse, as above.
ChatKit routes every request type, such as creating a thread, posting a message, listing threads or running a widget action, through a single endpoint. ChatKitServer.process takes the raw body and a context object, and that context is the only channel through which identity reaches your store and tools. The production guide is explicit that the Python SDK expects your app to handle authentication, so the pattern is: authenticate with whatever your framework already uses, build a typed context with the user id, tenant and locale, and pass it in.
# main.py — one endpoint; ChatKitServer routes every request type internally
from fastapi import Depends, FastAPI, HTTPException, Request, Response
from fastapi.responses import StreamingResponse
from chatkit.server import StreamingResult
from .auth import verify_jwt # your existing session or JWT check
from .context import RequestContext
from .server import OrderDeskServer
from .store import PostgresStore
app = FastAPI()
server = OrderDeskServer(store=PostgresStore())
async def current_user(request: Request) -> RequestContext:
claims = verify_jwt(request.headers.get("authorization", ""))
if claims is None:
# ChatKitServer does no authentication of its own. Skip this and the
# endpoint is an open, metered proxy to your model account.
raise HTTPException(status_code=401, detail="Unauthorized")
return RequestContext(
user_id=claims["sub"],
tenant_id=claims["tenant"],
# The ChatKit client sends exactly one locale in Accept-Language.
locale=request.headers.get("accept-language", "en"),
)
@app.post("/api/chatkit")
async def chatkit_endpoint(request: Request, ctx: RequestContext = Depends(current_user)):
result = await server.process(await request.body(), ctx)
if isinstance(result, StreamingResult):
return StreamingResponse(result, media_type="text/event-stream")
return Response(content=result.json, media_type="application/json")Locale comes for free: the ChatKit client sends a single locale in the Accept-Language header on every request, defaulting to the browser's language unless you override the locale option. For a bilingual audience like mine, that header is what decides whether tool output and error messages come back in English or Indonesian, and the production guide shows the gettext pattern for it. The other thing the guide asks for is that system instructions stay static and never include user-supplied values such as a thread title, because those are a prompt injection path.
chatkit.store.Store is an abstract class with twelve methods: load, save, list and delete for threads; add, save, load, list and delete for thread items; and save, load and delete for attachments. Every one of them receives the context object, which is the point. The quickstart's in-memory store is fine for a demo and wrong for anything shared, because its load_threads returns every thread in the dictionary regardless of who asked. The guide's advice for production is a durable database with the models stored as JSON blobs, so that a library upgrade which adds fields does not need a schema migration.
# store.py — two of the Store methods, showing where the tenant check lives
from chatkit.store import NotFoundError, Store
from chatkit.types import Page, ThreadMetadata
from .context import RequestContext
from .db import pool # an asyncpg pool
class PostgresStore(Store[RequestContext]):
async def load_thread(self, thread_id: str, context: RequestContext) -> ThreadMetadata:
row = await pool.fetchrow(
"SELECT data FROM chatkit_threads WHERE id = $1 AND user_id = $2",
thread_id, context.user_id,
)
if row is None:
# Same error for "missing" and "someone else's": never confirm that
# a thread id exists to a user who does not own it.
raise NotFoundError(f"Thread {thread_id} not found")
# Stored as a JSON blob, so a library upgrade that adds fields
# needs no migration.
return ThreadMetadata.model_validate_json(row["data"])
async def load_threads(self, limit, after, order, context) -> Page[ThreadMetadata]:
# Wrong (the quickstart's in-memory store): list(self.threads.values())
# returns every user's history to whoever opens the thread list.
rows = await pool.fetch(
f"SELECT id, data FROM chatkit_threads WHERE user_id = $1 "
f"ORDER BY created_at {'DESC' if order == 'desc' else 'ASC'} LIMIT $2",
context.user_id, limit + 1,
)
... # apply the "after" cursor, build Page(data=..., has_more=..., after=...)I would treat the Store as the real security boundary of the whole feature. The endpoint decides who you are; the Store decides what you can see. If load_thread does not check ownership, any user who learns or guesses a thread id can read another user's conversation, including whatever an ERP tool returned into it. Retention belongs here too: the production guide suggests implementing it in the Store, for example deleting threads older than a set number of days, because threads hold user text, attachments and tool output.
Install the bindings with npm install @openai/chatkit-react and load chatkit.js from OpenAI's CDN. For a custom server, the api option takes a url instead of the hosted getClientSecret callback, plus a domainKey, an optional fetch override and the upload strategy. The fetch override is the clean way to attach your own bearer token, because ChatKit calls it for every request it makes. Theme, start-screen prompts and composer settings all live in the same options object.
"use client";
// components/OrderDeskChat.tsx (Next.js App Router)
import Script from "next/script";
import { ChatKit, useChatKit } from "@openai/chatkit-react";
export function OrderDeskChat({ getToken }: { getToken: () => Promise<string> }) {
const { control } = useChatKit({
api: {
url: "/api/chatkit",
domainKey: process.env.NEXT_PUBLIC_CHATKIT_DOMAIN_KEY ?? "domain_pk_localhost_dev",
// Called for every ChatKit request, so the bearer token is always current.
fetch: async (input, init) => {
const headers = new Headers(init?.headers);
headers.set("Authorization", `Bearer ${await getToken()}`);
return fetch(input, { ...init, headers });
},
// Lives under api, not composer, in the published types.
uploadStrategy: { type: "two_phase" },
},
theme: {
colorScheme: "dark",
color: { accent: { primary: "#22d3ee", level: 1 } },
radius: "round",
density: "compact",
},
startScreen: {
greeting: "Ask about any of your sales orders",
prompts: [
{ label: "Late orders", prompt: "Which of my orders are past their delivery date?", icon: "search" },
],
},
composer: {
placeholder: "Order number or question",
attachments: {
enabled: true, // off by default
maxSize: 5 * 1024 * 1024, // default is 100 MB, far too generous
maxCount: 3, // default is 10
accept: { "application/pdf": [".pdf"], "image/*": [".png", ".jpg"] },
},
},
onError: ({ error }) => console.error("ChatKit error", error),
});
return (
<>
<Script src="https://cdn.platform.openai.com/deployments/chatkit/chatkit.js" strategy="afterInteractive" />
<ChatKit control={control} className="h-[600px] w-full" />
</>
);
}The theme options cover colour scheme, accent colour and level, a grayscale tint, corner radius from pill to sharp, density from compact to spacious, and typography. That was enough for me to match this site's dark palette with one accent colour. Attachment limits are worth setting explicitly: the published types default to 100 MB per file and 10 files per message, which is generous for a chat that forwards files to a model.
Trust the TypeScript types over the prose examples. The theming guide shows attachments configured with an uploadStrategy inside composer and no enabled flag, but the @openai/chatkit type definitions put uploadStrategy under api and make composer.attachments.enabled the switch, defaulting to false. In a Next.js App Router project, the component also needs the use client directive, since useChatKit is a hook.
Widgets are where ChatKit earns its place over a hand-rolled UI. You design a card, list or form visually at widgets.chatkit.studio, preview it with sample data, export a .widget file and commit it beside the server code. On the server, WidgetTemplate.from_file loads it and build fills its placeholders; a tool can then call ctx.context.stream_widget, and stream_agent_response forwards the widget to the client. Buttons and form controls carry an action with a type and a payload, which arrives at the action method on your ChatKitServer without the user typing a message.
# widgets.py — a card designed in widgets.chatkit.studio, exported as a .widget file
from datetime import datetime
from agents import RunContextWrapper, function_tool
from chatkit.agents import AgentContext
from chatkit.types import AssistantMessageContent, AssistantMessageItem, ThreadItemDoneEvent
from chatkit.widgets import WidgetTemplate
from .erp import create_cancel_request, fetch_order
order_card = WidgetTemplate.from_file("widgets/order_card.widget")
@function_tool(description_override="Show an order as a card with its lines and status.")
async def show_order_card(ctx: RunContextWrapper[AgentContext], order_no: str) -> str:
order = await fetch_order(order_no, tenant_id=ctx.context.request_context.tenant_id)
# The template's {{ }} placeholders are filled here, server-side, then streamed.
await ctx.context.stream_widget(order_card.build(order))
return f"Displayed order {order_no}." # what the model sees, not the user
# On OrderDeskServer: the card's "Request cancellation" button carries
# onClickAction = { type: "order.request_cancel", payload: { order_no } }
async def action(self, thread, action, sender, context):
if action.type != "order.request_cancel":
return
# Draft, never post: the button files a request that a human approves.
await create_cancel_request(action.payload["order_no"], requested_by=context.user_id)
yield ThreadItemDoneEvent(
item=AssistantMessageItem(
thread_id=thread.id,
id=self.store.generate_item_id("message", thread, context),
created_at=datetime.now(),
content=[AssistantMessageContent(text="Cancellation request filed for approval.")],
)
)In an ERP setting I keep one rule for actions: a button may create a draft or a request, never a posted document. The cancellation button above files a request that a person approves, which keeps the chat surface out of the approval chain. Actions can also be routed to the browser by setting the handler to client, which suits pure UI work such as opening a record in another panel. If a widget text field should stream as it is generated, give that Text or Markdown node an id; the widgets guide notes that only those nodes stream their text, and other changes re-render.
Attachments need an AttachmentStore passed to the ChatKitServer constructor. With the two_phase strategy, the client asks your server to create an attachment, your create_attachment returns an upload descriptor with a URL, and the browser uploads to it, which can be a signed URL on your object storage. The direct strategy posts the file to an uploadUrl you supply instead. The hosted strategy is for the Agent Builder path only. Getting the bytes to the model is a separate step: the official customer-support sample subclasses ThreadItemConverter and overrides attachment_to_message_content to turn a stored image into an input_image part.
Domain keys are the last step and the one that breaks deploys. You register each production hostname in the domain allowlist under your OpenAI organisation's security settings and copy the generated key into the domainKey option. On load, the ChatKit iframe calls api.openai.com to verify it, and if the key is missing or invalid ChatKit refuses to render. The official samples run locally on a placeholder key, so a staging hostname that was never registered is the first place this fails.
Self-hosting the back end does not remove OpenAI from the browser. The UI is fetched from OpenAI's CDN, the iframe verifies the domain key against api.openai.com, and the production guide states that the iframe sends its own telemetry to OpenAI-controlled endpoints, Datadog and chatgpt.com, which it says contains no PII or message content. If your compliance review requires that a page make no third-party requests, ChatKit does not fit, however the server is hosted.
Comparing against the chatbot I already run makes the trade concrete. It streams from a Groq-backed route and keeps everything in the browser; ChatKit would replace the client half and move history to the server, at the price of a Python service and a CDN dependency.
| Concern | ChatKit, self-hosted | Hand-rolled, like this site's chatbot |
|---|---|---|
| Streaming | StreamingResult over text/event-stream; the client parses events, including tool progress | A getReader and TextDecoder loop, plus abort handling, that you write and debug |
| Thread history | Server-side Store with a built-in thread list; you implement twelve methods | Whatever you build; mine keeps the last 20 messages in localStorage |
| Attachments | AttachmentStore with two_phase or direct uploads and previews | Upload route, storage, previews and model conversion all built by hand |
| Rich replies | Widgets designed in ChatKit Studio, with server or client actions | Your own components and your own message protocol to carry them |
| Where the UI runs | Iframe from OpenAI's CDN, domain key checked on load, OpenAI telemetry | Your bundle, your CSP, no third-party request |
| Server stack | Official server SDK is Python; respond can call any model or agent | Any language and any provider; this site uses a Next.js route and Groq |
My rule of thumb from that table: ChatKit pays off once you need two of server-side threads, attachments and widgets, because each is real, easy-to-get-wrong work when done by hand. For a single-purpose assistant that streams text, such as the one on this site, a hand-rolled UI on the Next.js stack you already run stays simpler, and keeps the page free of third-party requests. The deciding question is less about the chat UI than about whether a Python service is acceptable in your deployment.
ChatKit moves the hard part of agent chat from the browser to the server rather than removing it. The UI, streaming and widget rendering are done for you; authentication, the tenant-scoped Store, attachment storage, retention and domain keys are not. Build those five deliberately, load history newest-first, and keep every action a draft, and the self-hosted path gives you an agent chat whose data stays yours.
Sources and further reading