AI
OpenAI Realtime Voice Agent over SIP with gpt-realtime-2.1
October 202613 min read

OpenAI's Realtime guide names gpt-realtime-2.1 as the current model, with gpt-realtime-2.1-mini as the cheaper option. The legacy gpt-realtime, gpt-4o-realtime and their mini variants shut down on 20 January 2027, so new work should start on the 2.1 family. The cost guide recommends tuning on the full model first and testing mini afterwards.
Buy a number from a SIP trunking provider and point the trunk at sip:$PROJECT_ID@sip.api.openai.com;transport=tls. OpenAI then sends a realtime.call.incoming webhook to your project, and your server accepts the call by its call_id with the session configuration. With the Agents SDK you build that configuration with OpenAIRealtimeSIP.buildInitialConfig and attach a RealtimeSession using the OpenAIRealtimeSIP transport.
The GA interface drops the OpenAI-Beta: realtime=v1 header, requires session.type, moves output audio settings under session.audio.output and renames events to forms such as response.output_audio.delta. Browser credentials come from POST /v1/realtime/client_secrets and WebRTC sessions use /v1/realtime/calls. The beta interface was shut down on 12 May 2026, so beta code connects but its handlers no longer match.
No. A realtime handoff updates the live session with the new agent's instructions and tools, but the session keeps one model for its whole life and the voice cannot change after audio has been produced. When a step needs another model, delegate through a tool that calls a normal text agent on your server and returns the result to the voice session.
Audio is billed per token: caller audio counts 1 token per 100 ms and agent audio 1 token per 50 ms. On gpt-realtime-2.1 audio input is $32 and output $64 per million tokens, so a minute of agent speech is about $0.077 before context re-reads. Because the whole conversation is re-sent on each response, the cached input rate of $0.40 per million decides the real bill, so keep history stable and check response.done usage.

Key Takeaway
An OpenAI realtime voice agent in 2026 runs on gpt-realtime-2.1 through the GA Realtime API. The Agents SDK wraps it in RealtimeAgent and RealtimeSession, and OpenAIRealtimeSIP attaches that session to a phone call accepted from a realtime.call.incoming webhook. Legacy realtime models shut down on 20 January 2027.
The request was ordinary: a clinic wanted its phone answered after the front desk went home, in Bahasa Indonesia, able to find a free slot with a specific doctor and hold it. The first prototype I looked at was a year-old Realtime beta script. It connected, the caller said hello, and nothing came back. No error, no crash. The event names it listened for no longer exist, so every audio chunk went straight past the handler.
This post builds an OpenAI realtime voice agent the way the current documentation describes it: RealtimeAgent and RealtimeSession from the TypeScript Agents SDK on gpt-realtime-2.1, with tools, a handoff and guardrails, and a real phone number connected over SIP. It covers what the beta-to-GA move broke, which legacy models stop working on 20 January 2027, and how latency and cost behave for an Indonesian clinic-booking line. Every API name and price comes from the OpenAI docs listed at the end.
A realtime voice agent is one speech-to-speech model session that listens, decides and talks in one loop, with no separate speech-to-text and text-to-speech pipeline stitched together by you. OpenAI's Realtime guide names gpt-realtime-2.1 as the current model and offers three ways in: WebRTC from a browser, a WebSocket from a server, and SIP for phone calls. The Agents SDK puts the same agent abstraction over all three, so the choice of transport is mostly a choice of where audio and tools live.
| Transport | Where the session runs | Who handles audio | Fits |
|---|---|---|---|
| OpenAIRealtimeWebRTC | Browser, with an ephemeral client secret minted on your server | The SDK: microphone capture and playback are automatic | A booking widget on the clinic website |
| OpenAIRealtimeWebSocket | Your server | You: sendAudio in, the audio event out, interruption handling too | Custom media pipelines and telephony bridges you run yourself |
| OpenAIRealtimeSIP | Your server, attached to an existing call by its callId | The SIP call itself: media flows between the carrier and OpenAI | A real phone number through a SIP trunk |
Two limits from the SDK's build guide shape every design decision after this point. A session keeps one model for its whole life, and the voice can only change before the session has produced any audio. And the Realtime API currently caps a single session at 60 minutes, which no clinic call should reach, but a call-centre queue that parks callers on the agent can.
The silent prototype was not a bug in the script. OpenAI's deprecations page records that the Realtime beta interface, the one selected with the OpenAI-Beta: realtime=v1 header, was shut down on 12 May 2026. Code written against it does not fail loudly; it connects to an interface that answers in a different shape. The Realtime guide lists the changes that matter:
// Wrong: beta-era client. The beta interface was shut down on 2026-05-12,
// so this header and these event names belong to an API that no longer exists.
const ws = new WebSocket(url, {
headers: { Authorization: "Bearer " + key, "OpenAI-Beta": "realtime=v1" },
});
ws.on("message", (raw) => {
const event = JSON.parse(raw.toString());
if (event.type === "response.audio.delta") play(event.delta); // never fires on GA
});
// Right: GA interface. No beta header, session.type is required,
// output audio config lives under session.audio.output.
ws.send(JSON.stringify({
type: "session.update",
session: {
type: "realtime",
model: "gpt-realtime-2.1",
audio: { output: { voice: "marin" } },
},
}));
ws.on("message", (raw) => {
const event = JSON.parse(raw.toString());
switch (event.type) {
case "response.output_audio.delta": play(event.delta); break;
case "response.output_audio_transcript.delta": log(event.delta); break;
case "response.output_text.delta": log(event.delta); break;
}
});The second deadline is closer than it looks. On 20 July 2026 OpenAI announced that the legacy realtime models shut down on 20 January 2027. If a config file anywhere still says gpt-realtime or gpt-4o-realtime, it has under four months left, and the replacement names are not the obvious ones.
| Legacy model | Shutdown | Replacement |
|---|---|---|
| gpt-realtime, gpt-4o-realtime | 20 January 2027 | gpt-realtime-2.1 |
| gpt-realtime-mini, gpt-4o-mini-realtime | 20 January 2027 | gpt-realtime-2.1-mini |
| gpt-4o-realtime-preview and its dated snapshots | Already shut down on 7 May 2026 | gpt-realtime-1.5 at the time, now gpt-realtime-2.1 |
RealtimeAgent takes the same shape as a text agent: a name, instructions, tools defined with the tool() helper and a zod schema, and a list of handoffs. RealtimeSession then binds that agent to a model and a transport. The receptionist below has two tools, one that reads open slots and one that holds a slot, plus a billing specialist it can hand the call to.
import { RealtimeAgent, RealtimeSession, tool } from "@openai/agents/realtime";
import { z } from "zod";
import { clinicApi } from "./clinic-api"; // your backend, never the model's
const findSlots = tool({
name: "find_slots",
description: "List open appointment slots for a doctor on a date (Asia/Jakarta).",
parameters: z.object({
doctorId: z.string(),
date: z.string().describe("YYYY-MM-DD"),
}),
timeoutMs: 4000, // a caller hears dead air while a tool runs; fail fast
async execute({ doctorId, date }) {
return clinicApi.openSlots(doctorId, date);
},
});
const bookSlot = tool({
name: "book_slot",
description: "Hold a slot for a verified patient. Creates a booking, so ask first.",
parameters: z.object({ slotId: z.string(), patientPhone: z.string() }),
needsApproval: true, // on a phone there is no button: the server decides
async execute({ slotId, patientPhone }) {
return clinicApi.hold(slotId, patientPhone);
},
});
const billingAgent = new RealtimeAgent({
name: "Billing",
handoffDescription: "Questions about invoices, BPJS coverage and payment",
instructions: "Answer billing questions in the caller's language. Never quote a price you did not get from a tool.",
});
export const receptionist = new RealtimeAgent({
name: "Receptionist",
instructions:
"You answer the phone for a clinic in Jakarta. Speak Bahasa Indonesia unless the caller uses English. " +
"Before calling a tool, say a short filler such as 'sebentar, saya cek jadwalnya'.",
tools: [findSlots, bookSlot],
handoffs: [billingAgent],
});
export const sessionOptions = {
model: "gpt-realtime-2.1",
config: {
outputModalities: ["audio"],
reasoning: { effort: "low" }, // higher effort costs latency on every turn
audio: {
input: {
transcription: {
model: "gpt-live-transcribe",
languages: ["id", "en"],
keywords: ["BPJS", "dr. Sari", "poli anak"],
},
turnDetection: { type: "semantic_vad", eagerness: "medium", interruptResponse: true },
},
output: { voice: "marin" },
},
},
} as const;Three choices in that config are deliberate. The transcription block lists both id and en in languages and seeds keywords with the doctor names and BPJS, because callers mix languages mid-sentence and a misheard doctor's name produces a confident booking with the wrong person. Reasoning effort is set to low, since the build guide warns that higher effort adds latency and tokens on every turn. And each tool has a timeoutMs, because the build guide is explicit that while a tool executes the agent cannot process new requests from the caller, so a slow clinic database is heard as silence.
Packages: @openai/agents for RealtimeAgent, RealtimeSession, tool and the transports under @openai/agents/realtime, zod for the tool schemas, and openai for webhook verification and call control.
A realtime handoff is not the handoff you know from text agents. The SDK updates the live session in place with the new agent's instructions and tools, so the billing agent inherits the full conversation and input filters are not applied. The model cannot change, and because the receptionist has already spoken, neither can the voice. If a step genuinely needs a different model, for example a reasoning model checking a refund, the build guide's answer is delegation through a tool: the tool sends the request and the history snapshot to a normal text Agent on your server and speaks the result.
Guardrails run on what the agent says, not on what the caller says. In an audio session the SDK checks the output transcript as it streams, every 100 characters by default and once more on the final transcript. Speaking takes longer than generating the transcript, so a tripped guardrail usually cuts the response before the caller hears the offending phrase. For a clinic, the phrase that must never be spoken is a diagnosis.
import { RealtimeSession, type RealtimeOutputGuardrail } from "@openai/agents/realtime";
const noDiagnosis: RealtimeOutputGuardrail = {
name: "No medical diagnosis",
async execute({ agentOutput }) {
const diagnosing = /\b(diagnosis|diagnosa|you have|anda menderita)\b/i.test(agentOutput);
return { tripwireTriggered: diagnosing, outputInfo: { diagnosing } };
},
};
// callId is ours: it came from the webhook, not from the model. The From header
// is caller-supplied SIP metadata, so it is a lookup hint, never proof of identity.
export function createGuardedSession(callId: string) {
const session = new RealtimeSession(receptionist, {
...sessionOptions,
outputGuardrails: [noDiagnosis],
outputGuardrailSettings: { debounceTextLength: 100 }, // the default; -1 = only at the end
toolExecution: { preApprovalInputGuardrails: true },
});
session.on("guardrail_tripped", () => {
metrics.increment("voice.guardrail_tripped"); // the response is already cut off
});
// No UI on a phone line: approval is a policy check against state your server
// recorded (an OTP or date-of-birth match for this callId), never the model's arguments.
session.on("tool_approval_requested", async (_ctx, _agent, request) => {
const verified = await clinicApi.isCallVerified(callId);
if (verified) await session.approve(request.approvalItem);
else await session.reject(request.approvalItem, {
message: "Caller not verified. Offer to send a WhatsApp confirmation link instead.",
});
});
return session;
}Tool approval needs rethinking on a phone. In a browser, needsApproval pauses the call and shows a button. On a SIP call there is nobody to click it, so the tool_approval_requested handler becomes a server-side policy check: approve book_slot only if your backend has recorded a verification for this call. Setting preApprovalInputGuardrails runs the tool's input guardrails before the approval event as well as after, so an obviously bad argument is rejected without reaching the policy check at all.
Function tools run wherever the RealtimeSession runs. In the browser WebRTC setup that means the browser, where anyone can open the developer tools and call clinicApi.hold directly. Keep privileged tools behind your own authenticated backend, or run the session server-side with the SIP or WebSocket transport. Authorise against trusted session context, never against arguments the model supplied.
The SIP path keeps call audio between your carrier and OpenAI; your server only handles the webhook, the configuration and the tools. The setup, in the order the SIP guide gives it:
import express from "express";
import OpenAI from "openai";
import { OpenAIRealtimeSIP, RealtimeSession } from "@openai/agents/realtime";
import { receptionist, sessionOptions } from "./receptionist";
const openai = new OpenAI({
apiKey: process.env.OPENAI_API_KEY,
webhookSecret: process.env.OPENAI_WEBHOOK_SECRET,
});
const seen = new Set<string>(); // use Redis in production; webhooks are retried
const app = express();
app.post("/webhooks/openai", express.text({ type: "*/*" }), async (req, res) => {
// Signature check needs the raw body, so no express.json() on this route.
const event = await openai.webhooks.unwrap(req.body, req.headers);
if (event.type !== "realtime.call.incoming") return res.sendStatus(200);
const webhookId = String(req.headers["webhook-id"]);
if (seen.has(webhookId)) return res.sendStatus(200);
seen.add(webhookId);
const callId = event.data.call_id;
const config = await OpenAIRealtimeSIP.buildInitialConfig(receptionist, sessionOptions);
await openai.realtime.calls.accept(callId, config);
res.sendStatus(200);
const session = new RealtimeSession(receptionist, {
transport: new OpenAIRealtimeSIP(),
...sessionOptions,
});
await session.connect({ apiKey: process.env.OPENAI_API_KEY!, callId });
});
// Hand the caller to a human front desk: a SIP REFER, relayed to your trunk.
export async function transferToFrontDesk(callId: string) {
await fetch("https://api.openai.com/v1/realtime/calls/" + callId + "/refer", {
method: "POST",
headers: {
Authorization: "Bearer " + process.env.OPENAI_API_KEY,
"Content-Type": "application/json",
},
body: JSON.stringify({ target_uri: "tel:+622150000000" }),
});
}Two traps sit in that flow. buildInitialConfig throws locally if the turn-detection config contains threshold, prefixPaddingMs or silenceDurationMs, because the calls API rejects those fields for SIP sessions; semantic_vad with interruptResponse is the safe choice. And webhooks are retried, so deduplicate on the webhook-id header before accepting, or one ring becomes two attach attempts racing each other.
The network side is easy to forget until a call rings and has no audio. Signalling needs outbound TLS on port 5061 to whatever sip.api.openai.com resolves to, and media is SRTP over UDP to four published CIDR ranges listed in the SIP guide. When the agent should give up, a POST to the refer endpoint with a tel: or sip: target sends a SIP REFER to your trunk and the clinic's front desk takes the call; the hangup endpoint ends it cleanly.
I cannot give you a millisecond figure that holds for your carrier, and nobody honest can. What the documentation does establish is where the delay is added, and each item is something you control:
The cheapest fix is in the prompt: tell the agent to say a short filler before it calls a tool, as the receptionist's instructions do. The SDK also offers backgroundResult for a tool whose output should be delivered without immediately triggering another spoken response, which suits a confirmation that the caller does not need read back.
Record test calls through the real trunk from a real Indonesian mobile number before you tune anything. Calling the agent from a browser in the same city as your server measures the wrong path, and every decision you tune against it will be off by the international hop.
Realtime billing is per token, and audio has fixed token rates: the cost guide counts caller audio at 1 token per 100 ms and assistant audio at 1 token per 50 ms, so a minute of the agent talking is 1,200 tokens and a minute of the caller talking is 600. The model pages give the prices per million audio tokens:
| Item | gpt-realtime-2.1 | gpt-realtime-2.1-mini |
|---|---|---|
| Audio input, per 1M tokens | $32 | $10 |
| Cached audio input, per 1M tokens | $0.40 | $0.30 |
| Audio output, per 1M tokens | $64 | $20 |
| One minute of agent speech | about $0.077 | about $0.024 |
| One minute of caller speech, first time it is sent | about $0.019 | about $0.006 |
Those per-minute rows are the floor, not the bill, because the whole conversation is sent back to the model as input on every response. For a four-minute booking call with about two minutes of caller speech, ninety seconds of agent speech and twelve turns, my arithmetic on gpt-realtime-2.1 gives roughly $0.15 of fresh audio plus about 18,000 re-read audio tokens. If those re-reads hit the cache they add under one cent; if the cache misses every turn they add more than half a dollar. The two published input rates are 80 times apart, and that gap decides the economics. Cap the context so long calls cannot grow without bound:
{
"type": "session.update",
"session": {
"truncation": {
"type": "retention_ratio",
"retention_ratio": 0.8,
"token_limits": { "post_instructions": 8000 }
}
}
}Two things bust that cache. The cost guide says that editing or deleting conversation items invalidates it from the point of the change, and that instructions and tools sit at the start of the conversation, so changing them mid-session costs cache hits for every later turn. A handoff does exactly that: the SDK swaps the active agent's instructions and tools in place. I keep handoffs to one per call and put stable text first. The cost guide's own advice is to start on the full model, get the prompt right, then test the mini model, which is weaker at instruction following and function calling. My estimates exclude text tokens for instructions and tools, input transcription, and the trunk provider's per-minute charges, so treat them as a lower bound to check against the usage block in response.done.
The rule I took from that silent prototype: a voice agent fails quietly, so check it against the current names, not against whether it connects. Pin gpt-realtime-2.1 or gpt-realtime-2.1-mini before 20 January 2027, accept SIP calls with buildInitialConfig, authorise tools on the server, say something before every slow tool, and read response.done usage on real calls before you quote anyone a price per call.
Sources and further reading