AI
OpenAI Agents SDK TypeScript Tutorial: Build a Tool Agent
October 202612 min read

Yes. OpenAI publishes it as @openai/agents on npm, developed in the openai/openai-agents-js repository alongside the Python SDK. Its README lists Node.js 22 or later, Deno, Bun and Cloudflare Workers with nodejs_compat as supported runtimes.
The package declares zod ^4.0.0 as a peer dependency, and the quickstart says the SDK uses Zod v4 for tool schemas and structured outputs. A project still on zod 3 should upgrade before writing tools. Passing a zod schema as a tool's parameters also switches on strict argument validation.
The run throws MaxTurnsExceededError. The default limit is 10 turns, where each turn is one model call, and you can change it per run with the maxTurns option. If you would rather return a fallback answer, set errorHandlers.maxTurns to return a finalOutput instead.
Yes. Call run() with stream: true, then return stream.toTextStream() piped through a TextEncoderStream as the Response body. Pass the request's signal so an aborted request stops the run, and await or catch stream.completed, which settles only after the run and any persistence work finish.
Use promptfoo's openai:agents provider with your agent exported from a TypeScript file and tracing enabled. Trajectory assertions such as trajectory:tool-used, trajectory:tool-args-match and trajectory:tool-sequence check which tools ran and with what arguments. Note that the provider docs pin @openai/agents to ^0.11.8, and that mock mode does not support explicit handoff() objects or hosted tools.

Key Takeaway
The OpenAI Agents SDK for TypeScript, @openai/agents 0.18.0, builds a tool-using agent from three pieces: tool() with a zod v4 schema, an Agent with tools and handoffs, and run() with stream, context and maxTurns options. Authorise inside execute, pin the model explicitly, and test the tool-call trajectory with promptfoo's openai:agents provider.
Most OpenAI Agents SDK examples are written in Python. For a team whose ERP back end is NestJS and whose portal is Next.js, a Python sidecar for an order-status assistant would mean a second deployment, a second copy of the sales-order types, and a network hop between the agent and the database it exists to read.
This tutorial builds that agent with the OpenAI Agents SDK for TypeScript, published on npm as @openai/agents and at version 0.18.0 at the time of writing. It covers a zod-typed function tool, a hosted file search tool, a handoff to a refunds specialist, streaming through a Next.js route handler, a maxTurns cap with a fallback, and an eval suite in promptfoo. Every option name comes from the SDK's own guides, linked at the end, and where two official pages disagree I say so.
The SDK is one install: @openai/agents plus zod. Its README lists Node.js 22 or later, Deno, Bun and Cloudflare Workers with nodejs_compat enabled as supported runtimes, and the package declares zod ^4.0.0 as a peer dependency, so a codebase still on zod 3 has to settle that upgrade before writing its first tool. The API key is read lazily from OPENAI_API_KEY the first time the SDK needs a client.
npm install @openai/agents zod
# 0.18.0 on npm at the time of writing. It declares zod ^4.0.0 as a
# peer dependency, so a project still on zod 3 upgrades first.
# .env.local
OPENAI_API_KEY=sk-...
# Pin the model. The SDK default is documented as subject to change,
# and two official pages already disagree about what it is.
AGENT_MODEL=gpt-5.6-lunaThe setting I now treat as mandatory is the model. The SDK's models guide says an agent without a model uses the default, currently gpt-5.6-luna with reasoning effort none and low verbosity, and that OPENAI_DEFAULT_MODEL overrides it process-wide. promptfoo's provider page for the same SDK names gpt-5.4-mini as the default for v0.10 and later and warns that it can change over time. Two official sources disagreeing about a default is the clearest argument I know for never relying on one.
A function tool is tool() with a name, a description, a parameters schema and an execute function. With a zod schema the SDK enables strict mode, so arguments that fail validation go back to the model as an error instead of reaching your code. The second argument to execute is the RunContext, which carries whatever object you passed as context to run(). Here that is the tenant and user from the authenticated session: values the model never chooses and cannot overwrite.
// lib/agent/tools.ts
import { tool, type RunContext } from '@openai/agents';
import { z } from 'zod';
import { db } from '@/lib/db';
// App state the model never chooses: who is asking, for which company.
// It travels in run(..., { context }) and stays in this process.
export interface ErpContext {
tenantId: string;
userId: string;
}
export const getSalesOrder = tool({
name: 'get_sales_order',
description:
'Look up one sales order by its number, for example SO-2026-00412. ' +
'Returns status, promised delivery date and invoice state.',
// A zod schema switches on strict mode: bad arguments go back to the
// model as an error and never reach execute().
parameters: z.object({
orderNumber: z.string().startsWith('SO-').describe('Sales order number'),
}),
// A slow ERP query becomes a "timed out" tool result the model can
// explain, instead of a chat request that hangs.
timeoutMs: 5_000,
async execute({ orderNumber }, runContext?: RunContext<ErpContext>) {
const ctx = runContext?.context;
if (!ctx) throw new Error('get_sales_order called without ERP context');
// Authorise HERE, against the order number the model actually chose.
// isEnabled runs before any arguments exist, so it cannot do this.
const order = await db.salesOrder.findFirst({
where: { number: orderNumber, tenantId: ctx.tenantId },
select: { number: true, status: true, promisedDate: true, invoiceStatus: true },
});
// Return, do not throw, for "not found": the run continues and the
// agent can ask the user to check the number. A throw is rethrown.
return order ?? { error: 'No order ' + orderNumber + ' for this company' };
},
});Two details earn their place on any tool that touches a database. timeoutMs bounds each call, and in the default error_as_result mode a timeout returns a message to the model saying the tool timed out, so a slow ERP query becomes something the agent can explain rather than a request that hangs. Returning a plain object for the not-found case, instead of throwing, keeps the run going: the SDK serialises non-string results for the model, while an exception from execute is rethrown by default because the default errorFunction is disabled.
isEnabled decides whether the model can see a tool on this turn, but it runs before the model has produced any arguments, and the tools guide is explicit that it does not replace authorisation that depends on them. Check the tenant inside execute, against the order number the model actually chose, as the code above does.
The SDK offers more ways to give an agent a capability than the quickstart shows, and they differ in where the work runs and who ends up answering the user. These are the four that matter for a first agent.
| Primitive | Where it runs | Who answers the user | Reach for it when |
|---|---|---|---|
| tool() function tool | Your Node process, inside execute | The calling agent | Reading or writing your own systems, such as the ERP database or an internal API |
| Hosted tool: fileSearchTool, webSearchTool, codeInterpreterTool | OpenAI's servers, next to the model, on the Responses API | The calling agent | Searching a vector store you uploaded, web search, or sandboxed code execution |
| agent.asTool() | A nested run of another agent | The calling agent, which keeps control | A specialist step, such as a summary, whose output the parent reuses |
| handoff() | The same run, with a different active agent | The specialist, which takes over the conversation | A different job with different instructions, such as refunds |
Hosted tools are the easiest to add and the easiest to misuse. The tools guide lists them for the Responses API model, which is the SDK default, so switching the process to Chat Completions with setOpenAIAPI takes them off the table. They also run where your code cannot intercept them, which means there is no execute to log, time out or authorise. For a returns policy that is fine, because it is the same public document for every customer. For anything tenant-scoped, write a function tool.
// lib/agent/tools.ts (continued)
import { fileSearchTool } from '@openai/agents';
// Runs on OpenAI's servers against a vector store uploaded once from the
// returns-policy PDF. There is no execute() to log, time out or authorise,
// so nothing tenant-specific belongs in this store.
export const returnsPolicySearch = fileSearchTool(
process.env.RETURNS_POLICY_VECTOR_STORE_ID!,
{ maxNumResults: 3 },
);A handoff is presented to the model as a tool named transfer_to_ followed by the agent's name, so an agent called Refunds agent becomes transfer_to_refunds_agent. Wrapping the target in handoff() adds an inputType, a small zod schema the model fills in when it picks the handoff, and an onHandoff callback that receives the parsed value. I use it for the refund reason, which reaches the audit log before the specialist says a word.
// lib/agent/agents.ts
import { Agent, handoff } from '@openai/agents';
import { z } from 'zod';
import { getSalesOrder, returnsPolicySearch, type ErpContext } from './tools';
const MODEL = process.env.AGENT_MODEL ?? 'gpt-5.6-luna';
export const refundsAgent = new Agent<ErpContext>({
name: 'Refunds agent',
model: MODEL,
instructions:
'You handle refund and return requests. Quote the returns policy you ' +
'found with file search. Never promise a refund amount.',
tools: [returnsPolicySearch, getSalesOrder],
});
// The model fills this in when it picks the handoff; onHandoff gets it parsed.
const RefundReason = z.object({
reason: z.enum(['damaged', 'late_delivery', 'wrong_item', 'other']),
});
// Agent.create, not new Agent: TypeScript then infers the union of
// finalOutput types across every agent this one can hand off to.
export const orderAgent = Agent.create({
name: 'Order status agent',
model: MODEL,
instructions:
'Answer questions about sales orders using get_sales_order. ' +
'If the customer wants money back or a return, hand off.',
tools: [getSalesOrder],
handoffs: [
// The model sees this as a tool called transfer_to_refunds_agent.
handoff(refundsAgent, {
inputType: RefundReason,
toolDescriptionOverride:
'Transfer to the refunds agent when the customer asks for a refund or a return.',
onHandoff: async (_ctx, input) => {
// Lands in the audit log before the specialist says a word.
console.info('refund handoff', input?.reason);
},
}),
],
});Two details from the handoffs guide shape this code. Build the triage agent with Agent.create rather than new Agent, so TypeScript infers the union of final output types across the handoff graph. And by default the receiving agent sees the entire conversation; pass an inputFilter such as removeAllTools from @openai/agents-core/extensions if the specialist should not inherit the previous agent's tool calls. After the run, result.lastAgent tells you which agent produced the answer.
Passing stream: true makes run() return a StreamedRunResult instead of a finished result. For a chat box, toTextStream() is the right level of detail: it emits assistant text only, so tool calls, handoffs and approval requests stay on the server unless you choose to forward them from the full event stream.
// app/api/agent/route.ts
import { run } from '@openai/agents';
import { orderAgent } from '@/lib/agent/agents';
import { getSession } from '@/lib/auth';
// The SDK README lists Node.js 22+, Deno, Bun and Cloudflare Workers.
// The Edge runtime is not on that list, so say nodejs out loud.
export const runtime = 'nodejs';
export async function POST(req: Request) {
const session = await getSession(req);
if (!session) return new Response('Unauthorized', { status: 401 });
const { message } = (await req.json()) as { message: string };
const stream = await run(orderAgent, message, {
stream: true,
// Tenant comes from the session, never from the request body.
context: { tenantId: session.tenantId, userId: session.userId },
maxTurns: 6,
// Aborting this signal is how the streaming guide stops a run early.
signal: req.signal,
});
// Settles after the run AND any persistence work; log failures here.
stream.completed.catch((err) => console.error('agent run failed', err));
// toTextStream() emits assistant text only. Tool calls, handoffs and
// approval requests stay on the server.
return new Response(stream.toTextStream().pipeThrough(new TextEncoderStream()), {
headers: { 'Content-Type': 'text/plain; charset=utf-8' },
});
}The runner loop is simple: call the model, run any tool calls, append the results, call the model again, and stop when the model returns text with no tool calls. Each lap is a turn, maxTurns defaults to 10, and reaching it throws MaxTurnsExceededError. An agent that keeps re-querying an order because a result confused it will hit that ceiling, and the ceiling is what decides whether the mistake costs six model calls or an open-ended bill.
// A nightly job answering queued customer emails: failures should be loud.
import {
run,
MaxTurnsExceededError,
ModelBehaviorError,
ToolCallError,
} from '@openai/agents';
import { orderAgent } from '@/lib/agent/agents';
import type { ErpContext } from '@/lib/agent/tools';
import { alertOps } from '@/lib/ops';
export async function answerEmail(question: string, ctx: ErpContext) {
try {
// Default maxTurns is 10. One turn = one model call, so every
// "call tool, read result, call again" lap spends one.
const result = await run(orderAgent, question, { context: ctx, maxTurns: 6 });
return { answer: result.finalOutput, answeredBy: result.lastAgent?.name };
} catch (err) {
if (err instanceof MaxTurnsExceededError) {
// Six model calls without a final answer: the agent is looping.
return { answer: null, reason: 'max_turns' };
}
if (err instanceof ModelBehaviorError) {
// Malformed output, or a call to a tool that does not exist.
return { answer: null, reason: 'model_behaviour' };
}
if (err instanceof ToolCallError) {
// execute() threw: the database needs attention, not the prompt.
await alertOps('agent tool failure', err);
}
throw err;
}
}
// In user-facing chat, a polite fallback beats a 500:
const result = await run(orderAgent, message, {
context,
maxTurns: 6,
errorHandlers: {
maxTurns: () => ({
finalOutput: 'I could not finish that lookup. A colleague will follow up on this order.',
includeInHistory: false,
}),
},
});errorHandlers turns supported failures into a final answer instead of an exception. The keys are maxTurns, modelRefusal and invalidFinalOutput, with default as a fallback, and a handler returns a finalOutput that must match the agent's outputType plus an optional includeInHistory flag. I prefer the explicit catch in background jobs, where a failure should reach a person, and the handler in user-facing chat, where a polite fallback beats an HTTP 500.
Passing maxTurns: null disables the limit entirely. The running-agents guide documents the option, but on an agent with write tools it removes the only bound on how many times the model can call them in a single run.
Checking only the final message misses the failures that matter in an agent: the right answer reached through the wrong tool, or the right tool called with the wrong order number. promptfoo's openai:agents provider runs an agent exported from a TypeScript file, passes each test's vars into the run context, and with tracing enabled lets you assert on the path itself through trajectory:tool-used, trajectory:tool-args-match, trajectory:tool-sequence and trajectory:goal-success.
# promptfooconfig.yaml
# Run the project-local binary (npx promptfoo eval) so the eval and the
# agent load the same @openai/agents installation.
prompts:
- '{{query}}'
providers:
- id: openai:agents:order-agent
config:
agent: file://./lib/agent/order-agent.eval.ts # default export = orderAgent
model: gpt-5.6-luna # pin the baseline; do not inherit the SDK default
maxTurns: 6
tracing: true # trajectory:* assertions read the trace
tests:
- vars:
tenantId: tenant-acme # test vars arrive in runContext.context
userId: eval-user
query: 'Has SO-2026-00412 shipped yet?'
assert:
- type: trajectory:tool-used
value: get_sales_order
- type: trajectory:tool-args-match
value:
name: get_sales_order
args:
orderNumber: 'SO-2026-00412'
- type: trajectory:goal-success
value: 'Tell the user whether SO-2026-00412 has shipped'Two things on the promptfoo page are worth reading before the first run. Its install line pins @openai/agents to ^0.11.8, and a caret on a 0.x version stops short of 0.12.0, while the latest release on npm is 0.18.0; since the page also tells you to use the project-local promptfoo so the eval and the agent share one SDK installation, check which version your lockfile actually resolves. And mock mode, which replaces tool results for deterministic tests, fails closed for explicit handoff() objects and for hosted tools, both of which this agent uses. So I split the suite three ways.
@openai/agents bundles @openai/agents-core, @openai/agents-openai and @openai/agents-realtime at the same version, 0.18.0 today, and depends on openai ^7.2.0. If the app already calls the openai package directly elsewhere, check that both paths resolve to a single copy.
The TypeScript SDK is complete enough that a NestJS or Next.js team has no reason to run a Python sidecar for a tool agent. The rules I carry from this build are short: authorise in execute, not in isEnabled; keep tenant data out of hosted tools; pin the model; cap maxTurns and decide what happens when it trips; and test the trajectory, not just the reply.
Sources