DevOps
Codex SDK Tutorial: Automate Coding Tasks in TypeScript
October 202611 min read

The Codex SDK is a library that runs the Codex agent from your own code. The TypeScript package @openai/codex-sdk wraps the codex CLI: it spawns codex exec and exchanges JSONL events over stdin and stdout. You get the same agent as the CLI, but driven by a script, which suits CI jobs and internal tools.
For TypeScript run npm install @openai/codex-sdk, which needs Node.js 18 or newer and runs server-side only. For Python run pip install openai-codex, which needs Python 3.10 or newer. Both reuse existing Codex authentication, and the TypeScript client also accepts an apiKey option that is passed to the CLI as CODEX_API_KEY.
Read thread.id after the first turn has started, or capture thread_id from the thread.started event when streaming, and store it. Later call codex.resumeThread(threadId) and keep calling run(). Threads are persisted on local disk in the Codex sessions folder, so the resume must happen on a machine that has that folder.
Use read-only for triage, review and explanation, and workspace-write when the agent should draft a fix inside the checkout. Avoid danger-full-access on a CI runner because it holds tokens and deploy keys. Set approvalPolicy to never for unattended runs, since a script cannot answer an approval prompt.
If you pipe npm test into tee without declaring a shell, GitHub Actions runs bash -e without pipefail, so the step reports the exit code of tee, which is zero. Add shell: bash to the step, which runs bash with -eo pipefail. Then if: failure() fires and the Codex triage step actually runs.

Key Takeaway
The Codex SDK (@openai/codex-sdk for TypeScript, openai-codex for Python) runs the Codex agent from your own code by spawning the Codex CLI. Start a thread, call run() or runStreamed(), set sandboxMode and approvalPolicy explicitly, and request JSON with outputSchema so a CI job can act on the result instead of parsing prose.
The first time I wanted Codex inside a pipeline rather than inside my terminal, the question was concrete: a pull request has just gone red, nobody is looking at it yet, and the log is four hundred lines of install noise with one stack trace at the bottom. Could an agent read that log, open the code, say what broke, and stop there unless I asked for more? The Codex SDK is the piece that makes that a script instead of a chat session.
This Codex SDK tutorial builds exactly that job in TypeScript: a first thread, the sandbox and approval settings that matter in CI, streamed events, a structured triage result, a GitHub Actions workflow, and a resumed thread that drafts the fix. Every option name below comes from the SDK's own type definitions in the openai/codex repository and from OpenAI's SDK documentation, not from memory. It is a build guide; for choosing between Codex and Claude Code as a daily tool, the comparison post on this site covers that.
The SDK is not a new HTTP API. The TypeScript package wraps the codex CLI from @openai/codex: it spawns codex exec with experimental JSON output and exchanges JSONL events over stdin and stdout. That one design decision explains most of its behaviour. It is server-side only, it needs a filesystem and a working directory, it inherits the CLI's configuration, and threads persist on local disk in the sessions folder under the Codex home directory.
| Aspect | TypeScript | Python |
|---|---|---|
| Package and install | npm install @openai/codex-sdk | pip install openai-codex |
| Runtime | Node.js 18 or newer | Python 3.10 or newer, with a pinned Codex CLI runtime |
| Start and continue | startThread(), then run() repeatedly; resumeThread(id) | thread_start(), then run() repeatedly |
| Async style | Promises, plus runStreamed() as an async generator | Sync Codex client, or AsyncCodex with async with |
| Sandbox setting | sandboxMode string on the thread options | Sandbox enum, overridable per run() |
Because the SDK drives the CLI, the agent does real work on the machine it runs on: it executes shell commands, applies patches and reads files in the working directory. That is the point, and it is also why the sandbox section below is not optional reading.
Version check before you pin: at the time of writing the npm registry lists @openai/codex-sdk 0.159.3, which depends on @openai/codex at exactly the same version. Upgrade the two together by upgrading the SDK, never the CLI on its own, or the JSONL event shapes the SDK parses can drift away from what the CLI emits.
A thread is one conversation with the agent. Each run() call is a turn: the agent plans, runs commands, edits files if allowed, and finishes with a final message. The returned turn carries finalResponse, the full list of items, and token usage, so you get both the answer and the audit trail from one call.
npm install @openai/codex-sdk # pulls the matching @openai/codex CLI as a dependency
// first-thread.ts — Node 18+, server-side only
import { Codex } from "@openai/codex-sdk";
// The SDK spawns `codex exec --experimental-json` and reads JSONL from stdout.
// apiKey is forwarded to the CLI as CODEX_API_KEY.
const codex = new Codex({ apiKey: process.env.CODEX_API_KEY });
const thread = codex.startThread({
workingDirectory: process.cwd(), // must be a Git repo unless skipGitRepoCheck: true
sandboxMode: "read-only",
approvalPolicy: "never", // a script cannot answer an approval prompt
});
const turn = await thread.run("Explain how the test suite in this repo is organised.");
console.log(turn.finalResponse); // the agent's last message
console.log(turn.items.length); // every command, file change and message in the turn
console.log(turn.usage); // input, cached, output and reasoning tokens
// thread.id is null until the first turn has started — read it after run().
console.log("resume later with", thread.id);Two details here cost time if you skip them. The working directory must be a Git repository unless you pass skipGitRepoCheck, which is a deliberate guard so the agent's edits are always recoverable with git; in CI the checkout already satisfies it. And thread.id stays null until the first turn has started, so code that saves the id before calling run() saves nothing.
The sandbox decides what the agent's commands may touch; the approval policy decides when it stops to ask a human. In a terminal the defaults are tuned for a person watching. In CI nobody is watching, so set both explicitly on every thread rather than inheriting whatever the runner's Codex configuration says.
| TypeScript sandboxMode | Python preset | Where I use it |
|---|---|---|
| read-only | Sandbox.read_only | Triage, review and explanation. The agent can read and run read-only commands but cannot change the checkout. |
| workspace-write | Sandbox.workspace_write | Drafting a fix or a migration. Writes stay inside the working directory; network is a separate switch. |
| danger-full-access | Sandbox.full_access | Almost never on a CI runner, which holds deploy keys and tokens. Only inside a throwaway container you own. |
The approval policy accepts never, on-request, on-failure and untrusted. A script cannot click approve, so unattended threads use never and rely on the sandbox for safety. Network access is its own flag, networkAccessEnabled, and I leave it off for triage: diagnosing a failure from the code and the log should not need the internet, and a prompt-injected log line should not be able to reach it either.
run() buffers everything until the turn ends. For a ten-minute CI task that means ten minutes of silence in the job log. runStreamed() returns an async generator of typed events instead, so the job log shows each command as the agent runs it.
const { events } = await thread.runStreamed("Fix the failing test in src/invoice.test.ts");
for await (const event of events) {
switch (event.type) {
case "thread.started":
// The earliest moment the id exists. Persist it before anything can crash.
await saveThreadId(event.thread_id);
break;
case "item.completed":
if (event.item.type === "command_execution") {
console.log(`$ ${event.item.command} -> exit ${event.item.exit_code}`);
}
if (event.item.type === "file_change") {
for (const change of event.item.changes) console.log(`${change.kind} ${change.path}`);
}
break;
case "turn.completed":
console.log("tokens in/out", event.usage.input_tokens, event.usage.output_tokens);
break;
case "turn.failed":
// run() throws on this; runStreamed() hands it to you, so you must throw yourself.
throw new Error(event.error.message);
}
}That last point is the one I would underline. A streaming loop that only handles item.completed and turn.completed exits cleanly on a failed turn, and the CI step goes green with no result written. Treat turn.failed as an exception, every time.
A triage paragraph is pleasant to read and useless to the next workflow step. Passing outputSchema on a turn makes the final message a JSON string that conforms to the schema, so the job can branch on a verdict field. The schema below forces one of four verdicts, because the most useful thing triage can tell you is whether to touch the code at all.
// scripts/codex-triage.ts — runs only after the test step has failed
import { readFileSync, writeFileSync } from "node:fs";
import { Codex } from "@openai/codex-sdk";
const TRIAGE_SCHEMA = {
type: "object",
properties: {
verdict: { type: "string", enum: ["product_bug", "test_bug", "flaky", "environment"] },
failingTests: { type: "array", items: { type: "string" } },
rootCause: { type: "string" },
suggestedFix: { type: "string" },
},
required: ["verdict", "failingTests", "rootCause", "suggestedFix"],
additionalProperties: false,
} as const;
const TEN_MINUTES_MS = 10 * 60_000;
const LOG_TAIL_CHARS = 20_000; // the tail holds the stack traces; the head is install noise
const codex = new Codex({ apiKey: process.env.CODEX_API_KEY });
const thread = codex.startThread({
sandboxMode: "read-only", // triage reads; it never writes
approvalPolicy: "never",
networkAccessEnabled: false,
});
const logTail = readFileSync("test-output.log", "utf8").slice(-LOG_TAIL_CHARS);
const controller = new AbortController();
const timer = setTimeout(() => controller.abort(), TEN_MINUTES_MS);
try {
const turn = await thread.run(
"The test suite failed. Find the root cause without modifying any file.\n\n" + logTail,
{ outputSchema: TRIAGE_SCHEMA, signal: controller.signal },
);
// With outputSchema set, finalResponse is a JSON string that matches the schema.
const triage = JSON.parse(turn.finalResponse);
writeFileSync(
"triage.json",
JSON.stringify({ threadId: thread.id, usage: turn.usage, ...triage }, null, 2),
);
} finally {
clearTimeout(timer);
}The AbortSignal in TurnOptions is the cost ceiling. Without it a stuck turn runs until the job's own timeout, burning tokens the whole way; with it the script decides the budget. I also send only the tail of the log, because the stack traces sit at the end and the dependency install output at the top adds tokens without adding evidence.
The workflow runs the tests, and only if they fail does it run the triage script and upload triage.json as an artifact. The step that matters most is not the Codex step. It is the shell declaration on the test step, because the obvious way to capture a log silently defeats the failure condition.
# .github/workflows/test.yml
name: test
on: pull_request
jobs:
test:
runs-on: ubuntu-latest
permissions:
contents: read
steps:
- uses: actions/checkout@v4
- uses: actions/setup-node@v4
with:
node-version: 22
- run: npm ci
# Wrong: the default shell is `bash -e` WITHOUT pipefail, so tee's exit code
# (0) wins and a failing suite turns the step green.
# - run: npm test 2>&1 | tee test-output.log
# Right: an explicit `shell: bash` runs with -eo pipefail.
- run: npm test 2>&1 | tee test-output.log
shell: bash
- name: Triage with Codex
if: failure()
run: npx tsx scripts/codex-triage.ts
env:
CODEX_API_KEY: ${{ secrets.CODEX_API_KEY }}
- uses: actions/upload-artifact@v4
if: failure()
with:
name: codex-triage
path: triage.jsonThe pipe trap: GitHub Actions runs an unspecified shell as bash -e without pipefail, so npm test piped into tee reports tee's exit code, which is zero. The suite fails, the step passes, if: failure() never fires, and the agent never runs. Declaring shell: bash switches the step to bash with -eo pipefail and the failure propagates.
Drafting the fix is a second turn on the same thread, not a new conversation. resumeThread accepts fresh thread options, so the triage thread that ran read-only can be resumed with workspace-write and already knows what it found. Sessions live on the runner's disk, so this works within one job; across jobs you would need to carry the Codex sessions directory along.
// scripts/codex-draft-fix.ts — same job, after a human-readable triage exists
import { readFileSync } from "node:fs";
import { Codex } from "@openai/codex-sdk";
const { threadId, verdict } = JSON.parse(readFileSync("triage.json", "utf8"));
if (verdict !== "product_bug" && verdict !== "test_bug") process.exit(0); // never "fix" a flake
const codex = new Codex({ apiKey: process.env.CODEX_API_KEY });
// Same conversation, wider permissions: resumeThread takes fresh ThreadOptions.
const thread = codex.resumeThread(threadId, {
sandboxMode: "workspace-write",
approvalPolicy: "never",
});
await thread.run(
"Apply the fix you proposed. Keep the diff minimal, then run the failing tests again.",
);
// The workflow, not the agent, commits and opens the PR: git switch -c, git commit, gh pr create.The Python SDK covers the same ground with Python names: Codex as a context manager, thread_start() instead of startThread(), final_response instead of finalResponse, and an AsyncCodex client for asyncio code. Its most useful difference is that the sandbox can be overridden on an individual run() call, not only when the thread starts.
# pip install openai-codex (Python 3.10+)
from openai_codex import Codex, Sandbox
with Codex() as codex:
thread = codex.thread_start(sandbox=Sandbox.workspace_write)
thread.run("Migrate the date helpers in src/utils from moment to date-fns.")
# Same thread, narrower sandbox for this one run: the reviewer cannot edit
# the code it is reviewing.
review = thread.run("Review the diff only. List anything risky.", sandbox=Sandbox.read_only)
print(review.final_response)That per-run override makes a clean write-then-review loop: the same thread that made the change reviews its own diff under read_only, so the review pass physically cannot edit what it is judging. In TypeScript the equivalent is resuming the thread with different options, as in the CI example. The Python README also documents explicit login helpers, including API-key and device-code login, which matter on a headless server.
At DevDay on 29 September 2026 OpenAI gave Codex reusable cloud environments that can be shared across devices with approved settings and permissions, a CLI with voice control and a new /agents view for delegating and tracking several tasks, improvements to session resuming and worktrees, and a new code review experience in the ChatGPT desktop app that can review automatically while you are offline.
None of that replaces the SDK for pipeline work. The SDK still drives a local agent in a directory you control, which is exactly what a CI runner is. The cloud environments and the hosted code review are the better fit when the work should not run on your infrastructure at all, and the /agents view is for a person juggling tasks in a terminal. Pick by where the code should execute, not by which feature is newest.
Start every new automation in read-only with outputSchema, and keep it there for a week. Reading the verdicts tells you whether the agent understands your repository before you let it write a single line, and that week costs nothing to roll back.
The rule I carry from this: an unattended agent is a CI step like any other, so give it the same discipline. Set the sandbox and approval policy on every thread, ask for JSON instead of prose, put a ceiling on every turn with an AbortSignal, treat turn.failed as an exception, and keep commits and pushes in the workflow where reviewers can see them. Start read-only, earn the write access.
Sources