AI
OpenAI Skills API: SKILL.md Bundles in the Hosted Shell
October 202612 min read

It is the set of endpoints under /v1/skills that store reusable skill bundles: folders anchored by a SKILL.md manifest with a name and description in YAML front matter, plus optional scripts, references and assets. Skills are versioned on the server and attached to the Responses API shell tool, or discovered by the Agents API, so the model can follow a packaged procedure only when it needs one.
Send a POST to /v1/skills either as multipart, with one form part per file and the relative path in the filename, or as a single zip file. The documented limits are 50 MB for the zip, 25 MB uncompressed and 500 files per version, with exactly one SKILL.md per bundle. New versions of an existing skill go to the versions endpoint of that skill.
latest_version always points at the newest upload, while default_version is the version a request receives when it does not name one, and it only changes when you set it explicitly. This lets you upload and evaluate a new version without affecting production. OpenAI recommends pinning an explicit version in production requests.
OpenAI documents the hosted shell as currently based on Debian 12 with Python 3.11, Node.js 22.16, Java 17, PHP 8.2, Ruby 3.1 and Go 1.23, noting the base may change. Commands run without sudo or an interactive TTY, and /mnt/data is the supported location for downloadable output. Outbound network access is off unless you enable it.
Each domain_secrets entry in the network_policy binds a secret name and value to one allowed domain. The model and the container only see a placeholder such as $API_KEY, and OpenAI's auth-translation sidecar inserts the real value only on requests to that domain. It protects the credential itself, but you still need a tight allowlist and approvals because fetched content can carry prompt injection.

Key Takeaway
The OpenAI Skills API stores versioned SKILL.md bundles at /v1/skills and mounts them into the Responses API hosted shell, a Debian 12 container with no network by default. Pin a skill version in production, promote default_version like a deploy, and pass API keys through domain_secrets so the model only ever sees a placeholder.
The OpenAI Skills API answers a question that comes up constantly in ERP reporting jobs: where does a long, repeatable procedure live? An accounts-receivable aging report is not one tool call. It is a script, a set of bucketing rules, an output format and a list of things not to do, and pasting all of that into a system prompt on every request is both wasteful and impossible to version. OpenAI's answer, shipped to the Responses API in February 2026, is a skill: a folder anchored by a SKILL.md file, uploaded once, versioned on the server and mounted into a container the model can run commands in.
This post walks the whole path using OpenAI's own documentation as the authority: the manifest format, the upload endpoints and their limits, how default_version and latest_version interact, the four ways a skill reaches a runtime, what the hosted shell actually contains, and the network controls that decide whether a skill can call your ERP without the model ever holding the key. Claude Code skills use the same file format, and the sibling posts on the Agents API cover the session side; this one is about the upload, versioning and container model.
A skill is a directory with exactly one SKILL.md at its root. The file opens with YAML front matter holding two required fields, name and description, followed by Markdown instructions. Supporting files sit beside it by convention: scripts for repeatable actions, references for background material, assets for templates and sample data. OpenAI follows the open Agent Skills specification here, which caps name at 64 characters of lowercase letters, digits and single hyphens, requires it to match the folder name, and allows up to 1024 characters of description.
ar-aging-report/
├── SKILL.md
├── requirements.txt
├── scripts/
│ └── aging.py
└── references/
└── bucket-rules.md
# ar-aging-report/SKILL.md
---
name: ar-aging-report
description: Build an accounts-receivable aging report (0-30, 31-60, 61-90, 90+ days)
from an exported invoices CSV. Use when the user uploads open invoices and asks
for aging, overdue buckets or collection priorities. Do not use for AP or GL data.
---
# AR aging report
## How to run
python scripts/aging.py --input /mnt/data/invoices.csv --outdir /mnt/data/aging
## Rules
- Bucket by due_date, not invoice_date. See references/bucket-rules.md.
- If a dependency is missing, report it. Do not attempt a network install.
## Outputs
- /mnt/data/aging/report.md
- /mnt/data/aging/aging.csvThe description is the part that does the work. When a skill is attached, the platform adds each skill's name, description and path to the prompt context, and the model only reads the full SKILL.md through that path once it decides the skill applies. So the description has to say when to use the skill and, just as usefully, when not to. OpenAI's cookbook recommends negative examples for exactly this reason, and the specification suggests keeping the body under 500 lines and pushing detail into reference files the model opens on demand.
There are two upload shapes for POST /v1/skills. Multipart sends one form part per file, with the filename field carrying the relative path so the folder structure survives. A zip sends the whole bundle as a single part. I would use the zip in CI, because it is one artifact you can hash, store and re-upload, and because it is the shape the documented size limit is stated against.
# Option 1: multipart, one part per file. The filename= carries the folder path,
# so the bundle keeps its layout on the server.
curl --fail-with-body 'https://api.openai.com/v1/skills' \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-F 'files[]=@./ar-aging-report/SKILL.md;filename=ar-aging-report/SKILL.md;type=text/markdown' \
-F 'files[]=@./ar-aging-report/scripts/aging.py;filename=ar-aging-report/scripts/aging.py;type=text/plain'
# Option 2: one zip. Easier in CI, and it is what the 50 MB limit applies to.
(cd ar-aging-report/.. && zip -r ar-aging-report.zip ar-aging-report)
curl -X POST 'https://api.openai.com/v1/skills' \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-F 'files=@./ar-aging-report.zip;type=application/zip'
# The response carries the skill id plus default_version and latest_version.
# Store the id in config; it is what every request references.The 25 MB uncompressed ceiling is the one that bites. A skill that vendors a Python wheel or ships a sample dataset can be well under 50 MB zipped and still fail. Keep large inputs out of the bundle and pass them as files at request time; the skill should describe how to process data, not carry it.
Every skill carries two pointers. latest_version always tracks the newest upload, made with POST /v1/skills/skill_id/versions. default_version is the one a request gets when it does not name a version, and it only moves when you explicitly set it on the skill. That separation is the whole release model: upload freely, evaluate the new version by pinning it in a test request, then promote it.
# Ship a new version. This moves latest_version, NOT default_version.
curl -X POST "https://api.openai.com/v1/skills/$SKILL_ID/versions" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-F 'files=@./ar-aging-report.zip;type=application/zip'
# Promote it once the eval run passes. Requests that omit "version"
# switch over at this moment, so treat it like a deploy.
curl -X POST "https://api.openai.com/v1/skills/$SKILL_ID" \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-d '{"default_version": 3}'
# Rollback is the same call with the old number: {"default_version": 2}A skill_reference accepts an integer version or the string latest. OpenAI's cookbook advises pinning versions in production rather than following latest, and pinning the model alongside the skill so a run is reproducible. I would go one step further and treat a default_version change as a deploy with a changelog entry, because every caller that omits the version changes behaviour at that instant, with no code change anywhere to point at.
Deletion has two rules worth knowing before you script cleanup. You cannot delete the version that is currently default_version until you point default_version somewhere else, and deleting the last remaining version deletes the skill itself, id included. A careless loop that prunes old versions can take a production skill_id with it.
Uploading is optional. The same SKILL.md bundle can reach the model through four routes, and the choice decides where the files live and how long they last.
| Route | How the skill is attached | Lifetime | Use it when |
|---|---|---|---|
| Hosted shell, container_auto | skill_reference with skill_id and optional version in tools.environment.skills | A fresh container per request, provisioned by OpenAI | One-shot jobs such as generating a report from an upload |
| Hosted shell, container_reference | Skills mounted on a container created through the Containers API, then referenced by container_id | Until deleted or expired, 20 minutes after last activity in the documented example | Multi-turn work that needs files and installed packages to persist |
| Inline bundle | Base64-encoded zip in the request, type inline, with name and description | Only the container it is sent with | Prototyping, or skills generated per tenant that you do not want stored |
| Local shell | name, description and a filesystem path; skill_reference is not supported | Your own machine or server | Data that must stay on premises; you execute the commands yourself |
from openai import OpenAI
client = OpenAI()
response = client.responses.create(
model="gpt-6-astra",
tools=[
{
"type": "shell",
"environment": {
"type": "container_auto",
"skills": [
# Wrong in production: no version means default_version,
# which someone can move without touching this code.
# {"type": "skill_reference", "skill_id": AGING_SKILL_ID},
# Right: pin the version you evaluated.
{"type": "skill_reference", "skill_id": AGING_SKILL_ID, "version": 3},
# A curated first-party skill, referenced by its id.
{"type": "skill_reference", "skill_id": "openai-spreadsheets", "version": "latest"},
],
},
}
],
input="Build the AR aging report from the attached invoices export.",
)
print(response.output_text)The commented-out line in that request is the mistake I would expect most teams to make first: it works, it passes review, and it silently follows whatever default_version becomes next month. The curated openai-spreadsheets skill shows the other end of the trade, where following latest is a reasonable choice because OpenAI maintains it and you are not the one evaluating each release. Remember too that hosted skills are discarded when the container expires, so nothing you mount is a durable store.
OpenAI documents the hosted runtime as currently based on Debian 12, with Python 3.11, Node.js 22.16, Java 17, PHP 8.2, Ruby 3.1 and Go 1.23 preinstalled, and says the base may change over time. Commands run without sudo and without an interactive TTY, so a skill script that prompts for input or installs system packages will stall or fail. /mnt/data is always present and is the supported location for anything the user should be able to download; artifacts there are retrieved with the same container files APIs used by code interpreter.
// What the model emits: a batch of commands with its own limits.
{
"type": "shell_call",
"call_id": "call_9d14...",
"action": {
"commands": ["python scripts/aging.py --input /mnt/data/invoices.csv --outdir /mnt/data/aging"],
"timeout_ms": 120000,
"max_output_length": 4096
},
"status": "in_progress"
}
// What comes back. Log the exit code, not just stdout: a non-zero exit
// with a confident final answer is the failure you want an alert on.
{
"type": "shell_call_output",
"call_id": "call_9d14...",
"output": [
{ "stdout": "...", "stderr": "...", "outcome": { "type": "exit", "exit_code": 0 } }
]
}Each shell_call can carry several commands with a per-call timeout_ms and max_output_length, and each result reports stdout, stderr and an exit outcome. When you create a container yourself you also choose memory_limit and expires_after, for example one gigabyte and 20 minutes from last_active_at, and you can DELETE it as soon as the job finishes instead of paying for idle time. The shell tool is only available through the Responses API, not Chat Completions.
Write skill scripts to print one short, structured summary line on success and a clear error on failure. Output is truncated at max_output_length, and the model decides its next step from what survives the cut. A script that dumps ten thousand rows to stdout teaches the model nothing and wastes the turn.
Hosted containers have no outbound network by default. Turning it on takes two layers: an admin sets an organisation allowlist in the dashboard, and each request adds a network_policy of type allowlist inside the shell environment. The request list can only narrow the org list; a request naming a domain outside it fails rather than quietly widening access, which is the behaviour you want from a security control.
"tools": [
{
"type": "shell",
"environment": {
"type": "container_auto",
"network_policy": {
"type": "allowlist",
// Must be a subset of the org allowlist set in the dashboard,
// or the request fails rather than silently widening access.
"allowed_domains": ["pypi.org", "files.pythonhosted.org", "erp.example.co.id"],
"domain_secrets": [
{
"domain": "erp.example.co.id",
"name": "ERP_API_KEY",
// The model and the container only ever see $ERP_API_KEY.
// The real value is applied by OpenAI's auth sidecar on the way
// out, and only for this domain.
"value": "<read from your secret manager at request time>"
}
]
}
}
}
]domain_secrets is what makes a skill that calls a private API tolerable. Each entry binds a name and a value to one domain. The model and the runtime see only the placeholder, such as $ERP_API_KEY, and OpenAI's auth-translation sidecar substitutes the real value only on requests to the approved destination. A skill script can therefore authenticate to your ERP while a prompt-injected model that prints its environment prints a placeholder, and a request to any other host carries nothing useful.
domain_secrets protects the key, not the data. OpenAI's guide warns that any content fetched over the network may contain hidden instructions aimed at the model, and its skills guide lists prompt injection and data exfiltration as the main risks. Keep the allowlist to the domains a skill actually needs, prefer read-only credentials, and put write actions behind explicit approval.
The Agents API discovers skills differently. Instead of a list of references, a session's environment names capability_directories, and the platform searches them for SKILL.md files. The paths must be absolute with no dot segments, must already exist in the sandbox, and a session can register at most 32 of them.
{
"environment": {
"type": "self_hosted",
"workspace_directory": "/workspace",
"capability_directories": [
"/workspace/capabilities/finance",
"/workspace/capabilities/inventory"
]
}
}
// Rules the platform enforces:
// absolute paths only, no "." or ".." segments
// at most 32 directories per session
// each directory must already exist in the sandbox
// Any SKILL.md found under them is registered as a skill.That limit shapes how you organise skills. Thirty-two directories is plenty if each one is a domain, such as finance, inventory or engineering, holding several skills, and tight if every skill gets its own top-level directory. Grouping by domain also gives you a natural unit of permission: a sales agent's session simply never mounts the finance directory, so those skills are not in its prompt at all.
OpenAI's cookbook draws the line clearly: the system prompt holds global, always-on behaviour; tools do something in the world with side effects; skills are packaged procedures the model invokes conditionally, with scripts and templates run in a sandbox. An AR aging report, a month-end reconciliation checklist or a standard export format are skills. Posting a journal entry is a tool, with its own approval gate. My shipping order for a new skill is:
The same SKILL.md format is used by OpenAI, Claude Code and other tools that follow the Agent Skills specification, so a well-written skill folder is portable. What does not port is the plumbing around it: skill ids, version pointers, container modes and network policy are specific to OpenAI's API.
The useful mental model is that a skill is a deployable artifact, not a prompt. It has an id, numbered versions, a promotion step, a runtime with known languages and limits, and a network boundary you configure. Treat it with the same discipline as any other deploy: pin what you tested, promote on purpose, and never let a credential reach the model as text.