AI
OpenAI DevDay 2026 for Developers: Sol, Ultrafast, Agents API
October 202611 min read

At DevDay on 29 September 2026, OpenAI released GPT-6.1 Sol, an Ultrafast service tier for GPT-6 Astra, and computer use in the Agents API public beta. It also announced Amazon Bedrock Managed Agents, a Decisions API in limited preview, Codex Cloud and CLI updates, and plugin extensions for ChatGPT.
GPT-6.1 Sol costs $2 per million input tokens and $10 per million output tokens, against $10 and $50 for GPT-6 Astra, so its list price is one-fifth of Astra's. Prompts above 272K input tokens cost 2x on input and 1.5x on output for both. Per-task cost can differ because Astra often uses fewer output tokens.
Set model to gpt-6-astra and service_tier to ultrafast on a Responses API request, ideally over a persistent WebSocket. The Ultrafast price for Astra is $60 input and $300 output per million tokens, six times the standard rate. It supports only US data residency and global processing.
GPT-6.1 Sol rejects the none and minimal reasoning efforts, so move to low, and you must remove temperature, top_p and top_logprobs while reasoning is on. Tool calling requires the Responses API rather than Chat Completions. If you come from GPT-5.5 or earlier, replace prompt_cache_retention with prompt_cache_options and a 30m ttl.
No, the Decisions API is in limited preview and OpenAI has published no pricing for it. It uses the Luna model to choose from a finite set of predefined answers for routing, classification or picking an agent's next action. Until it leaves preview, keep your own classifier in production.

Key Takeaway
OpenAI DevDay 2026, held on 29 September 2026, gave developers GPT-6.1 Sol at one-fifth of GPT-6 Astra's token price, an Ultrafast service tier billed at six times Astra's standard rate, computer use in the Agents API public beta, Bedrock Managed Agents on AWS, and a Decisions API that is still a limited preview.
The OpenAI DevDay 2026 keynote on 29 September announced more than twenty things, and by the next morning my feeds held three different versions of what had shipped. One recap said Ultrafast was available to everyone, another said it was only for a new Pro 500 plan, and a third listed GPT-6.1 Sol under Ultrafast when the developer docs still name a different Sol model there. For anyone who builds on the OpenAI API, the question is not what was announced but what you can call today, at what price, and what will break if you switch.
This is the developer-side roundup. I read the official community recap against the API changelog, the GPT-6.1 Sol model page, the GPT-6 migration guide, the Ultrafast guide and the Bedrock Managed Agents page, and only kept a claim when the docs backed it. Dots, the consumer headline, gets one section here because it has its own deep-dive post on this site. The rest is models, pricing arithmetic, migration traps and a short plan for the quarter.
The useful way to read the keynote is by status. Some things are generally available with a model ID and a price, some are public beta with a stable shape, and some are previews that you should not put on a roadmap yet. This table sorts the developer-facing announcements by that axis, using the wording in OpenAI's own recap and changelog.
| Announcement | Status on 1 October 2026 | What it changes for you |
|---|---|---|
| GPT-6.1 Sol | Released 29 September as gpt-6.1-sol in Responses and Chat Completions | Near-Astra quality at one-fifth of Astra's standard token price; the new default to evaluate |
| Ultrafast service tier | Available for gpt-6-astra at low default rate limits; Sol is preview-only | Up to 6x faster in the API by OpenAI's figure, at 6x Astra's standard price |
| Agents API computer use | Added 29 September to the Agents API, which has been public beta since 10 September | Hosted agents can operate websites in an OpenAI-hosted browser, with approvals and sign-in handled by your app |
| Amazon Bedrock Managed Agents | Available, with access, Regions and pricing set by AWS | The Agents API harness and model inference running inside Amazon Bedrock, authenticated with IAM |
| Decisions API | Limited preview, no public pricing | Luna picks one of your predefined answers for routing and classification; do not build on it yet |
| Codex Cloud, CLI, Code Review, Security Cloud | Shipping in ChatGPT and the Codex CLI | Cloud tasks that survive a closed laptop, voice and an agents view in the CLI, scheduled vulnerability scans |
| Plugin extensions and MCP events | Documented in the plugins docs; web support for Free and Go users is listed as coming soon | Plugins get sidebar entries, side panels and file viewers; MCP events let connected apps trigger automations |
| Private Inference | Preview planned for this fall | Nothing to adopt yet; ZDR with Private Safety Processing is the available option today |
OpenAI also reported 45% lower API time to first token and over 30% faster tool calls and workflows. Those are the vendor's numbers from the keynote, not something I measured, and they describe platform averages rather than your traffic. Treat them as a reason to re-run your latency dashboards this month, not as a figure to quote in a proposal.
The GPT-6 family now has three tiers that OpenAI's own guide names plainly: Astra for the highest intelligence, GPT-6.1 Sol for balanced speed, cost and intelligence, and Luna for focused, high-volume work. GPT-6.1 Sol follows the GPT-6 Sol model released one week earlier on 22 September, at the same input and output price but with half the cached-input price, and OpenAI asks gpt-6-sol users to read the migration notes before switching.
| Model | Standard price per 1M tokens | Context and output | Notes from the docs |
|---|---|---|---|
| gpt-6-astra | $10 input, $1 cached, $50 output | 1,050,000 context, 128,000 max output | No none reasoning effort; tool calling needs Responses; Ultrafast and Fast mode available |
| gpt-6.1-sol | $2 input, $0.10 cached, $2.50 cache write, $10 output | 1,050,000 context, 922,000 max input, 128,000 max output | Efforts low to max, no none or minimal; US and EU data residency |
| gpt-6-sol | $2 input, $0.20 cached, $10 output | See the model catalog | Released 22 September; review the migration notes before moving to 6.1 Sol |
| gpt-6-luna | $0.10 input, $0.01 cached, $0.50 output | See the model catalog | Supports the none reasoning effort; the model behind the Decisions API preview |
| Long prompts, all GPT-6 models | 2x input and cache rates, 1.5x output | Applies to the full request once input exceeds 272K tokens | A long-context job crosses the line silently; watch input size |
| Batch, Flex and Fast | Batch and Flex 50% of Standard; Fast mode 2x | Priced off the Standard rate for each model | Fast mode is unavailable with EU data residency |
Sol's price is exactly one-fifth of Astra's on both input and output, which is where the keynote's claim comes from. The catch is that per-token price is not per-task cost. OpenAI's GPT-6 guide says Astra reached stronger results using substantially fewer output tokens in several evaluations, so the ratio on your workload could be smaller than five. The arithmetic below uses identical token counts on purpose, to show the list-price gap before you measure anything.
// List-price arithmetic for ONE job: 200K input tokens, 20K output tokens,
// prompt under the 272K threshold, no cache hits. Not a benchmark.
const PRICES_PER_MILLION = {
"gpt-6-astra": { input: 10, output: 50 },
"gpt-6.1-sol": { input: 2, output: 10 },
"gpt-6-luna": { input: 0.1, output: 0.5 },
"gpt-6-astra/ultrafast": { input: 60, output: 300 },
};
function jobCost(model: keyof typeof PRICES_PER_MILLION, inTok: number, outTok: number) {
const p = PRICES_PER_MILLION[model];
return (inTok / 1e6) * p.input + (outTok / 1e6) * p.output;
}
jobCost("gpt-6-astra", 200_000, 20_000); // 3.00
jobCost("gpt-6.1-sol", 200_000, 20_000); // 0.60
jobCost("gpt-6-luna", 200_000, 20_000); // 0.03
jobCost("gpt-6-astra/ultrafast", 200_000, 20_000); // 18.00
// The trap: same token counts on every row. OpenAI says Astra often finishes
// with far fewer output tokens, so the only honest number is one you measure
// on your own tasks. Over 272K input, every row costs 2x input, 1.5x output.Most migration failures here are request-shape failures, not quality failures. GPT-6.1 Sol and Astra do not accept the none reasoning effort, and minimal is gone too, so a pipeline tuned for the cheapest GPT-5.x setting has to move to low. While reasoning is on, temperature, top_p and top_logprobs must be removed. And if you are coming from GPT-5.5 or earlier, the caching field changes name.
import OpenAI from "openai";
const client = new OpenAI();
// Wrong: a GPT-5.x-era request copied onto GPT-6.1 Sol.
await client.responses.create({
model: "gpt-6.1-sol",
reasoning: { effort: "minimal" }, // not supported on Astra or 6.1 Sol
temperature: 0.2, // rejected while reasoning is on
prompt_cache_retention: "24h", // pre-GPT-5.6 caching field
input: prompt,
tools,
});
// Right: same intent, GPT-6 parameters.
await client.responses.create({
model: "gpt-6.1-sol",
reasoning: { effort: "low" }, // low | medium (default) | high | xhigh | max
prompt_cache_options: { ttl: "30m" },
input: prompt,
tools, // tool calling needs Responses, not Chat Completions
});The other quiet change is tool calling. Astra and GPT-6.1 Sol support Chat Completions only for requests without tools; anything with function calls needs the Responses API. If your integration still sends tools through Chat Completions, the model upgrade is really an API migration, and that is the part to budget time for. OpenAI also ships an openai-docs skill for Codex that applies these changes, which is worth running on a branch and reviewing line by line rather than trusting.
Astra is more sensitive to instruction files than earlier models. OpenAI's guide says it can pause and block early on unclear or conflicting guidance in a skill or AGENTS.md, and strongly recommends auditing those files. Before you swap the model on an agent that loads many skills, read every instruction file it can see and remove contradictions, or expect approval pauses you did not have on GPT-5.6.
Ultrafast is a service_tier value on a normal Responses request. OpenAI's figure is up to 8x faster token generation in Codex and 6x in the API, and the Ultrafast pricing table lists gpt-6-astra at $60 input and $300 output per million tokens for prompts up to 272K, exactly six times the standard Astra rate. Speed is bought linearly with money.
# Ultrafast is a service tier, not a model. Only gpt-6-astra is broadly
# available; Sol is preview-only at the time of writing.
curl https://api.openai.com/v1/responses \
-H "Authorization: Bearer $OPENAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "gpt-6-astra",
"service_tier": "ultrafast",
"input": "Summarise the failing test output in three lines."
}'
# For agent loops with many tool calls, OpenAI recommends one persistent
# WebSocket (client.responses.connect() in Python) and previous_response_id
# per turn; over plain HTTP the connection overhead eats the speed-up.The honest use case is a human waiting on a long generation: an engineer watching Codex write a large diff, or a support tool where seconds of latency cost a conversation. A nightly batch job is the opposite case, where Batch at half price is the right tier. I would gate Ultrafast behind an explicit flag per call site, so nobody discovers the six-fold bill at the end of the month.
The Agents API is not new at DevDay. The changelog dates its public beta to 10 September: a managed Codex harness where OpenAI runs session orchestration, context compaction and recovery, and your sandbox can be OpenAI-hosted or your own. What DevDay added on 29 September is computer use inside it, so a hosted agent can complete tasks in an OpenAI-hosted browser, with website access approvals and sign-in handled by your application. Model usage, tools and hosted sandboxes bill at their normal rates.
Amazon Bedrock Managed Agents is the same idea on AWS. OpenAI's page describes it as the Agents API adapted for AWS: the harness and model inference run in Amazon Bedrock, commands and tools run on AgentCore Runtime or your own compute, and calls are authenticated with IAM credentials and SigV4 signing instead of an OpenAI project key. For teams whose data and procurement already live in AWS, that removes the main objection to the hosted harness.
OpenAI's own Bedrock page warns that shared concepts do not imply identical API contracts or feature availability. Do not port Agents API examples to Bedrock and assume they work; check the AWS documentation for the current endpoint, supported models, tools, Regions and pricing first.
Three smaller announcements matter depending on what you build. None of them needs action this week, but each changes a build-or-wait decision.
The Decisions API is the one I would watch most closely. Routing and classification are the high-volume, low-margin calls in most production systems, and a purpose-built endpoint could undercut a general model call. Until it has a price and a stable contract, it is a design option to keep in mind, not a dependency.
Dots are the consumer headline: persistent agents powered by Astra, each with connected apps and its own cloud computer, rolling out to Pro, Business and Enterprise in beta, with the European Economic Area, Switzerland and the UK excluded at launch. Specialist dots are an enterprise preview, and a Microsoft Agent 365 integration is planned. The in-depth treatment lives in this site's separate Dots post, so I will not repeat it here.
For developers, the relevant parts are distribution and plans. ChatGPT Space gives teams and agents a shared place for files and project context, and plugins are how your product shows up inside it. On plans, Pro 500 adds 25 times the Plus usage allowance plus Ultrafast access, and Pro 200 subscriptions are open again with different allowances for new subscribers than for grandfathered ones. Those are ChatGPT plans, not API pricing, so they change what your users can do in ChatGPT, not your API bill.
If I had one sprint to spend on DevDay, this is the order I would spend it in. Each step is cheap, and each one protects you from the expensive version of the same mistake later.
The rule I take from this DevDay is the same one every model launch teaches: separate what has a model ID and a price from what has a keynote slide. GPT-6.1 Sol, Ultrafast on Astra and computer use in the Agents API are real today and worth an evaluation sprint. The Decisions API, Private Inference and Ultrafast on Sol are announcements. Measure cost per task before you switch, and budget the migration as an API change, because that is where the time goes.
Sources