Cloud
Cloudflare Containers: 6x Faster Starts for Agent Sandboxes
October 202614 min read

On ComputeSDK's Burst TTI benchmark, which launches 100 sandboxes concurrently and measures time-to-interactive from the client, the median start fell from 4.049 seconds to 648 milliseconds, a 6.2x improvement. The 95th percentile went from 5.839 seconds to 910 milliseconds and the 99th from 6.717 seconds to 1129 milliseconds. Cloudflare's own preliminary burst test also started 100,000 Containers in 5.387 seconds across six locations on a single account.
Yes, under the durable_object scheduling policy, which is in public beta. You declare named images in the images field of the containers entry in wrangler.jsonc, each becomes a property on this.ctx.container.images, and you pass an image and an instance such as standard-1 or standard-2 to this.ctx.container.start(). If you omit the instance, the Container runs as lite.
Cloudflare maintains the Container class and the legacy Sandbox class through December 31, 2026. Existing deployments keep running after that date, but the classes receive no further updates, and new capabilities such as runtime image selection and snapshots are available only through ctx.container. Moving to the faster durable_object scheduling policy also requires a new Container application, because the policy of an existing application cannot be changed.
Call this.ctx.container.snapshotContainer() with a name, store the returned handle in Durable Object storage, and later pass it to start() as containerSnapshot instead of an image. Snapshots are immutable, tied to the image they were created from, capped at 20 GB, and retained for 30 days from creation or the most recent restore. They are in public beta and work only with the durable_object scheduling policy.
Yes. Reported on 4 September 2026 by Oren Yomtov of Accomplish, the flaw let a Workers Paid customer read residual 64 KiB disk blocks left by other tenants' Containers, because a shared dm-thin pool was configured with skip_block_zeroing. Cloudflare removed that option fleet-wide, retired running disks and cleared cached image snapshots by 19 September, and found no evidence of exploitation beyond authorised testing. Customers did not need to take any action.

Key Takeaway
Cloudflare Containers' durable_object scheduling policy, in public beta since 30 September 2026, cut median sandbox start from 4.049 seconds to 648 milliseconds on ComputeSDK's Burst TTI benchmark, lets a Durable Object pick each sandbox's image and instance type at runtime, and adds filesystem snapshots. Adopting it requires a new Container application, not a configuration change.
Cloudflare published two posts about Containers six days apart. On 24 September 2026 it disclosed that a customer on a Workers Paid account could read residual disk blocks left behind by other tenants' Containers on the same host, a flaw it had already fixed. On 30 September it announced that Containers had been rebuilt for agent sandboxes: starts more than six times faster, image and instance type chosen at runtime, and filesystem snapshots in public beta. Read together, they describe one engineering pressure from both ends — reuse hosts, disks and prepared state as aggressively as possible, and keep tenants apart while doing it.
I run my own agent jobs headless on a single VPS, so I read the announcement as an infrastructure engineer deciding where sandboxes should live. I have not deployed the new scheduling policy, and nothing here is a measurement of mine. This post covers what changed, the numbers and what they measure, the API shape with code quoted from Cloudflare's post, the snapshot model, the disk-isolation disclosure, the migration that the 31 December 2026 date forces, and what the announcement leaves unsaid. Every figure traces to Cloudflare's post or its documentation.
Until now a Cloudflare Containers application fixed its image and compute size at deploy time, and every combination of the two was its own application with its own Durable Object namespace. Cloudflare's own example shows the cost: a small Node.js sandbox and a large Python build sandbox meant two applications, two namespaces and routing logic in the Worker. The new durable_object scheduling policy turns both choices into arguments that code passes when a sandbox starts. Four things ship together:
All of it is exposed only through the native ctx.container API inside a Durable Object, not through the Container wrapper class that Containers launched with. Cloudflare says it hid the Durable Object behind that class to make sandboxes feel familiar, and that nearly every team it worked with then needed something the generic lifecycle did not give them. Base44 and Kilo Code are quoted as customers; Cursor Cloud Agents, Devin Outposts, the OpenAI Agents API and Claude Managed Agents are named as integrations. The scheduling policy itself is in public beta.
The headline figures come from ComputeSDK's Burst TTI benchmark, which launches 100 sandboxes concurrently and measures time-to-interactive from the client. Cloudflare calls it independent, and the comparison is the previous scheduling path against the new policy:
| Measurement | Previous scheduling path | durable_object policy | Improvement |
|---|---|---|---|
| Median | 4.049 s | 648 ms | 6.2x faster |
| 95th percentile | 5.839 s | 910 ms | 6.4x faster |
| 99th percentile | 6.717 s | 1129 ms | 5.9x faster |
The median is the number in the headline, but the tail is the one I care about. An agent that fans a task out to ten sandboxes waits for the slowest of them, so the gap between median and p99 is what a user actually feels. On the old path that gap was 2.668 seconds; on the new one it is 0.481 seconds. That compression matters more than the 6x, because it decides whether a fan-out of ten feels like one start or like the worst of ten.
The second figure is different in kind. In what Cloudflare calls a preliminary burst test of its own, a single account started 100,000 Containers in 5.387 seconds across six locations. That is a vendor measurement, not the independent one, and the post gives neither the instance type nor the image used. Treat it as evidence that the scheduler does not fall over under burst load, not as a throughput you can plan against.
The old path put Cloudflare's global control plane in front of every start: resolve the application configuration, find capacity, coordinate placement. Under the new policy the request begins at the Durable Object, and the scheduler looks for capacity on the same machine first, then widens the search within the same location. It favours hosts that already hold the Container's image or snapshot in local storage, and instead of booting a new virtual machine it restores a prepared one that is not yet assigned, reusing networking and filesystem setup.
The system image is the same idea pushed further. Because Cloudflare controls cloudflare/debian-trixie, it distributes and prepares it on eligible hosts before any request arrives, so a start never downloads or unpacks the base image while the user waits. Starting it needs no image declaration in Wrangler at all — this is the call as Cloudflare's post shows it:
// Quoted from Cloudflare's announcement. No named image in Wrangler:
// the Cloudflare-managed image is already prepared on eligible hosts.
this.ctx.container.start({
image: "cloudflare/debian-trixie",
instance: "standard-2",
enableInternet: true,
entrypoint: ["/bin/sleep", "infinity"]
});The design trade is plain once it is laid out: the fast path is a cache-hit path. A host that already has your image, or your snapshot, is what turns four seconds into well under one. The post does not say which image the benchmark used, so the case I would measure before committing is the first start of a large custom image that few hosts have cached yet.
For a new application, opting in is a Wrangler change. You set the scheduling policy and declare the images the Durable Object may choose from; each key becomes a property on this.ctx.container.images. This is the configuration from Cloudflare's post:
// wrangler.jsonc
{
"containers": [
{
"class_name": "AgentSandbox",
"scheduling_policy": "durable_object",
"images": {
"node": { "dockerfile": "./images/node/Dockerfile" },
"python": { "dockerfile": "./images/python/Dockerfile" }
}
}
],
"durable_objects": {
"bindings": [{ "name": "SANDBOX", "class_name": "AgentSandbox" }]
}
}The selection then happens when a task arrives. Cloudflare's example picks the toolchain and the compute size from the task itself — what used to need a separate application and a separate wrangler deploy is now a conditional:
import { DurableObject } from "cloudflare:workers";
export class AgentSandbox extends DurableObject {
async startWorkspace(workspace) {
if (this.ctx.container.running) {
return;
}
const image =
workspace.toolchain === "python"
? this.ctx.container.images.python
: this.ctx.container.images.node;
const instance =
workspace.workload === "build"
? "standard-2"
: "standard-1";
this.ctx.container.start({
image,
instance,
enableInternet: true,
});
}
}Two details in that example deserve attention. First, start() returns before the Container is ready to accept requests, so the documentation tells you to add your own readiness check before sending traffic or calling exec(). Second, the example sets enableInternet: true, while the scheduling-policy documentation's examples set it to false. For an agent sandbox, outbound access is the leg that turns a prompt injection into data exfiltration, so decide it per task rather than copying the example.
The instance names map to fixed sizes. The announcement never states them; Cloudflare's limits page does, and the runtime accepts exactly these five names:
| Instance type | vCPU | Memory | Disk |
|---|---|---|---|
| lite | 1/16 | 256 MiB | 2 GB |
| standard-1 | 1/2 | 4 GiB | 8 GB |
| standard-2 | 1 | 6 GiB | 12 GB |
| standard-3 | 2 | 8 GiB | 16 GB |
| standard-4 | 4 | 12 GiB | 20 GB |
If you omit instance, the Container gets lite. The runtime does not accept basic or the legacy dev and standard aliases. A custom size is allowed as an object of vcpu, memoryMib and diskMb, between 1 and 4 vCPU, up to 12 GiB of memory and 20 GB of disk, with at least 3 GiB of memory per vCPU. So in Cloudflare's example, standard-1 means half a vCPU with 4 GiB, and standard-2 a full vCPU with 6 GiB.
The durable_object policy does not support max_instances. Running instances count only against your account limits, and Cloudflare's migration guide says to enforce any per-application cap in application code. An agent loop that retries by starting fresh sandboxes will keep starting them until the account ceiling stops it — put the cap in the Durable Object before the first production deploy, not after the first bill.
Cloning a repository and installing dependencies can take much longer than starting the Container, and an agent's output is worth keeping between sessions. Filesystem snapshots, in public beta, address both: capture the filesystem, store the handle in the Durable Object's own storage, and pass it back to start() later. Cloudflare's example:
async saveWorkspace() {
const snapshot = await this.ctx.container.snapshotContainer({
name: "project-ready",
});
await this.ctx.storage.put("workspace-snapshot", snapshot);
}
async restoreWorkspace() {
const snapshot = await this.ctx.storage.get("workspace-snapshot");
if (!snapshot) {
throw new Error("No workspace snapshot found");
}
this.ctx.container.start({
containerSnapshot: snapshot,
instance: "standard-2",
enableInternet: true,
});
}Cloudflare describes two patterns. A single workspace can continue across sessions — save when the user stops, and restore the repository, dependencies, build caches and edits the next day. And because snapshots are immutable and reusable, one snapshot can be the shared starting point for many sandboxes, which is what agent evaluations need: the same repository, tools and inputs for every run, so the only variable is the prompt, the model or the agent version.
Cloudflare calls them filesystem snapshots, and the name is the scope: the files come back, so a development server or a warm interpreter has to be started again after a restore. That start cost belongs in any estimate of how long a resume takes.
Take a shared snapshot before any credential touches the disk. A snapshot captures the whole filesystem, so a token written to a .env file or a git credential store rides into every sandbox started from it. Cloudflare's design gives you a better place for secrets anyway: the Durable Object can inject credentials through the Container's outbound request handler, so they never have to live on the sandbox's disk.
On 4 September 2026, Oren Yomtov of Accomplish reported through Cloudflare's HackerOne programme that a customer with a Workers Paid account could recover residual disk blocks previously used by other customers' Containers on the same host. Cloudflare published the details on 24 September. Each Container runs inside its own Firecracker virtual machine, but its writable root disk came from a Linux device-mapper thin pool (dm-thin) shared across customer accounts and configured with skip_block_zeroing.
The mechanism is small and instructive. The pool used 64 KiB thin blocks. A 4 KiB write into an unmapped region made dm-thin allocate a whole reused 64 KiB block, and with zeroing disabled the other 60 KiB kept whatever the previous owner had written, readable through a raw read of the device. The researchers observed residual material on 18 of 24 placements and 20 of 22 underlying nodes across four continents, including directory structures, database pages and structurally complete SQLite databases. They could not choose a victim or touch an attached disk.
The response was fast and, to my reading, thorough. Cloudflare merged a runtime fix the same day, finished removing skip_block_zeroing across the fleet on 7 September, and then went further than the patch strictly required: zeroing new allocations did not clean blocks already mapped into running disks or into each host's cache of prepared dm-thin snapshots for image layers, so it retired those disks and cleared those caches, finishing on 19 September. Its review of retained disk-I/O telemetry found no activity matching the technique beyond the researchers and its own engineers, and customers had nothing to do.
I raise it here not as a mark against the new release but because it locates the risk precisely. The flaw lived in shared storage pools and host-local caches of prepared disk state — the same layer the new start path leans on harder, through cached images, cached snapshots and prepared virtual machines. The announcement does not describe how prepared VMs or snapshot storage are sanitised between tenants. Given how this disclosure was handled, I expect a good answer exists; it is still the first question I would put to an account team before storing customer workspaces in snapshots.
Cloudflare will maintain the Container class and the legacy Sandbox class through 31 December 2026. Deployments keep running after that, but the classes get no further updates, and every new capability is native-only. The trap is that migrating means two separate changes, and only the second unlocks the faster start and snapshots:
// Wrong: flipping the policy on the live entry. The policy is immutable,
// and the deploy fails only after the new Worker version is already serving.
{
"class_name": "Sandbox",
"scheduling_policy": "durable_object", // was "default"
"image": "./container/Dockerfile"
}
// Right: keep the old application untouched and add a replacement with a
// NEW Durable Object class, so traffic can be routed back. (Shape follows
// Cloudflare's migration guide.)
{
"containers": [
// Existing application. Keep this entry unchanged.
{
"class_name": "Sandbox",
"image": "./container/Dockerfile",
"instance_type": "standard-2",
"max_instances": 10,
},
// Replacement application.
{
"name": "sandbox-durable-object",
"class_name": "DurableSandbox",
"scheduling_policy": "durable_object",
"images": {
"base": {
"dockerfile": "./container/Dockerfile",
},
},
},
],
"durable_objects": {
"bindings": [
{ "name": "SANDBOX", "class_name": "Sandbox" },
{ "name": "DURABLE_SANDBOX", "class_name": "DurableSandbox" },
],
},
"exports": {
"Sandbox": { "type": "durable-object", "storage": "sqlite" },
// The new class must be SQLite-backed. No storage moves across with it.
"DurableSandbox": { "type": "durable-object", "storage": "sqlite" },
},
}Do not flip scheduling_policy on the existing Container entry. Cloudflare's migration guide warns that wrangler deploy ships the new Worker version before it configures the Container application, so the deploy fails with the new code already live and no durable_object application behind it. Add a new entry with a new class, and leave the old one untouched so traffic can be routed back.
Two packaging changes sit alongside this. Sandbox SDK 1.0 becomes a set of utilities that work inside your own Durable Object class rather than a base class you extend, and @cloudflare/computer is offered as a higher-level environment that combines Dynamic Workers and Containers with a synchronised filesystem.
An announcement is written to answer the questions its authors expect. These are the ones it leaves open, and each is worth an answer before a production commitment:
None of these is a reason to dismiss the release. They are the difference between a promising beta and something I would put customer code in, and every one of them can be answered with a short test or a direct question.
The rule I take from this release: on Cloudflare, an agent sandbox is now a Durable Object that decides its own Container, and everything worth having — the faster start, runtime images, snapshots — lives only on that path. If you run the Container or Sandbox class today, plan the move as a new application before 31 December 2026, cap instances in code, keep credentials off the snapshot, and measure a cold start of your own image before trusting someone else's median.
Related reading on this site:
Sources and further reading