Cloudflare Containers, rebuilt to scale agent sandboxes¶
Summary¶
Cloudflare rearchitected Containers from an
application-deployment model into an on-demand agent-sandbox substrate. The
central change is a new durable_object scheduling policy that moves the
two decisions that matter most for an agent sandbox — which image it runs and
how much compute it gets — out of deploy-time configuration and into
application code that runs at request time, controlled by the Container's
attached Durable Object. On top of that,
a redesigned runtime restores a prepared, not-yet-assigned virtual machine
instead of booting one from scratch, prefers hosts that already hold the image
or snapshot locally, and starts the Container wherever the Durable Object is
already running. On ComputeSDK's independent Burst TTI benchmark, median
time-to-interactive fell from 4.049 s to 648 ms (6.2× faster); a single
account started 100,000 Containers in 5.387 s across six locations. Two more
capabilities ship alongside: a Cloudflare-managed prepared system image
(cloudflare/debian-trixie) that is pre-distributed to hosts, and native
filesystem snapshots (public beta) that let a workspace be saved and
restored — including immutable snapshots forked into many isolated environments
for evals / reinforcement learning. All new capabilities are native-only on
ctx.container; the legacy Container / Sandbox base classes are maintained
only through 2026-12-31.
Key takeaways¶
- The application was the wrong unit of configuration for agents. Until
now, each
(image, instance-type)combination was its own Containers application with its own Durable Object namespace, deployed ahead of time withwrangler deploy; routing logic in a Worker sent each task to the right one. Agent sandboxes are created on demand, per task, and the task determines image/resources/tools/starting-filesystem — so those decisions "need to live with the application code handling the task." (Source: sources/2026-09-30-cloudflare-containers-rebuilt-to-scale-agent-sandboxes) durable_objectscheduling policy = infrastructure as request-time code. Images are declared inwrangler.jsonc("images": { "node": {...}, "python": {...} }), exposed asthis.ctx.container.images.<name>, and the DO picks image + instance ("standard-1"vs"standard-2") insidectx.container.start({ image, instance, enableInternet })after the task is known. "What used to take a separate application and a separatewrangler deployis now anifstatement."- Rollouts collapse into ordinary code. With no platform-driven grace
periods / percentage splits / config pushes, a Container keeps its start-time
image until the DO's code stops it; the next start chooses whatever image the
code selects. Canary (hash the DO ID), pin (store a
pinned-imagein DO storage), migrate at a checkpoint, and roll back all become a few lines in the DO — e.g.(await ctx.storage.get("pinned-image")) ?? (isCanary(ctx.id) ? images.nodeV2 : images.node). - Startup got 6× faster by moving demand to the Durable Object and skipping boot work. The old path routed every first command through a global control plane to resolve app config, find capacity, and coordinate placement. The new path starts demand at the DO, searches for capacity on the same machine first (then widens within the location), and favors hosts that already have the image/snapshot in local storage. The runtime restores a prepared VM that isn't yet assigned rather than booting a fresh one, reuses networking
- filesystem setup, batches repeated operations, and skips services the first command doesn't need.
- Measured numbers (ComputeSDK Burst TTI, 100 concurrent sandboxes, client- measured time-to-interactive): median 4.049 s → 648 ms (6.2×); p95 5.839 s → 910 ms (6.4×); p99 6.717 s → 1129 ms (5.9×). Burst: 100,000 Containers started in 5.387 s across six locations from one account.
cloudflare/debian-trixieis a pre-warmed base image. A ready-to-use system image (Debian Trixie Slim + Node.js 24.20.0 LTS) that agents can start without authoring a Dockerfile, building, or pushing. Because Cloudflare controls it, it is distributed and unpacked onto eligible hosts before requests arrive, removing image download/unpack from the user-visible path. Agents thenexec()to clone repos, install packages, and configure at runtime.- Native filesystem snapshots (public beta) enable persistence + fan-out.
ctx.container.snapshotContainer({ name })saves a workspace; a later start with{ containerSnapshot, instance, enableInternet }restores it. Snapshots enable (a) one workspace continued across many sessions (repo, deps, build caches, config, edits survive without rebuild), and (b) a shared immutable checkpoint forked into many independent sandboxes — the eval / RL use case, where fixing everything but the variable under test prevents environment drift from contaminating results. - The Durable Object is now the explicit Container controller. The three
improvements (faster scheduling, runtime image selection, snapshots) all come
from "leaning further into the Durable Object" as the stateful controller
attached to each Container.
exec()runs directly in the Workers runtime; outbound request interception, image/instance selection, and snapshots are all onctx.containerwith no wrapper class. This makes the Container a compute extension of the DO, which retains identity, state, policy, and lifecycle. - Three named DO+Container patterns. (a) Agent available while workspace sleeps — the agent loop runs in the DO (session state, WebSockets, model calls), waking the Container only for a shell/compiler/dev-server and paying nothing for idle Linux compute (scale-to-zero). (b) DO programs the security boundary — it remembers authorized services/repos/operations and updates the Container's Outbound Request Handler to inject credentials / enforce policy, described as the Cloudflare OS Gatekeeper pattern applied per agent computer. (c) Evals/RL supervised from outside — a coordinator snapshots a base workspace, forks N attempts (each its own DO + Container), grades them, snapshots the best, and forks again.
- "Decouple the brain from the hands." The post explicitly cites Anthropic's managed-agents framing: run the agent in the DO and use the Container as its workspace, or run the agent in the Container and use the DO to supervise. Because DO and Container are separate, the agent stays available while its sandboxes/tools start, stop, fail, or get replaced independently (patterns/specialized-agent-decomposition).
- Migration + deprecation. New capabilities (
durable_objectpolicy, faster startup, runtime image/instance selection, filesystem snapshots) are native-only viactx.container. TheContainerclass and legacySandboxclass are maintained only through 2026-12-31 (existing deployments keep running but get no updates). Sandbox SDK 1.0 is reframed as a set of utilities, not a base class — its helpers work inside your own Durable Object class alongsidectx.container. Migration is mostly changingextends Containertoextends DurableObjectand callingthis.ctx.containerdirectly.@cloudflare/computeris offered as a higher-level environment combining Dynamic Workers + Containers with a synchronized filesystem.
Operational numbers¶
| Metric | Previous path | New durable_object policy |
Improvement |
|---|---|---|---|
| Startup — median | 4.049 s | 648 ms | 6.2× |
| Startup — p95 | 5.839 s | 910 ms | 6.4× |
| Startup — p99 | 6.717 s | 1129 ms | 5.9× |
| Burst create | — | 100,000 Containers in 5.387 s (6 locations) | — |
- Base image:
cloudflare/debian-trixie= Debian Trixie Slim + Node.js 24.20.0 LTS. - Instance sizes referenced:
standard-1,standard-2. - Named production users: Base44 (app-building workspaces), Kilo Code (cloud-agent sessions). Integrations cited: Cursor Cloud Agents, Devin Outposts, OpenAI Agents API, Claude Managed Agents.
Systems¶
- systems/cloudflare-containers — the rearchitected substrate;
durable_objectscheduling policy, runtime image/instance selection, snapshots. - systems/cloudflare-durable-objects — now the explicit Container controller
via
ctx.container; owns identity, state, rollout policy, lifecycle. - systems/cloudflare-sandbox-sdk — reframed at 1.0 as utilities, not a base
class; legacy
Sandboxclass maintained only through 2026-12-31. - systems/dynamic-workers — the isolate tier complementing Containers; combined
by
@cloudflare/computer. - systems/cloudflare-debian-trixie — new Cloudflare-managed prepared base image.
- systems/cloudflare-computer — higher-level env combining Dynamic Workers + Containers with a synchronized filesystem.
- systems/cloudflare-workers — the runtime
exec()now runs in directly. - systems/firecracker — micro-VM monitor; the "restore a prepared VM" path.
- systems/nodejs — bundled in the
debian-trixieimage (24.20.0 LTS).
Concepts¶
- concepts/cold-start — the core problem: every second of startup is user wait time; the redesign attacks it via prepared-VM restore + image locality + DO-local scheduling.
- concepts/copy-on-write-storage-fork — immutable, reusable snapshots forked into many isolated sandboxes from one baseline.
- concepts/scale-to-zero — the DO stays available and stops Container compute while work pauses, "paying nothing for idle Linux compute."
- concepts/control-plane-data-plane-separation — the redesign explicitly removes the global control plane from the first-command path.
- concepts/micro-vm-isolation — each Container is a Firecracker micro-VM tier.
Patterns¶
- patterns/warm-pool-zero-create-path — "restore a prepared virtual machine that isn't yet assigned" is the canonical zero-work create path.
- patterns/snapshot-replay-agent-evaluation — one immutable snapshot forked into N eval/RL attempts, holding everything fixed but the variable under test.
- patterns/specialized-agent-decomposition — "decouple the brain from the hands": agent in DO, workspace in Container (or vice versa).
- patterns/central-proxy-choke-point — DO programs the Outbound Request Handler as the per-agent egress security boundary (Gatekeeper pattern).
- patterns/progressive-configuration-rollout — canary/pin/migrate/rollback, now expressed as ordinary DO code instead of a platform config push.
Caveats¶
- Startup numbers are (a) an independent ComputeSDK benchmark for the TTI percentiles and (b) Cloudflare's own preliminary burst test for the 100k-in-5.4s figure. Different harnesses; treat the burst number as vendor- reported.
- Filesystem snapshots are public beta at time of writing.
- The legacy
Container/Sandboxclasses stop receiving updates after 2026-12-31; new features are native-only, so this is a soft-forced migration.
Source¶
- Original: https://blog.cloudflare.com/faster-agent-sandboxes/
- Raw markdown:
raw/cloudflare/2026-09-30-cloudflare-containers-rebuilt-to-scale-agent-sandboxes-bde6015f.md