Skip to content

CLOUDFLARE 2026-09-24 Tier 1

Read original ↗

How Cloudflare addressed a cross-tenant data exposure vulnerability in Containers

Summary

On 2026-09-04, security researcher Oren Yomtov (of Accomplish) reported, through Cloudflare's HackerOne bug-bounty program, a cross-tenant data-exposure vulnerability in Cloudflare Containers (and the Sandboxes built on them). A customer on a Workers Paid account could recover residual disk blocks previously used by other customers' Containers on the same host — a classic data-remanence leak arising from a storage-allocation optimization. Cloudflare presents the root cause (Linux device-mapper thin provisioning with block zeroing disabled), the exploit mechanics, the researcher's validation methodology, the two-stage mitigation, and a telemetry-based investigation that found no evidence of malicious exploitation. The whole fix landed the same day it was reported; cache-snapshot cleanup completed 2026-09-19. This is a textbook coordinated-disclosure vendor-first-patch instance on a substrate the vendor fully controls.

Key takeaways

  1. The leak was data remanence through storage-block reuse, not a logic/auth bug. Cloudflare Containers give each container a writable root disk backed by dm-thin thin provisioning with a 64 KiB thin-block size, drawn from a pool shared across multiple customer accounts. When a container's thin volume is deleted, its physical blocks return to the shared pool. The pool was configured with skip_block_zeroing, so dm-thin did not clear a block before handing it to a new owner. (Source: this article)

  2. A sub-block write is what exposed the remainder. For a full 64 KiB write, the previous contents are entirely replaced. But a smaller write (e.g. 4 KiB) into a freshly-allocated 64 KiB block replaced only the written 4 KiB — the remaining 60 KiB retained the previous owner's data and became readable via a raw device read of /dev/vdc. Reading an unmapped region revealed nothing: dm-thin returns zeroes for unmapped regions without allocating a physical block, so the exploit had to trigger allocation with a small write first. (Source: this article)

  3. The proof-of-concept, step by step. (1) Create a container on a Workers Paid account. (2) Open the writable root disk /dev/vdc. (3) Read + record a baseline. (4) Write one 4 KiB block into each 64 KiB-aligned region that corresponds to ext4 free space in the guest filesystem. (5) Re-read those blocks. (6) Examine only the 60 KiB not overwritten by the new container. (Source: this article)

  4. Validation used ext4 directory-block checksums to attribute blocks to foreign filesystems. With ext4's metadata_csum feature, directory-block checksums incorporate filesystem- and inode-specific values, letting the researchers distinguish their own test blocks from foreign ones. Across six production placements they reported all 5,614 testable directory blocks were foreign (zero attributable to their own filesystem) and identified 2,700 distinct foreign directory inodes. A control test on 162 deliberately created-then-deleted blocks in their own filesystem attributed all 162 correctly. Overall they observed residual material on 18 of 24 placements and 20 of 22 underlying nodes across four continents, including directory structures, database pages, and structurally complete SQLite databases. (Source: this article)

  5. The exploit could not target a victim — it crossed the boundary but not at will. Cloudflare assigns Containers to eligible hosts automatically; customers can't select the host. Exposure depended on placement luck and which released blocks dm-thin happened to reassign — the attacker could not pick a customer, workload, host, or data, and residual data wasn't guaranteed present. A successful exploitation crossed the tenant-isolation boundary (disclosing filesystem metadata, directory structures, database pages, application data) but did not allow modifying another tenant's active data or affecting availability — a read-only, non-targetable blast-radius. (Source: this article)

  6. The fix was two stages, because zeroing new allocations doesn't sanitize already-mapped blocks. Stage 1: remove skip_block_zeroing fleet-wide, restoring dm-thin's default of clearing newly allocated blocks before exposing them — this stopped the reported technique (the researchers confirmed the PoC no longer worked). But blocks already mapped into running container disks and into each host's cache of prepared dm-thin snapshots for OCI image layers were untouched — a new container could inherit a cached-layer mapping without re-allocating, leaving residual bytes readable. Stage 2: retire all running container disks and remove pre-mitigation cached image snapshots — drain hosts off-peak, restart VMs, clear each host's image cache so disks + cached layers are recreated with zeroed allocations. (Source: this article)

  7. Telemetry-based exploitation hunt found nothing beyond the researchers. The PoC produces a characteristic write/read signature (a small 4 KiB write triggering allocation, followed by reads recovering substantially more than was written). Cloudflare built detection signatures from that signature and applied them to retained historical disk-I/O telemetry, attributing all matching activity to the researchers and Cloudflare's own authorized validation — no evidence of any other exploitation. (Source: this article)

  8. Same-day response, fully coordinated. Reported 2026-09-04 15:26 UTC → incident opened 18:45 → runtime fix + reuse test merged 21:27 → new/live pool changes merged 22:03 → rollout begun 23:15 → rollout complete + old-pool clearing begun 2026-09-07 06:13 → researchers confirm PoC dead 2026-09-14 10:50 → bounty awarded 2026-09-14 12:52 → all pre-mitigation cached snapshots cleared fleet-wide 2026-09-19 15:03. Post co-authored with the Accomplish research team; recovered data confirmed securely deleted per HackerOne policy. (Source: this article)

Systems

  • Cloudflare Containers — the affected multi-tenant Docker-container runtime; each container runs inside a dedicated Firecracker VM with a dm-thin-backed writable root disk drawn from a host-level pool shared across accounts. Sandboxes are built on Containers, so both were affected.
  • Firecracker — the micro-VM monitor hosting each container; it presents the dm-thin root disk to the guest as /dev/vdc, the raw device the PoC read residual bytes from.
  • Linux device-mapper (dm-thin) — the thin-provisioning target providing per-container writable root disks; skip_block_zeroing on the shared pool is the root-cause configuration.
  • ext4 (recorded as prose/tag, not a page) — the guest filesystem; its metadata_csum directory-block checksums were the researchers' attribution primitive, and ext4 free-space regions were where the PoC wrote its probe blocks.

Concepts

  • Tenant isolation — the boundary that was crossed; this is a storage-layer breach (freed-block reuse), distinct from the identity/authorization/network layers that most tenant-isolation shapes on the wiki defend, and adjacent to the microarchitectural boundary-break on Cloudflare Workers (Spectre).
  • Coordinated disclosure — a bug-bounty-initiated, vendor-first-patch instance: because Cloudflare owns the entire deployment substrate, the fix, fleet rollout, and cache cleanup all happened before public disclosure, with the reporter co-authoring the writeup.
  • Blast radius — the exposure was bounded to read-only, non-targetable recovery of residual blocks; no write access, no victim selection, no availability impact.
  • Data remanence (recorded as prose/tag, not a page per the taxonomy gate) — the textbook mechanism: data persisting in storage after logical deletion, exposed when freed media is reallocated without sanitization (cf. NIST SP 800-88 media-sanitization framing).
  • Side-channel (adjacent, referenced) — the leak is direct residual-data recovery, not a timing/microarchitectural side channel, but it sits in the same isolation-boundary-break family as the Workers Spectre work.

Operational numbers

  • Thin-block size: 64 KiB; probe write size: 4 KiB (so 60 KiB residual per triggered block).
  • Validation: 5,614 testable directory blocks across 6 placements, 0 attributable to the researchers, 2,700 distinct foreign directory inodes; control test 162/162 correctly attributed.
  • Reach: residual material on 18/24 placements, 20/22 nodes, across four continents.
  • Timeline: report → runtime fix merged in ~6 hours; rollout complete 2026-09-07; all cached snapshots cleared 2026-09-19.
  • Query transport of the substrate is not disclosed; no evidence of third-party exploitation in historical disk-I/O telemetry.

Caveats

  • No customer data confirmed compromised. Cloudflare states it found no evidence of malicious exploitation and that no customer action is required.
  • Non-targetability was a mitigating factor, not a control. The inability to select a victim reduced practical risk but was incidental to placement, not a designed defense — the designed fix is block zeroing + snapshot cleanup.
  • The materials submitted contained no third-party content or identifiers. The researchers' scripts emitted only aggregate counts, block offsets/sizes, checksum results, and truncated hash prefixes; recovered data was securely deleted post-submission.
  • The cached-image-snapshot subtlety is the load-bearing lesson. Fixing the allocation path (zero-on-allocate) does not retroactively sanitize blocks already mapped into live disks or cached OCI-layer snapshots — a full remediation required draining hosts and rebuilding disks + caches.

Source

Last updated · 766 distilled / 2,225 read