Skip to content

SYSTEM Cited by 2 sources

Linux Device Mapper (DM)

The Linux Device Mapper is the kernel block-layer proxy mechanism that lets targets interpose on struct bio — the main unit of I/O for the Linux block layer — and either dispatch it, drop it, or mutate it and ask the kernel to resubmit. It's the backend for userland LVM2 and the substrate on which a long list of targets are built:

  • dm-linear — carve one big device into smaller ones.
  • dm-stripe — combine smaller devices into one striped device.
  • dm-raid1 — software RAID mirroring.
  • dm-snap — snapshots of arbitrary devices.
  • dm-verity — cryptographic verification of boot devices.
  • dm-clone — block-level async clone.
  • dm-crypt — full-disk encryption.
  • dm-thin — thin provisioning: lazily allocate physical blocks only on first write, from a shared pool (see below).

Why it looks like a network protocol

The 2024-07-30 Fly.io Making Machines Move post frames the block layer as "organized as if your computer was a network running a protocol that basically looks just like that" — a struct bio carries an opcode (read, write, flush, discard, secure erase, write-same, write-zeroes), a device, and a page/len/offset vector. DM targets plug into this bio stream the way a proxy plugs into a network protocol. "No nerd has ever looked at a fixed-format message like this without thinking about writing a proxy for it, and struct bio is no exception. The proxy system in the Linux kernel for struct bio is called device mapper, or DM."

Composition

DM targets can stack on top of each other. Fly's fleet stacks DM like this on the target worker during a migration:

  1. Source Volume mounted over the network via iSCSI appears as a block device.
  2. dm-crypt wraps it to produce the plaintext view.
  3. dm-clone takes that plaintext source device + a fresh local target device + a metadata device, and presents the cloned plaintext volume.
  4. The new Fly Machine mounts the cloned plaintext volume.

dm-thin thin provisioning and skip_block_zeroing

dm-thin is the device-mapper target for thin provisioning: a thin volume presents a large virtual address space but physical storage is allocated lazily — a physical block from a backing pool is mapped only when the thin device first writes to a previously-unmapped region. Reading an unmapped region does not allocate; dm-thin simply returns zeroes. Multiple thin volumes can be carved from one shared pool, and freeing a thin volume returns its physical blocks to that pool for reuse by other volumes. This is the block- layer machinery behind fast copy-on-write container/VM root disks (paired with CoW snapshots of image layers).

The pool has a configurable thin-block size and a skip_block_zeroing option. By default dm-thin zeroes a newly-allocated block before exposing it to a device, so a block can never carry a previous owner's data across a reuse. With skip_block_zeroing set, that clearing step is skipped (a throughput optimization) — and a block reassigned from the shared pool retains whatever the previous owner wrote, except for the portion the new owner overwrites. If the new owner's write is smaller than the block size, the un-written remainder of the block is readable residual data from the prior tenant. This is the exact data-remanence failure mode Cloudflare hit in 2026-09 (see Seen in): a 64 KiB thin-block pool with skip_block_zeroing set, where a 4 KiB write left 60 KiB of a reused block exposed. The fix was to remove skip_block_zeroing, restoring the zero-on-allocate default — but that only protects future allocations, not blocks already mapped into live devices or cached snapshots.

Seen in

  • sources/2026-09-24-cloudflare-how-cloudflare-addressed-a-cross-tenant-data-exposure-vulner-4224c10f — canonical wiki instance of a dm-thin skip_block_zeroing data-remanence vulnerability. Cloudflare Containers back each container's writable root disk with a dm-thin pool (64 KiB thin-block size) shared across customer accounts; the pool had skip_block_zeroing set, so freed blocks were handed to new containers un-cleared. A sub-block (4 KiB) write triggered allocation of a reused 64 KiB block and left the remaining 60 KiB readable as residual data from another tenant via a raw read of Firecracker's /dev/vdc. Note the precise semantics the exploit relied on: an unmapped region returns zeroes (no leak), so the attacker first had to trigger allocation with a small write. Cloudflare removed skip_block_zeroing fleet-wide (restoring zero-on-allocate) and additionally rebuilt already-mapped disks + purged the per-host cache of prepared dm-thin snapshots for OCI image layers, since zero-on-allocate does not retroactively sanitize existing mappings. See systems/cloudflare-containers and concepts/tenant-isolation.

  • sources/2024-07-30-flyio-making-machines-move — Load-bearing description of DM as the kernel's block-layer proxy system; anchors every DM target the post relies on.

Last updated · 766 distilled / 2,225 read