Skip to content

SYSTEM Cited by 1 source

Cache Transcoding

Cache Transcoding is a Cloudflare prototype (built during a 1.1.1.1-Intern-Program internship) that expands effective CDN cache capacity by compressing eligible cached responses with Zstandard inside the Pingora-based proxy. When an eligible response enters the cache it is encoded with zstd before being written to disk; it stays compressed while it lives in cache and while it moves between data centers over Tiered Cache; it is decoded back to its original identity representation only on the final client-facing hop. It is a prototype, not a shipped product. (Source: sources/2026-09-01-cloudflare-how-we-could-save-petabytes-of-cache-storage-with-zstandard)

What problem it solves

RAM and HDD prices rose sharply over the year preceding the post. Instead of buying more hardware, Cache Transcoding changes how an asset is represented on disk so that existing hardware stores more objects and the backbone carries fewer bytes between POPs. Cloudflare traditionally stores an asset in whatever content encoding the origin supplied (an uncompressed origin response is stored and moved uncompressed); Cache Transcoding adds a compression layer inside the cache, independent of what the origin sent.

Design

  • Encode once, decode many. Encoding is paid a single time when an asset fills the cache; decoding is paid on every serve. Encode is ~2.8× costlier per byte than decode, but runs far less often — see encode-once-decode-many. This asymmetry is the core rationale.
  • Eligibility gate. Only 200 OK responses with Content-Encoding unset, a compressible-text Content-Type, and a known Content-Length ≥ 4 KiB are transcoded (eligibility-gate-before-compression). Media (images/video/fonts), already-compressed responses, slice subrequests, range requests, unknown-length bodies, and binary content pass through unchanged.
  • zstd level 3. Chosen as the conservative speed/ratio default so cache fills don't become a CPU bottleneck. Both the level and the 4 KiB threshold are tunable parameters, not permanent limits.
  • Compressed across tiers, decoded at the edge. The zstd form is preserved on disk and on the wire between cache tiers; decode happens only on the client-facing hop (store-compressed-transfer-compressed-decode-at-edge).
  • Encoding marker prevents double work. Cache metadata records that the stored form is zstd and preserves the original content length; a tier receiving an already-zstd object sees the marker and does not re-encode (encoding-marker-prevents-double-work).

Request paths

Scenario Behavior
Cache miss (single hop) Proxy encodes body with zstd → writes to disk → decodes to identity before serving
Cache hit Read zstd object from disk → decode on the client-facing hop → serve
Full miss with Tiered Cache Upper tier fetches identity from origin → encodes once → stores zstd → transfers compressed to lower tier → lower tier stores zstd, decodes for request
Lower-tier miss, upper-tier hit Origin not involved; compressed object moves directly between tiers, stays compressed on wire+disk, decoded once at lower tier
Lower-tier hit No transfer/encode; read zstd, decode, serve

Measured behavior (prototype)

Measure Value
Compression ratio (test corpus) 2.834× (≈⅓ on-disk size)
Encode cost 4.31 ns/byte ≈ 232 MB/s, once per fill
Decode cost 1.56 ns/byte ≈ 641 MB/s, per serve
Eligible text arriving uncompressed ~71%
4 KiB threshold excludes only ~1% of otherwise-eligible bytes
Extra CPU (modeled, level 3) "a few percent" under tested traffic/reuse

Validated over >1 million requests across 10 cache servers (half with Tiered Cache enabled, half disabled), with request logs + Prometheus + Jaeger traces confirming where encode/decode happened. The test corpus (assets ~195 KiB and ~272 KiB, both ~2.8× compressible) was deliberately compressible — a clean validation signal, not a fleet-wide constant.

Why not just transcode hot content

Cloudflare considered restricting transcoding to popular/hot assets (more reuse per fill) but found it counterproductive: decoding runs on every serve regardless of popularity, so limiting to hot content cut the storage saving without cutting CPU proportionally. The simpler blanket policy — transcode all eligible compressible text ≥ 4 KiB — captured nearly all of the storage benefit within the CPU budget.

Status and future work

Prototype only. Named future directions: higher zstd levels, broader content types and object sizes, tuning the eligibility parameters, and handling range requests, pre-compressed origin responses, and passing the compressed object directly to downstream components that already support it (skipping decode).

Seen in

Last updated · 766 distilled / 2,225 read