Cloudflare — How we could save petabytes of cache storage with Zstandard and Pingora¶
Summary¶
Cloudflare prototyped Cache Transcoding, a scheme built inside its Pingora-based proxy that encodes eligible cached responses with Zstandard (zstd) before writing them to disk, keeps them compressed while they live in cache and travel between data centers over Tiered Cache, and decodes them back to their original identity representation only on the final client-facing hop. In initial testing, eligible assets shrank to ⅓ of their on-disk size on average (≈2.8× compression ratio on the test corpus), trading a small, one-time CPU cost paid once per cache fill for petabytes of effective cache capacity and reduced cross-data-center backbone traffic. The encode cost is paid once; the storage and bandwidth savings recur every time the asset is reused. Built during an intern's stint on the 1.1.1.1 Intern Program; a prototype, not yet a shipped product. (Source: sources/2026-09-01-cloudflare-how-we-could-save-petabytes-of-cache-storage-with-zstandard)
Problem framing¶
RAM and HDD prices have "exploded" over the past year. Cloudflare runs several massively distributed storage products (the CDN cache among them) whose value depends on packing as much customer content as possible into deployed memory and disk. Rather than buy more hardware, the prototype changes how an asset is represented on disk — compressing eligible text at rest — to store more objects on the same hardware and move fewer bytes between POPs.
Cloudflare traditionally stores an asset using the content encoding the origin supplied: if the origin sends an uncompressed response, Cloudflare stores those uncompressed bytes and transfers them between data centers in the same form. Cache Transcoding adds compression inside the cache itself, independent of what the origin sent.
What is worth compressing (eligibility)¶
Transcoding does not mean compressing everything — that would burn CPU for no gain on already-compressed media. The prototype's traffic-sample split:
| Content class | Share of requests | Share of bytes | Transcode? |
|---|---|---|---|
| Images / video / fonts (already compressed) | 21.4% | 63.3% | No — re-compression wastes CPU |
| Compressible text (HTML, JSON, CSS, JS) | 67.3% | 22.3% | Yes, when eligible |
Within the text slice, ~71% arrived uncompressed (Content-Encoding unset)
and compresses well. The prototype transcodes a response only when all of
these hold (eligibility gate):
- Status is
200 OK. Content-Encodingis unset (not already compressed).Content-Typeis compressible text.- Known
Content-Lengthof at least 4 KiB.
Excluded and passed through unchanged: slice subrequests, responses already using active upstream compression, range requests, precompressed responses, unknown-length bodies, and binary content.
The 4 KiB threshold removed a large number of tiny requests while leaving out only ~1% of the otherwise-eligible bytes; lowering it would add per-object overhead for little extra storage saving. Both the threshold and the zstd level are tunable parameters, not permanent limits.
Why Zstandard, and why level 3¶
zstd is a lossless compressor (every decoded byte identical to the original — "we can change how an asset is represented on disk without changing the asset itself"). Cloudflare cites its own earlier browser- compression testing: zstd compressed 42% faster than Brotli at nearly the same size, and produced files 11.3% smaller than gzip at comparable speed. Because Cache Transcoding would touch a large fraction of traffic, both encode and decode must stay fast — so the prototype uses zstd level 3 (the common speed/ratio default), capturing most of the compression benefit without turning cache fills into a CPU bottleneck.
The asymmetric cost that makes it work¶
| Measure | Value |
|---|---|
| Compression ratio (test corpus) | 2.834× |
| Encode cost | 4.31 ns/byte ≈ 232 MB/s, paid once per fill |
| Decode cost | 1.56 ns/byte ≈ 641 MB/s, paid on every serve |
Encoding is ~2.8× more expensive per byte than decoding, but assets are served far more often than they are filled — so the expensive half runs rarely and the cheap half runs on the hot path. This is the encode-once, decode-many economics that underwrites the whole trade (compression codec trade-off applied at CDN-cache altitude rather than the usual streaming-producer altitude).
Cloudflare initially considered limiting transcoding to popular/hot content (more reuse per fill), but it didn't help: decoding happens on every serve regardless of popularity, so restricting to hot content cut the storage saving without cutting CPU proportionally. The simpler policy — transcode all eligible compressible text ≥ 4 KiB — captured nearly all of the storage benefit within the CPU budget. A worked instance of preferring the simpler blanket policy over a popularity-tiered one.
How it works (request paths)¶
On a cache miss, the Pingora proxy encodes the body with zstd before writing it to disk. Cache metadata records that the stored representation is compressed and preserves the original content length. Before the response leaves the proxy, the body is decoded back to identity for the client.
On a cache hit, the stored zstd object is read and decoded; decoding happens only on the client-facing hop.
Interaction with Tiered Cache — the compressed form is preserved on the wire between tiers (store compressed, transfer compressed, decode at the edge):
- Full cache miss: upper tier fetches identity bytes from origin → encodes once → stores as zstd → transfers the compressed form to the lower tier → lower tier also stores zstd, then decodes for the request path.
- Lower-tier miss but upper-tier hit: origin not involved; the compressed object moves directly between tiers, stays compressed on wire and disk, decoded once at the lower tier.
- Lower-tier hit: no network transfer or encoding; read zstd from disk, decode, serve.
A storage encoding marker prevents an object from being encoded more than once — a cache layer receiving an already-zstd object from another tier sees the marker and preserves it rather than re-encoding (encoding marker prevents double work).
Testing over one million requests¶
Correctness campaign correlated each request across request logs, Prometheus metrics, and Jaeger traces, covering cache misses, cache hits, single-hop fills, and Tiered Cache fills. Cache keys were varied so each request took a specific path, and traces confirmed where encode/decode occurred.
One performance campaign sent >1 million requests across 10 cache servers — half with Tiered Cache disabled, half enabled — to separate local-cache behavior from cross-tier transfers. The two test assets were ~195 KiB and ~272 KiB and both compressed ~2.8×. Cloudflare is explicit that this was a deliberately compressible test corpus that gives a clean validation signal but does not represent every text object on the Internet — a broader corpus is required before treating the measured ratio as a fleet-wide constant.
Key takeaways¶
-
Compress inside the cache, independent of origin encoding. The CDN traditionally stores whatever encoding the origin sent; Cache Transcoding adds a zstd layer at rest for eligible uncompressed text, so storage efficiency no longer depends on origins compressing their own responses. (Source: sources/2026-09-01-cloudflare-how-we-could-save-petabytes-of-cache-storage-with-zstandard)
-
The cost is asymmetric and the asymmetry is the whole design. Encode is ~2.8× costlier per byte than decode (232 vs 641 MB/s), but encode runs once per fill and decode runs per serve — so the expensive operation is rare and the cheap one is on the hot path. (encode-once-decode-many)
-
Eligibility filtering avoids wasted CPU. Only
200 OK,Content-Encodingunset, compressible-textContent-Type, andContent-Length ≥ 4 KiBqualify; media (21.4% of requests / 63.3% of bytes) is skipped because re-compressing it burns CPU for nothing. (eligibility-gate-before-compression) -
The simplest blanket policy beat the popularity-tiered one. Limiting transcoding to hot content reduced storage savings without proportionally cutting CPU, because decode happens on every serve regardless of popularity. Transcode-all-eligible captured nearly all the benefit within budget.
-
Fewer bytes on disk raises cache density and cuts backbone traffic. Smaller representations mean each server retains more objects (higher effective capacity, less eviction of useful content — cache-density), and the compressed form moving through Tiered Cache reduces cross-data-center bandwidth.
-
The compressed form is preserved across cache tiers; decode happens only at the client-facing hop. An encoding marker stops any tier from re-encoding an already-zstd object. (store-compressed-transfer-compressed-decode-at-edge)
-
zstd level 3 is a conservative starting point, not a ceiling. The level and the 4 KiB threshold are parameters; with the CPU budget now understood, higher levels can be evaluated for better ratios at higher cost.
Operational numbers¶
- Compression ratio on eligible test assets: 2.834× (assets shrink to ~⅓ of on-disk size on average).
- Encode: 4.31 ns/byte ≈ 232 MB/s, paid once per cache fill.
- Decode: 1.56 ns/byte ≈ 641 MB/s, paid on every serve.
- Traffic sample: media = 21.4% of requests / 63.3% of bytes (skipped); compressible text = 67.3% of requests / 22.3% of bytes; ~71% of text arrives uncompressed.
- 4 KiB minimum size leaves out only ~1% of otherwise-eligible bytes.
- Estimated extra CPU under tested traffic/reuse assumptions: "a few percent" at zstd level 3.
- Performance campaign: >1M requests across 10 cache servers, half with Tiered Cache on / half off; test assets ~195 KiB and ~272 KiB.
- Projected fleet-level upside: petabytes of effective cache capacity.
Caveats¶
- Prototype, not shipped. Cache Transcoding is described as prototyped and under evaluation, built during an internship — not a production feature.
- Test corpus is deliberately compressible. The ~2.8× ratio is a clean validation signal, not a fleet-wide constant; a broader content mix is needed before generalizing.
- CPU estimate is model-based. "A few percent" extra CPU holds under the specific traffic and reuse assumptions tested; real fleet traffic may differ.
- Future work named in the post: higher zstd levels, broader content types / object sizes, parameter tuning of the eligibility criteria, and handling range requests, pre-compressed origin responses, and passing the compressed object directly to downstream components that already support it (skipping the decode).
Source¶
- Original: https://blog.cloudflare.com/cache-transcoding/
- Raw markdown:
raw/cloudflare/2026-09-01-how-we-could-save-petabytes-of-cache-storage-with-zstandard-0a4287f3.md
Related¶
- systems/cache-transcoding — the prototype system this post introduces.
- systems/zstandard-zstd — the compressor used (level 3).
- systems/pingora — the proxy framework the encode/decode runs inside.
- systems/cloudflare-cache · systems/cloudflare-smart-tiered-cache — the cache substrate and tiered-cache topology the compressed form flows through.
- systems/brotli — the codec zstd is benchmarked against.
- concepts/compression-codec-tradeoff · cache-density · tiered-caching · concepts/cache-hit-rate · concepts/compression-codec-tradeoff
- encode-once-decode-many · eligibility-gate-before-compression · store-compressed-transfer-compressed-decode-at-edge · encoding-marker-prevents-double-work
- companies/cloudflare