Skip to content

Cloudflare — How we saved 100 terabytes of memory by optimizing 1.1.1.1's DNS cache

Summary

Cloudflare's Big Pineapple — the in-house recursive resolver behind systems/cloudflare-1-1-1-1-resolver|1.1.1.1, Gateway DNS, DNS Firewall, and AS112 — stores over 250 billion DNS cache entries across the fleet at any moment. At that scale a single wasted byte per entry costs >250 GB of RAM. Five successive changes to how each cache entry is laid out in memory cut the per-entry footprint by 56% (953 → 420 bytes), freeing roughly 100 terabytes across the fleet (≈130 Gen 13 servers worth of RAM). The optimizations also made the cache faster — insert throughput +43%, lookup latency −19% — because fewer allocations and better memory locality traded space for speed rather than against it. Every change is a Rust struct-layout technique applied to an immutable-after-insert data structure. (Source: sources/2026-08-27-cloudflare-how-we-saved-100-terabytes-of-memory-by-optimizing-1111s-dns)

What is cached

Each cache item is a key–value pair. The key (CacheKey) is {qname, qtype, authenticated, tag}; the value (CacheEntry) holds the DNS response's answer/authority/additional record sections plus metadata (creation timestamp, inception: Instant, ttl, hits: u32, extended errors). Once an entry is stored it is never modified again — this immutability is what unlocks every optimization. Cache size varies by data center; EDNS Client Subnet (ECS) locations cache many versions of the same query (one per client network), multiplying both entry count and per-entry memory, so ECS-heavy POPs benefit most.

Benchmarking method

They fill the cache with randomly generated entries matching production mix (56% A, 25% AAAA, 19% TXT; TXT stands in for all variable-length types, sized 64–224 bytes; 1–4 records per entry). A custom global allocator wrapping Rust's System allocator records allocation count + size per entry, alongside insert throughput and lookup latency. Benchmark numbers approximate but don't reproduce production, so they also measured resident memory across production instances during rollout.

Key takeaways

  1. At 250 billion entries, per-entry byte-shaving is a fleet-scale memory lever. One byte per entry ≈ 250 GB fleetwide; the whole project reclaimed ~100 TB by removing ~533 bytes per entry. (Source: sources/2026-08-27-cloudflare-how-we-saved-100-terabytes-of-memory-by-optimizing-1111s-dns)

  2. Immutable-after-insert means the growth machinery is pure waste. Vec<T> and String each carry an 8-byte capacity field plus over-allocated heap slack for future pushes. Replacing them with Box<[T]> / Box<str> drops the capacity field and the slack. 8 such fields × 8 bytes = 64 bytes/entry, >15 TB fleetwide. (box-slice-over-vec-for-immutable-data)

  3. Collapsing separate lists into one buffer + small offsets removes pointers. Instead of three Box<[Record]> for answer/authority/ additional (each an 8-byte pointer + 8-byte length), store one list plus two u16 (2-byte) section offsets — record counts fit in a u16. Removing two lists (32 bytes) for two offsets (4 bytes) saved 28 bytes/entry. (single-buffer-with-section-offsets)

  4. Removing a small field can remove more than its own size, via alignment padding. Rust rounds a struct up to a multiple of its alignment and inserts padding; packing several booleans into one bitflag shrank the struct by more than the booleans' raw bytes because it eliminated surrounding padding.

  5. A field that's usually redundant can be inferred instead of stored. Every DNS record has an owner (the domain it belongs to). Usually the owner equals the queried name; only when a CNAME is involved does it differ. Change owner to Option<Box<Name>>: None means "same as the query," reconstructed from the cache key at read time (no heap alloc); Some stores the differing name. Most records need no owner allocation. (infer-field-from-context-key)

  6. A Rust enum is always as big as its largest variant — box the big ones. RecordData is a sum type over record types; NAPTR (136 B, 144 B with tag+padding) is the largest, but **A (4 B) + AAAA (16 B) are

    80% of traffic, so most records wasted >120 bytes on padding. Boxing only the large variants (Txt(Box<Txt>), Naptr(Box<Naptr>), …) shrinks the inline enum to 24 bytes and sizes each heap allocation to its actual data. Saves 120 bytes per A/AAAA record**; rare NAPTR pays slightly more. (box-large-enum-variants)

  7. Boxing has two costs — allocator size-class rounding and lost locality — and storing records in wire format eliminates both. Big Pineapple uses jemalloc, which rounds each allocation up to a fixed size-class bin (an MX record asks for 40 B, rounds to a 48 B bin, wastes 8). And each boxed variant is a separate heap region, so records scatter — a lookup follows pointers to cold cache lines. The fix: store the record data as a single Box<[u8]> of length-prefixed raw bytes, one contiguous allocation for all records, no per-variant boxing. (records-in-wire-format-in-cache)

  8. Storing records already in wire format also removes serialization work on the read path. A, AAAA, TXT, and all DNSSEC record types can be memcpy'd straight from the buffer into the outgoing DNS message; only records containing domain names (CNAME, NS, MX, SOA) still need parsing to apply DNS name compression. Since direct-copyable types dominate traffic, this cut lookup latency ~5% on its own. The trade-off: records can no longer be randomly indexed — you iterate the buffer sequentially — but record counts per entry are small so the cost is negligible (mildly complicates round-robin rotation).

  9. A reusable scratchspace buffer + one memcpy beats per-record allocation and beats shrinking a Vec. Records are serialized into a persistent scratch buffer that survives across insertions (so it rarely reallocates), then a single Box<[u8]> is allocated and the bytes memcpy'd in. This replaces N per-record allocations with one, and avoids the wasted tail an allocator may not reclaim when a Vec<u8> is shrunk. This change alone lifted insert throughput +13%. (reusable-scratchspace-buffer)

  10. Deliberately traded memory for lookup speed where it mattered. DNS wire format compresses repeated owner names (RFC 1035 §4.1.4) with 2-byte pointers, but following those pointers on the hot path is expensive — so the cache stores full owner names (trading memory for speed) except the common redundant-owner case (takeaway 5). They also rejected caching the full wire-format message because DNSSEC records are only included when the client sets the DO flag, forcing either two cached variants or filtering an already-built message.

Operational numbers

Metric Before After Change
Per-entry net footprint 953 bytes 420 bytes −56%
Per-entry allocations 1.1 KB 461 bytes −58%
Cache insert throughput 625,000 entries/s 893,000 entries/s +43%
Cache lookup latency 828 ns 670 ns −19%
  • Fleet scale: >250 billion cache entries; 1 byte/entry ≈ 250 GB.
  • Aggregate reclaimed: ~100 TB working-set memory (≈130 Gen 13 servers).
  • Per-instance resident memory: p99 9.3 → 5.3 GB (−43%), p90 6.5 → 3.8 GB (−42%); fuller-cache instances saved the most absolute.
  • Rollout window: 2026-05-18 (start) → 2026-07-06 (complete across all services); memory dropped in steps, one or more optimizations per release. Restarted instances start with empty caches and climb as caches refill; steady-state plateaus are the true measure, not initial dips.
  • Individual wins: box-large-variants saves 120 B/record for A/AAAA; Box<[T]>/Box saves 64 B/entry (>15 TB fleetwide); single-buffer + offsets saves 28 B/entry; wire-format lookup path −5% latency; scratch buffer +13% insert throughput.
  • Future work: reinvest freed memory into larger cache capacity (higher hit rate, fewer upstream queries) without increasing memory usage.

Caveats

  • Benchmark memory numbers approximate production (which also depends on traffic mix, occupancy, allocator state, and non-cache memory); production resident-memory reduction (−43% p99) is smaller than the benchmarked per-entry reduction (−56%) because resident memory includes everything besides the cache.
  • Techniques are Rust-specific in expression (Vec/Box<[T]>, enum sizing, Option<Box<T>>) but the principles — immutable data needs no growth headroom, box large sum-type variants, store hot data contiguously in a serialization-friendly layout, infer redundant fields — generalize.
  • The wire-format layout gives up random indexing of records; only viable because per-entry record counts are small.

Source

Last updated · 766 distilled / 2,225 read