Skip to content

SYSTEM Cited by 1 source

Pingora Backend Router (PBR)

Pingora Backend Router (PBR) is Cloudflare's internal load-balancing service, built on the Pingora Rust proxy framework. Its job is to route cacheable requests to servers by URL using consistent hashing, so that each data center keeps only one copy of a given file and every server has a stable, well-known place to find it (Source: sources/2026-09-18-cloudflare-saving-another-100tb-of-ram-with-math-and-rust).

What it does

  • Maps servers (by representative value, e.g. IP address) and tasks (by cache key / URL) onto a shared 32-bit hash number line via systems/pingora-ketama.
  • A task is served by the first server to its left on the ring (equivalently, the next server clockwise on the ring topology).
  • Uses ketama weighting by disk space — servers with more storage get proportionally more hashes and therefore more of the cacheable request load.

Why it got memory-hungry

Not every server can serve every request — compliance requirements and enabled caching features mean only a subset of servers is eligible per request. Because eligibility is a combination of features, PBR needs a separate consistent-hash ring per feature combination: 2^handful = dozens of rings, each holding base-160 × weight hashes per server. Weighted hash counts can reach k = 160 × 625 = 100,000 per server. The aggregate blew up to as much as 6 GB per instance — the finding in Ivan's original Performance-team ticket.

The 100 TB fix

Two changes to the underlying systems/pingora-ketama library reclaimed >100 TB of RAM globally:

  1. Struct packing — shrank each hash Point from 8 bytes to 6 bytes (index fits in u16; raw-byte-array storage sidesteps Rust alignment padding). A flat 25 % cut.
  2. 90 % fewer hashes per server — a first-principles derivation of the load-balancing error formula (coefficient of variation) plus 32-bit hash-collision analysis showed the ring could carry an order-of-magnitude fewer hashes with no appreciable increase in load imbalance.

Migration without melting origins

Changing the ring moves where cacheable requests land, which would invalidate cached content. PBR therefore carried both the old and new rings in memory simultaneously and chose per request via the normal migration framework (stable per request hash → clean rollback). Rollout was controlled on two independent dimensions — traffic fraction on the new ring, and which data centers were allowed to move — starting at small validation locations and expanding by DC groups. See patterns/shadow-migration and patterns/progressive-configuration-rollout.

Scale / numbers

  • Up to 6 GB ketama memory per instance before the fix.
  • Dozens of rings (one per feature combination) held in memory.
  • Weighted hash counts up to ~100,000 per server.
  • >100 TB RAM reclaimed globally after the fix.

Seen in

Last updated · 766 distilled / 2,225 read