SYSTEM Cited by 1 source
ZGateway (Meta ZippyDB proxy)¶
Definition¶
ZGateway is Meta's stateless proxy tier that sits between ZippyDB clients and the ZippyDB database (ZServer) fleet, unifying ZippyDB client traffic through one managed layer. It is described as "a ZippyDB client run as a managed service": each ZGateway host runs Meta's thick C++ ZippyDB client as its engine (one internal client per use case), which made moving capability onto the gateway natural. It handles >1 billion operations/second and carries ~40% of all ZippyDB traffic (projected past 60%) while adding only ~6% computational overhead to an average use case. (Source: sources/2026-09-03-meta-zgateway-learnings-from-putting-a-proxy-in-front-of-zippydb)
Why ZippyDB needed a proxy layer¶
In the direct-access model every client connected to every database host it needed, producing a dense many-to-many mesh of TLS connections — a typical client held tens of thousands of outbound connections and a typical database host accepted tens of thousands of inbound ones. That mesh was wasteful (idle connections consuming memory, CPU, file descriptors on both ends) and, critically, fan-in scaled with the client population, so every new client cohort degraded every database host. Reconnection storms (a cohort restart, a routing bug that made every client open a connection per shard) breached file-descriptor limits and drove hosts into reboot loops. A proxy tier decouples the two independently-evolving fleets and moves connection management to the one place Meta can solve it. See collapse-many-to-many-mesh-into-two-bounded-hops and persistent-connection-amplification. (Source: sources/2026-09-03-meta-zgateway-learnings-from-putting-a-proxy-in-front-of-zippydb)
Direct access was the right design for ZippyDB's early life; the means to build a shared tier arrived on their own schedule — ServiceRouter feature improvements, Thrift overload protection, a thin client — and building ZGateway earlier would have meant building each of those first.
Architecture¶
ZGateway is a stateless tier deployed as regional tiers discovered through ServiceRouter, keeping every client near its gateway. It runs in two flavors sharing one pipeline: a pure proxy and a read-through cache.
Request path: a client sends over its sticky connection to a regional ZGateway host, which terminates TLS (in the Thrift/ServiceRouter stack), authorizes against the use case's ACLs, and applies per-tenant admission control, validation, and shaping. ZGateway resolves the shard, checks the local cache (on a caching tier), batches and coalesces misses and writes with other in-flight requests for that shard, then sends them to the correct replicas. Responses are demultiplexed back to callers, with per-use-case metrics, traces, and quota usage recorded along the way.
The key property is the asymmetry of connection counts: each client needs only a sticky pool to its regional ZGateway hosts, and each ZServer sees connections only from the ZGateway fleet (whose size Meta controls). Some responsibilities deliberately stay put — TLS in the Thrift/ServiceRouter stack, key-to-shard mapping in the shard locator, replica selection and hedging in the embedded client. ZGateway owns traffic management, not a reimplementation of the database client. (Source: sources/2026-09-03-meta-zgateway-learnings-from-putting-a-proxy-in-front-of-zippydb)
Capabilities¶
- Fan-in/fan-out reduction — collapses the many-to-many mesh into two bounded hops; per-host connections drop ~97–98% and total persistent connections ~19x (model estimate). Database-host fan-in becomes
R · S_host, independent of the client population. See collapse-many-to-many-mesh-into-two-bounded-hops. - Batching and coalescing across clients — a shared batcher groups same-destination requests (keyed by use case + physical shard) into one backend RPC; a coalescer collapses simultaneous same-key reads into a single fetch. Ships with idle-eviction (TTL) and an in-flight cap for OOM safety. Let Meta retire fragile client-side batching libraries. See shared-batcher-across-clients and concepts/request-collapsing.
- Tenant isolation via Discriminant Load Shedding (DLS) — per-tenant priority-split buckets drained round-robin; a flooding tenant sheds only its own excess. Fronted by an AIMD CPU concurrency controller and a memory handler. See discriminant-load-shedding.
- Read caching with live invalidation — in-process cache with a per-key fill lock (thundering-herd protection), refreshed by a CDC stream of write/checkpoint events under a bounded-staleness contract; keyspace sliced by consistent hashing. See concepts/change-data-capture, concepts/thundering-herd.
- Control-plane load balancing — a balancer reads per-host CPU on a fixed cadence and nudges ServiceRouter weighted-consistent-hash weights (damped, clamped, re-centered, change-throttled) to even out load across heterogeneous ~26–126-core hosts; becoming adaptive by classifying tier state.
- Cross-region resilience — global routing, mega-regions, and rings let a saturated regional tier fail over to healthy capacity, each per-tier/region behind a percentage knob, keyed off a sharp overload signal.
- Transactions and richer operations — consolidated client-side transaction bookkeeping (read set, scanned ranges, pending writes) onto one store shared with the engine, rolled out behind a flag in nine phases to 100% of transaction traffic with no reliability regression.
- Safe migration — client-side config flags (percentage knob, region filter, global kill switch) scoped per service and shard prefix; no client code change. See patterns/staged-rollout.
Operating in production¶
ZGateway runs as a large volume of servers across dozens of regions, in a handful of tiers by workload: one large general-purpose tier for the long tail of use cases, dedicated tiers for the largest customers, and a separate high-throughput proxy tier. Tiers differ in footprint by more than an order of magnitude and are non-uniform internally, largely because of stacking (multiple tasks packed onto one machine at varying densities, next to full-size dedicated hosts). Rich per-use-case observability is what makes the admission control and load balancing safe on shared infrastructure.
What's next (stated directions)¶
- Agent-operated heuristics — expose the many hand-tuned control loops (load-shedding thresholds, balancer params, failover triggers, batch flush windows, cache staleness bounds) as a structured control surface so AI agents can diagnose tier state and remediate behind guardrails; the adaptive balancer is "already an agent in all but name."
- Co-location — push part of the gateway down beside the ZServer host so the gateway↔server leg becomes a local call, while the connection-management/admission-control front stays a shared regional tier (avoid re-coupling the fleets deliberately decoupled).
- A multi-process gateway — split the single process into cooperating processes (connection/TLS front-end, request workers, separate cache and transaction components) for hard fault isolation and independent lifecycles.
Seen in¶
- 2026-09-03 Meta — ZGateway: Learnings from Putting a Proxy in Front of ZippyDB (sources/2026-09-03-meta-zgateway-learnings-from-putting-a-proxy-in-front-of-zippydb) — canonical; first wiki disclosure of ZGateway.
Caveats¶
- Connection-collapse figures (~97–98% per host, ~19x total) are a model-based estimate with round mock inputs, illustrating scaling behavior rather than measured fleet numbers.
- Agent-operated heuristics, co-location, and the multi-process split are roadmap directions, not shipped as of the post.
Related¶
- systems/zippydb — the key-value store ZGateway fronts (ZServer = its database fleet).
- systems/servicerouter — Meta's service mesh underpinning discovery, weighted-consistent-hash routing, and cross-region failover.
- systems/fbthrift — the Thrift/ServiceRouter stack where TLS termination lives.
- discriminant-load-shedding — DLS tenant isolation.
- persistent-connection-amplification — the failure mode direct access amplified.
- concepts/stateless-compute — statelessness is what enables free request placement + weight-based balancing.
- collapse-many-to-many-mesh-into-two-bounded-hops — the core structural move.
- shared-batcher-across-clients — cross-client batching + coalescing.
- patterns/central-proxy-choke-point — the general "one managed tier owns shared work" posture.
- companies/meta — canonical company page.