SYSTEM Cited by 1 source
Databricks Network Configuration Delivery¶
Definition¶
Databricks Network Configuration Delivery is the internal system that supplies each Databricks serverless VM with the network configuration it needs before running a customer workload — allowed storage destinations, private-link endpoints, Unity Catalog grants, and Delta Sharing destinations. Because serverless compute launches tens of millions of VMs per day and each node fetches config at startup and polls for updates over its lifetime, the system serves billions of network-config requests per day across AWS, Azure, and GCP. (Source: sources/2026-08-12-databricks-network-configuration-delivery-to-tens-of-millions-of-serverless-vms)
Network configuration is not stored in any single place; it must be assembled from multiple upstream services, each owning one piece of the picture. The engineering story of this system is the migration from assembling that picture synchronously on the cluster-launch critical path to assembling it asynchronously in the background and serving a pre-computed snapshot.
Old architecture (synchronous aggregation)¶
Every serverless cluster start caused the network-config service to synchronously call all upstream services, aggregate their responses, compute the per-workspace config, and return it — on the critical path of cluster creation. Two structural problems:
- Latency. With multiple upstream services in series, RPC p99 for serving network config was ~5,000 ms, inflating cluster startup.
- Compound availability. Each upstream has its own availability; several in series multiply down fast, raising the annual rate of cluster-launch failures.
It also re-did expensive, often duplicated computation on every request, so load grew with the number of tenants and their configured resources.
New architecture (event-driven precomputation)¶
A ground-up re-architecture on three principles:
- Event-driven pipeline. Instead of synchronous fan-out, the system subscribes to change events via a message queue. When a customer creates a Unity Catalog connection or edits a network policy, the upstream service emits an event; the system processes it and updates the pre-computed config. See concepts/event-driven-architecture.
- Snapshot pre-computation. Configs are computed asynchronously in the background and stored in a pre-computed snapshot store. The serving path becomes a single thin storage fetch, fully decoupled from upstream services. See snapshot-precomputation-serving-path.
- Static stability. On any upstream outage the system keeps serving a static (last-known-good) config, so serverless clusters keep launching.
Two paths¶
- Management path (async, background): upstream services emit change events (IDs only) to a message queue → an event processor resolves which workspaces are affected and fans out per-workspace update notifications → a per-partition event manager fetches the relevant upstream details, recomputes the workspace's config, and writes it to the snapshot store with a new version mark. A periodic reconciler re-syncs all workspaces as a backstop (reconciler-as-safety-net).
- Serving path (critical, fast): a starting cluster reads its config directly from the snapshot store — one read, no upstream calls — which also sheds load off the upstream services.
This is a textbook control / data plane separation: the slow, changeable assembly logic is kept off the fast, high-volume serving path.
Key design decisions¶
- Push events + low-frequency reconciler = reliability of a synchronous framework with the efficiency of push.
- Partition-colocated computation. Configs are computed and stored locally within each service partition, co-located with the workspaces they serve — distributing compute, shrinking blast radius, and removing cross-partition dependencies from the serving path.
- Identifier-only events. Events carry only workspace and resource IDs, keeping them lightweight, idempotent (replayable in any order), and free of sensitive customer data.
- Stage-based extensibility. Adding a new upstream data source is a new stage implementation with zero changes to the core pipeline.
Impact¶
| Metric | Before | After |
|---|---|---|
| RPC latency (p99) | ~5,000 ms | 125 ms (−97.5%) |
| Service success rate | 99.8% | 99.99% |
| Upstream call volume | baseline | −86% |
The legacy synchronous framework was fully deprecated. Today the system serves billions of requests/day at ~125 ms latency and 99.99% availability.
Related¶
- systems/databricks-serverless-compute
- systems/unity-catalog
- snapshot-precomputation-serving-path
- reconciler-as-safety-net
- concepts/event-driven-architecture
- concepts/static-stability
- event-driven-snapshot-precomputation
- partition-colocated-precomputation