RADAR: Catch gray failures with anomaly detection¶
Summary¶
Databricks describes RADAR (Reliability Anomaly Detection, Alerting,
and Root-cause analysis), an internal reliability-monitoring system built
to catch gray failures — partial, silent outages where "every dashboard
is green" (CPU, latency, liveness all nominal) while a slice of customers
quietly fails. The canonical example: a subtle deploy bug makes ~1-in-20
credit-card checkouts fail for hours, and the only signal is a spike in
"user-fault"-looking errors (e.g. INVALID_ARGUMENT) that individually look
legitimate but collectively are the provider's fault. RADAR is a four-stage
pipeline — reliability metrics → anomaly detection → alerting →
root-cause analysis — that turns "many users hitting the same error at the
same time" into a fired alert and a routed ticket in minutes. It runs an
unsupervised, streaming SPOT (Streaming Peaks-Over-Threshold, Extreme
Value Theory) model that learns "normal" from the last 14 days and needs
only a single risk parameter instead of hand-tuned thresholds. Databricks
reports a 95% reduction in incident-discovery time at >90% precision
with no human needed to spot the pattern, and stresses the design is
metric-agnostic (point it at billing, checkout conversion, model
performance/drift, etc.). The post ships a public GitHub scaffold that maps
each RADAR stage to a Databricks component and lets an AI agent build the
whole system from a single prompt.
Key takeaways¶
- Gray failures are defined by differential observability. The failure detectors read healthy while users clearly see failures — the framing borrowed from Microsoft's Gray Failure: The Achilles' Heel of Cloud-Scale Systems (SOSP/HotOS 2017). Two properties make them sneaky: they are partial (one slice — e.g. a single card type — fails while most users are fine, so there's no crash to catch) and they grow (the affected slice widens over time, expanding blast radius). (Source: sources/2026-09-19-databricks-radar-catch-gray-failures-with-anomaly-detection)
- Waiting for customer reports fails three ways: manual, delayed, silent. Someone has to notice the same complaint across a pile of tickets (manual); by the time enough people complain, hours-to-days have passed (delayed); and most affected customers never file a ticket — they just leave (silent). RADAR's thesis: keep reading tickets, but add always-on automatic detection that fires "the moment a lot more customers than usual start hitting the same issue at the same time."
- The load-bearing signal is user-error rate, broken down by dimension. At every timestamp record two numbers per error — how many errors are happening AND how many distinct users hit each one — sliced by error code and region. Counting distinct users (not just raw error volume) is what separates a real gray failure (many users, one error) from a single misbehaving client retrying. This yields a rich set of time series describing service health.
- SPOT gives thresholdless, streaming anomaly detection. RADAR runs anomaly detection on each series using SPOT — an unsupervised streaming model grounded in Extreme Value Theory (Siffer et al., Anomaly Detection in Streams with Extreme Value Theory, KDD 2017). It learns the distribution of "normal" from the past 14 days and needs only a single risk parameter rather than a pile of hand-tuned per-metric cutoffs — the key to running detection across many series without manual threshold maintenance. (Source: sources/2026-09-19-databricks-radar-catch-gray-failures-with-anomaly-detection)
- The alerting layer is enrich → filter → dedupe → route, not raw fire. When a series fires, the alerting stage enriches with context, filters out the insignificant, and dedupes so on-call isn't buried under copies of the same event — then files a ticket routed to the right engineering team for that error. This is the explicit answer to alert fatigue: unsupervised detection over many series would be noisy without a dedupe/filter stage.
- Every ticket ships with root-cause context. Each ticket arrives with anomaly deep-dive details plus a link to a dashboard backed by an AI assistant (AI/BI Genie), so the on-call goes straight to "what actually broke" instead of first reconstructing the incident.
- Results: 95% reduction in incident-discovery time at >90% precision. Running RADAR on Databricks itself changed the shape of these incidents — from "days of delay waiting on customer tickets" to minutes, with no human needed to spot the pattern, keeping the gray-failure blast radius contained. (Databricks' own numbers; no external validation.)
- RADAR is metric-agnostic. The same detect-alert-diagnose pipeline applies anywhere a number can quietly go wrong: financial services (payment/transaction failures, billing anomalies, fraud signals), retail (checkout conversion, cart errors, delivery times), healthcare (patient throughput, claims processing), and AI products (model-performance and data-distribution drift that shows up before a model visibly breaks).
- The whole system maps onto stock Databricks components and deploys as one Declarative Asset Bundle (DAB). See stage → component mapping below. Databricks distilled the internal system into a single markdown "recipe" + demo prompt so an AI agent builds a working RADAR live on the platform from your own metric.
Architecture — the four stages of RADAR¶
| Stage | What it does | Databricks components |
|---|---|---|
| 1. Reliability metrics | Record, per timestamp, error count and distinct-user count, broken down by error code + region → a set of health time series. | Zerobus (low-latency ingest), Unity Catalog + Metric View (governance/semantics), Delta Lake (storage) |
| 2. Anomaly detection | Run SPOT (unsupervised, streaming, EVT-based; 14-day window, single risk parameter) on each series. | MLflow (train), Model Serving (serve endpoint), Workflows (orchestrate recurring jobs) |
| 3. Alerting | Enrich → filter → dedupe → file a ticket routed to the owning team. | Databricks SQL Alerts (Databricks SQL) |
| 4. Root-cause analysis | Ticket carries anomaly deep-dive + link to an AI-assistant-backed dashboard. | AI/BI Genie, AI/BI Dashboards |
Deployment: the whole thing ships as a single unit through a Declarative Asset Bundle (DAB). Collect+store → a Delta table; detect → a job; alert + dedupe → a ticket; visualize → a dashboard.
Systems / concepts / patterns extracted¶
- Systems: RADAR (subject; new page), Zerobus Ingest, Unity Catalog, Metric Views, Delta Lake, MLflow, Model Serving, Databricks SQL, AI/BI Dashboards, AI/BI Genie.
- Concepts: anomaly detection (new page — SPOT / Extreme Value Theory as the reusable primitive), gray failure (core topic; enriched with the differential-observability framing + user-error-spike detector), observability (RADAR is an advanced observability capability), alert fatigue (answered by the enrich/filter/dedupe stage), blast radius (gray failures grow; early detection contains it).
- Patterns: alerts as code (RADAR is a code-deployed alerting pipeline shipped as a DAB), telemetry to lakehouse (reliability metrics land in Delta via Zerobus), patterns/closed-loop-remediation (RADAR stops at routed ticket to a human, not auto-fix — the open-loop end of the maturity spectrum; cited as contrast, not instance).
- Non-page mentions (prose/tags only): differential observability (Microsoft's term for the gray-failure root cause), distinct-user count as the discriminating dimension, SPOT single-risk-parameter tuning, DAB one-unit deploy, AI-agent-builds-the-system-from-a-prompt scaffold.
Operational numbers¶
- 95% reduction in incident-discovery time (days → minutes).
- >90% precision on fired anomalies.
- SPOT learns "normal" from a 14-day rolling window; tuned by one risk parameter.
- Illustrative gray failure: ~1 in 20 credit-card checkouts failing for ~6.5 hours (9:30 deploy → 16:00 fix) before manual escalation in the motivating scenario.
Caveats¶
- Tier-3 source; vendor post. All headline numbers (95% / >90%) are Databricks' own, self-reported, with no external validation and no absolute baselines (what was the prior discovery time? which service?).
- Product-adjacent. The back half is a "build it yourself on Databricks" mapping + GitHub scaffold CTA — included because the front half is genuine reliability-detection architecture (gray-failure framing, SPOT/EVT model choice, the enrich/filter/dedupe alerting stage, the distinct-user-count design decision) well above the 20% architecture bar.
- SPOT internals not detailed. The post names SPOT + the KDD 2017 paper but does not give the peaks-over-threshold math, the drift-handling variant (DSPOT), or how multi-series alerts are correlated beyond "dedupe."
- Open-loop. RADAR detects and routes a ticket to a human; it does not auto-remediate. Not an instance of patterns/closed-loop-remediation.
Source¶
- Original: https://www.databricks.com/blog/radar-catch-gray-failures-anomaly-detection
- Raw markdown:
raw/databricks/2026-09-19-radar-catch-gray-failures-with-anomaly-detection-31b58922.md
Related¶
- concepts/grey-failure — the failure mode RADAR targets.
- concepts/anomaly-detection — the detection primitive (SPOT / EVT).
- concepts/observability — RADAR is an advanced observability capability.
- concepts/alert-fatigue — answered by enrich/filter/dedupe.
- systems/databricks-radar — the system.
- companies/databricks