Skip to content

DATABRICKS 2026-09-19

Read original ↗

RADAR: Catch gray failures with anomaly detection

Summary

Databricks describes RADAR (Reliability Anomaly Detection, Alerting, and Root-cause analysis), an internal reliability-monitoring system built to catch gray failures — partial, silent outages where "every dashboard is green" (CPU, latency, liveness all nominal) while a slice of customers quietly fails. The canonical example: a subtle deploy bug makes ~1-in-20 credit-card checkouts fail for hours, and the only signal is a spike in "user-fault"-looking errors (e.g. INVALID_ARGUMENT) that individually look legitimate but collectively are the provider's fault. RADAR is a four-stage pipeline — reliability metrics → anomaly detection → alerting → root-cause analysis — that turns "many users hitting the same error at the same time" into a fired alert and a routed ticket in minutes. It runs an unsupervised, streaming SPOT (Streaming Peaks-Over-Threshold, Extreme Value Theory) model that learns "normal" from the last 14 days and needs only a single risk parameter instead of hand-tuned thresholds. Databricks reports a 95% reduction in incident-discovery time at >90% precision with no human needed to spot the pattern, and stresses the design is metric-agnostic (point it at billing, checkout conversion, model performance/drift, etc.). The post ships a public GitHub scaffold that maps each RADAR stage to a Databricks component and lets an AI agent build the whole system from a single prompt.

Key takeaways

  1. Gray failures are defined by differential observability. The failure detectors read healthy while users clearly see failures — the framing borrowed from Microsoft's Gray Failure: The Achilles' Heel of Cloud-Scale Systems (SOSP/HotOS 2017). Two properties make them sneaky: they are partial (one slice — e.g. a single card type — fails while most users are fine, so there's no crash to catch) and they grow (the affected slice widens over time, expanding blast radius). (Source: sources/2026-09-19-databricks-radar-catch-gray-failures-with-anomaly-detection)
  2. Waiting for customer reports fails three ways: manual, delayed, silent. Someone has to notice the same complaint across a pile of tickets (manual); by the time enough people complain, hours-to-days have passed (delayed); and most affected customers never file a ticket — they just leave (silent). RADAR's thesis: keep reading tickets, but add always-on automatic detection that fires "the moment a lot more customers than usual start hitting the same issue at the same time."
  3. The load-bearing signal is user-error rate, broken down by dimension. At every timestamp record two numbers per error — how many errors are happening AND how many distinct users hit each one — sliced by error code and region. Counting distinct users (not just raw error volume) is what separates a real gray failure (many users, one error) from a single misbehaving client retrying. This yields a rich set of time series describing service health.
  4. SPOT gives thresholdless, streaming anomaly detection. RADAR runs anomaly detection on each series using SPOT — an unsupervised streaming model grounded in Extreme Value Theory (Siffer et al., Anomaly Detection in Streams with Extreme Value Theory, KDD 2017). It learns the distribution of "normal" from the past 14 days and needs only a single risk parameter rather than a pile of hand-tuned per-metric cutoffs — the key to running detection across many series without manual threshold maintenance. (Source: sources/2026-09-19-databricks-radar-catch-gray-failures-with-anomaly-detection)
  5. The alerting layer is enrich → filter → dedupe → route, not raw fire. When a series fires, the alerting stage enriches with context, filters out the insignificant, and dedupes so on-call isn't buried under copies of the same event — then files a ticket routed to the right engineering team for that error. This is the explicit answer to alert fatigue: unsupervised detection over many series would be noisy without a dedupe/filter stage.
  6. Every ticket ships with root-cause context. Each ticket arrives with anomaly deep-dive details plus a link to a dashboard backed by an AI assistant (AI/BI Genie), so the on-call goes straight to "what actually broke" instead of first reconstructing the incident.
  7. Results: 95% reduction in incident-discovery time at >90% precision. Running RADAR on Databricks itself changed the shape of these incidents — from "days of delay waiting on customer tickets" to minutes, with no human needed to spot the pattern, keeping the gray-failure blast radius contained. (Databricks' own numbers; no external validation.)
  8. RADAR is metric-agnostic. The same detect-alert-diagnose pipeline applies anywhere a number can quietly go wrong: financial services (payment/transaction failures, billing anomalies, fraud signals), retail (checkout conversion, cart errors, delivery times), healthcare (patient throughput, claims processing), and AI products (model-performance and data-distribution drift that shows up before a model visibly breaks).
  9. The whole system maps onto stock Databricks components and deploys as one Declarative Asset Bundle (DAB). See stage → component mapping below. Databricks distilled the internal system into a single markdown "recipe" + demo prompt so an AI agent builds a working RADAR live on the platform from your own metric.

Architecture — the four stages of RADAR

Stage What it does Databricks components
1. Reliability metrics Record, per timestamp, error count and distinct-user count, broken down by error code + region → a set of health time series. Zerobus (low-latency ingest), Unity Catalog + Metric View (governance/semantics), Delta Lake (storage)
2. Anomaly detection Run SPOT (unsupervised, streaming, EVT-based; 14-day window, single risk parameter) on each series. MLflow (train), Model Serving (serve endpoint), Workflows (orchestrate recurring jobs)
3. Alerting Enrich → filter → dedupe → file a ticket routed to the owning team. Databricks SQL Alerts (Databricks SQL)
4. Root-cause analysis Ticket carries anomaly deep-dive + link to an AI-assistant-backed dashboard. AI/BI Genie, AI/BI Dashboards

Deployment: the whole thing ships as a single unit through a Declarative Asset Bundle (DAB). Collect+store → a Delta table; detect → a job; alert + dedupe → a ticket; visualize → a dashboard.

Systems / concepts / patterns extracted

Operational numbers

  • 95% reduction in incident-discovery time (days → minutes).
  • >90% precision on fired anomalies.
  • SPOT learns "normal" from a 14-day rolling window; tuned by one risk parameter.
  • Illustrative gray failure: ~1 in 20 credit-card checkouts failing for ~6.5 hours (9:30 deploy → 16:00 fix) before manual escalation in the motivating scenario.

Caveats

  • Tier-3 source; vendor post. All headline numbers (95% / >90%) are Databricks' own, self-reported, with no external validation and no absolute baselines (what was the prior discovery time? which service?).
  • Product-adjacent. The back half is a "build it yourself on Databricks" mapping + GitHub scaffold CTA — included because the front half is genuine reliability-detection architecture (gray-failure framing, SPOT/EVT model choice, the enrich/filter/dedupe alerting stage, the distinct-user-count design decision) well above the 20% architecture bar.
  • SPOT internals not detailed. The post names SPOT + the KDD 2017 paper but does not give the peaks-over-threshold math, the drift-handling variant (DSPOT), or how multi-series alerts are correlated beyond "dedupe."
  • Open-loop. RADAR detects and routes a ticket to a human; it does not auto-remediate. Not an instance of patterns/closed-loop-remediation.

Source

Last updated · 766 distilled / 2,225 read