SYSTEM Cited by 1 source
RADAR (Databricks)¶
RADAR — Reliability Anomaly Detection, Alerting, and Root-cause analysis — is a Databricks internal reliability-monitoring system built to catch gray failures: partial, silent outages where every conventional dashboard reads green (CPU, latency, liveness all nominal) while a slice of customers quietly fails. RADAR points anomaly detection at reliability metrics — especially user-error rates — and fires within minutes when many users start hitting the same error at once, filing a routed ticket with root-cause context. (First wiki disclosure: 2026-09-19.)
The problem it targets¶
A textbook gray failure: a routine deploy slips a subtle bug into checkout; ~1-in-20 credit-card payments silently fail; customers retry, give up, and leave; the first support ticket lands hours later and looks like a mistyped card. For ~6.5 hours monitoring insists everything is fine while revenue leaks. The failure is partial (one card type) and grows (widening blast radius) — so it never trips an aggregate threshold. The root cause is differential observability (Microsoft's term): the failure detectors don't notice what users clearly do. (Source: sources/2026-09-19-databricks-radar-catch-gray-failures-with-anomaly-detection)
Architecture — four stages¶
- Reliability metrics. Per timestamp, record error count and distinct-user count, sliced by error code and region. The distinct-user dimension is the discriminator: many users × one error = the provider's fault, not the user's. Yields a rich set of health time series.
- Anomaly detection. Run SPOT (Streaming Peaks-Over-Threshold, Extreme Value Theory — Siffer et al., KDD 2017) on each series: unsupervised, streaming, learns "normal" from the past 14 days, tuned by a single risk parameter rather than per-metric thresholds. This is what lets one model cover many series without hand-tuning.
- Alerting. On fire: enrich with context → filter out the insignificant → dedupe so on-call isn't buried → file a ticket routed to the owning engineering team. The dedupe/filter stage is the alert-fatigue defense that unsupervised multi-series detection requires.
- Root-cause analysis. Each ticket carries anomaly deep-dive details + a link to an AI/BI Genie-backed dashboard so the on-call goes straight to "what broke."
Stage → Databricks component mapping¶
- Reliability metrics — Zerobus (low-latency ingest), Unity Catalog + Metric View (governance/semantics), Delta Lake (storage). See also patterns/telemetry-to-lakehouse.
- Anomaly detection — MLflow (train), Model Serving (serve the model endpoint), Workflows (orchestrate recurring jobs).
- Alerting — Databricks SQL Alerts (Databricks SQL).
- Root-cause analysis — AI/BI Genie + AI/BI Dashboards.
The whole system deploys as a single unit via a Declarative Asset Bundle (DAB). Databricks published a public GitHub scaffold (one markdown recipe + demo prompt) so an AI agent builds a working RADAR live on the platform from a user-supplied metric.
Reported results¶
- 95% reduction in incident-discovery time (days → minutes).
- >90% precision on fired anomalies.
- No human needed to spot the pattern; gray-failure blast radius stays contained. (Databricks' own numbers; no external validation, no absolute baseline given.)
Metric-agnostic by design¶
RADAR "doesn't care what the metric is." Beyond user errors it applies to financial services (payment/transaction failures, billing anomalies, fraud signals), retail (checkout conversion, cart errors, delivery times), healthcare (patient throughput, claims processing), and AI products (model-performance / data-distribution drift that surfaces before a model visibly breaks).
Caveats¶
- Tier-3 vendor post; headline numbers self-reported and unbaselined.
- Open-loop: RADAR detects and routes a ticket to a human — it does not auto-remediate, so it is not an instance of patterns/closed-loop-remediation.
- SPOT internals (POT math, drift-aware DSPOT variant, multi-series correlation beyond "dedupe") are not detailed in the post.
Seen in¶
- sources/2026-09-19-databricks-radar-catch-gray-failures-with-anomaly-detection — first and canonical wiki disclosure; the four-stage pipeline, SPOT/EVT model choice, distinct-user-count design decision, DAB packaging, and the 95%/>90% results all attributed here.
Related¶
- concepts/anomaly-detection — the detection primitive (SPOT / EVT).
- concepts/grey-failure — the failure mode RADAR targets.
- concepts/observability — RADAR is an advanced observability capability.
- concepts/alert-fatigue — answered by the enrich/filter/dedupe stage.
- systems/zerobus-ingest — the ingest layer for reliability metrics.
- systems/databricks-genie — the root-cause-analysis assistant.
- companies/databricks