Skip to content

SYSTEM Cited by 1 source

RADAR (Databricks)

RADAR — Reliability Anomaly Detection, Alerting, and Root-cause analysis — is a Databricks internal reliability-monitoring system built to catch gray failures: partial, silent outages where every conventional dashboard reads green (CPU, latency, liveness all nominal) while a slice of customers quietly fails. RADAR points anomaly detection at reliability metrics — especially user-error rates — and fires within minutes when many users start hitting the same error at once, filing a routed ticket with root-cause context. (First wiki disclosure: 2026-09-19.)

The problem it targets

A textbook gray failure: a routine deploy slips a subtle bug into checkout; ~1-in-20 credit-card payments silently fail; customers retry, give up, and leave; the first support ticket lands hours later and looks like a mistyped card. For ~6.5 hours monitoring insists everything is fine while revenue leaks. The failure is partial (one card type) and grows (widening blast radius) — so it never trips an aggregate threshold. The root cause is differential observability (Microsoft's term): the failure detectors don't notice what users clearly do. (Source: sources/2026-09-19-databricks-radar-catch-gray-failures-with-anomaly-detection)

Architecture — four stages

  1. Reliability metrics. Per timestamp, record error count and distinct-user count, sliced by error code and region. The distinct-user dimension is the discriminator: many users × one error = the provider's fault, not the user's. Yields a rich set of health time series.
  2. Anomaly detection. Run SPOT (Streaming Peaks-Over-Threshold, Extreme Value Theory — Siffer et al., KDD 2017) on each series: unsupervised, streaming, learns "normal" from the past 14 days, tuned by a single risk parameter rather than per-metric thresholds. This is what lets one model cover many series without hand-tuning.
  3. Alerting. On fire: enrich with context → filter out the insignificant → dedupe so on-call isn't buried → file a ticket routed to the owning engineering team. The dedupe/filter stage is the alert-fatigue defense that unsupervised multi-series detection requires.
  4. Root-cause analysis. Each ticket carries anomaly deep-dive details + a link to an AI/BI Genie-backed dashboard so the on-call goes straight to "what broke."

Stage → Databricks component mapping

The whole system deploys as a single unit via a Declarative Asset Bundle (DAB). Databricks published a public GitHub scaffold (one markdown recipe + demo prompt) so an AI agent builds a working RADAR live on the platform from a user-supplied metric.

Reported results

  • 95% reduction in incident-discovery time (days → minutes).
  • >90% precision on fired anomalies.
  • No human needed to spot the pattern; gray-failure blast radius stays contained. (Databricks' own numbers; no external validation, no absolute baseline given.)

Metric-agnostic by design

RADAR "doesn't care what the metric is." Beyond user errors it applies to financial services (payment/transaction failures, billing anomalies, fraud signals), retail (checkout conversion, cart errors, delivery times), healthcare (patient throughput, claims processing), and AI products (model-performance / data-distribution drift that surfaces before a model visibly breaks).

Caveats

  • Tier-3 vendor post; headline numbers self-reported and unbaselined.
  • Open-loop: RADAR detects and routes a ticket to a human — it does not auto-remediate, so it is not an instance of patterns/closed-loop-remediation.
  • SPOT internals (POT math, drift-aware DSPOT variant, multi-series correlation beyond "dedupe") are not detailed in the post.

Seen in

Last updated · 766 distilled / 2,225 read