Skip to content

CONCEPT Cited by 1 source

Anomaly detection

Definition

Anomaly detection is the automated identification of observations that deviate from a system's learned "normal" behavior — flagging the moment a metric goes off without a human hand-tuning a threshold for it. In a reliability/observability context it is the primitive that turns a stream of health metrics into "something is wrong right now," and it is the load-bearing mechanism for catching gray failures that boolean health checks and static thresholds structurally miss.

The core problem it solves is threshold sprawl: at scale you have thousands of time series (per error code, per region, per shard), and hand-setting a static cutoff for each is both unmaintainable and wrong — "normal" differs per series and drifts over time. Anomaly detection replaces "alert when value > X" with "alert when value is statistically unlikely given this series' recent history."

Why static thresholds fail

  • Per-series normal. A p99 of 200 ms is fine for one endpoint and a disaster for another; one global threshold can't serve both.
  • Drift. Traffic, seasonality, and deploys move the baseline; a threshold set last quarter is stale this quarter.
  • The gray-failure gap. A gray failure is partial — one slice fails while aggregates stay green — so it never crosses an aggregate threshold. Detection must operate on high-cardinality slices and flag statistical deviation, not threshold crossing. (Source: sources/2026-09-19-databricks-radar-catch-gray-failures-with-anomaly-detection)

Design axes

  • Supervised vs unsupervised. Labeled anomalies are rare and go stale; production reliability detection is almost always unsupervised — learn normal from history, flag outliers.
  • Batch vs streaming. Reliability use cases want streaming: score each new point against a model of the recent past so detection latency is seconds, not the next batch run.
  • Number of tuning knobs. The fewer per-metric parameters, the more series you can cover. The ideal is one global sensitivity/risk parameter that generalizes across series (see SPOT below).
  • What you count. Choosing the right signal matters as much as the detector. Databricks' RADAR counts not just error volume but distinct users per error — the dimension that separates a real incident (many users, one error) from one client retrying. (Source: sources/2026-09-19-databricks-radar-catch-gray-failures-with-anomaly-detection)
  • Detection ≠ alerting. Raw per-series anomaly flags are noisy; production systems bolt an enrich → filter → dedupe → route stage on top so on-call isn't buried (the alert-fatigue defense).

SPOT / Extreme Value Theory

SPOT (Streaming Peaks-Over-Threshold) is the canonical thresholdless streaming detector cited in the RADAR post — from Siffer et al., Anomaly Detection in Streams with Extreme Value Theory (KDD 2017). Instead of a fixed cutoff, it models the tail of a series' distribution using Extreme Value Theory (the Peaks-Over-Threshold / Generalized Pareto approach) and computes, per new point, whether it is an extreme value given the learned tail. Key properties that make it operationally attractive:

  • Unsupervised — no labeled anomalies needed.
  • Streaming — scores points online; a drift-aware variant (DSPOT) detrends first.
  • Single risk parameter — one knob (roughly, the acceptable false-positive probability) instead of per-metric thresholds; this is what lets one model cover many series.
  • Learns "normal" from a rolling window — RADAR uses the past 14 days. (Source: sources/2026-09-19-databricks-radar-catch-gray-failures-with-anomaly-detection)

Where it fits in the stack

Anomaly detection is a stage in a broader reliability pipeline: reliability metrics → anomaly detection → alerting → root-cause analysis (RADAR's four stages). It sits downstream of metric collection (SLIs / high-cardinality series) and upstream of the alerting/routing layer. Detection can be deployed as code (model trained + served, run on a recurring job) — the same "alert artifact as code" discipline as patterns/alerts-as-code.

Beyond reliability: drift detection

The same primitive detects data-distribution / model-performance drift in ML products — an anomaly in prediction distribution or input features that shows up before a model visibly breaks. RADAR is explicitly pitched as metric-agnostic for exactly this reason. (Source: sources/2026-09-19-databricks-radar-catch-gray-failures-with-anomaly-detection)

Seen in

  • sources/2026-09-19-databricks-radar-catch-gray-failures-with-anomaly-detection — canonical wiki instance. Databricks' RADAR runs unsupervised streaming SPOT (Extreme Value Theory, KDD 2017) over per-error / per-region reliability series (error count + distinct-user count), 14-day normal window, single risk parameter, to catch gray failures in minutes; 95% reduction in incident-discovery time at >90% precision (vendor-reported).
Last updated · 766 distilled / 2,225 read