CONCEPT Cited by 1 source
Anomaly detection¶
Definition¶
Anomaly detection is the automated identification of observations that deviate from a system's learned "normal" behavior — flagging the moment a metric goes off without a human hand-tuning a threshold for it. In a reliability/observability context it is the primitive that turns a stream of health metrics into "something is wrong right now," and it is the load-bearing mechanism for catching gray failures that boolean health checks and static thresholds structurally miss.
The core problem it solves is threshold sprawl: at scale you have thousands of time series (per error code, per region, per shard), and hand-setting a static cutoff for each is both unmaintainable and wrong — "normal" differs per series and drifts over time. Anomaly detection replaces "alert when value > X" with "alert when value is statistically unlikely given this series' recent history."
Why static thresholds fail¶
- Per-series normal. A p99 of 200 ms is fine for one endpoint and a disaster for another; one global threshold can't serve both.
- Drift. Traffic, seasonality, and deploys move the baseline; a threshold set last quarter is stale this quarter.
- The gray-failure gap. A gray failure is partial — one slice fails while aggregates stay green — so it never crosses an aggregate threshold. Detection must operate on high-cardinality slices and flag statistical deviation, not threshold crossing. (Source: sources/2026-09-19-databricks-radar-catch-gray-failures-with-anomaly-detection)
Design axes¶
- Supervised vs unsupervised. Labeled anomalies are rare and go stale; production reliability detection is almost always unsupervised — learn normal from history, flag outliers.
- Batch vs streaming. Reliability use cases want streaming: score each new point against a model of the recent past so detection latency is seconds, not the next batch run.
- Number of tuning knobs. The fewer per-metric parameters, the more series you can cover. The ideal is one global sensitivity/risk parameter that generalizes across series (see SPOT below).
- What you count. Choosing the right signal matters as much as the detector. Databricks' RADAR counts not just error volume but distinct users per error — the dimension that separates a real incident (many users, one error) from one client retrying. (Source: sources/2026-09-19-databricks-radar-catch-gray-failures-with-anomaly-detection)
- Detection ≠ alerting. Raw per-series anomaly flags are noisy; production systems bolt an enrich → filter → dedupe → route stage on top so on-call isn't buried (the alert-fatigue defense).
SPOT / Extreme Value Theory¶
SPOT (Streaming Peaks-Over-Threshold) is the canonical thresholdless streaming detector cited in the RADAR post — from Siffer et al., Anomaly Detection in Streams with Extreme Value Theory (KDD 2017). Instead of a fixed cutoff, it models the tail of a series' distribution using Extreme Value Theory (the Peaks-Over-Threshold / Generalized Pareto approach) and computes, per new point, whether it is an extreme value given the learned tail. Key properties that make it operationally attractive:
- Unsupervised — no labeled anomalies needed.
- Streaming — scores points online; a drift-aware variant (DSPOT) detrends first.
- Single risk parameter — one knob (roughly, the acceptable false-positive probability) instead of per-metric thresholds; this is what lets one model cover many series.
- Learns "normal" from a rolling window — RADAR uses the past 14 days. (Source: sources/2026-09-19-databricks-radar-catch-gray-failures-with-anomaly-detection)
Where it fits in the stack¶
Anomaly detection is a stage in a broader reliability pipeline: reliability metrics → anomaly detection → alerting → root-cause analysis (RADAR's four stages). It sits downstream of metric collection (SLIs / high-cardinality series) and upstream of the alerting/routing layer. Detection can be deployed as code (model trained + served, run on a recurring job) — the same "alert artifact as code" discipline as patterns/alerts-as-code.
Beyond reliability: drift detection¶
The same primitive detects data-distribution / model-performance drift in ML products — an anomaly in prediction distribution or input features that shows up before a model visibly breaks. RADAR is explicitly pitched as metric-agnostic for exactly this reason. (Source: sources/2026-09-19-databricks-radar-catch-gray-failures-with-anomaly-detection)
Seen in¶
- sources/2026-09-19-databricks-radar-catch-gray-failures-with-anomaly-detection — canonical wiki instance. Databricks' RADAR runs unsupervised streaming SPOT (Extreme Value Theory, KDD 2017) over per-error / per-region reliability series (error count + distinct-user count), 14-day normal window, single risk parameter, to catch gray failures in minutes; 95% reduction in incident-discovery time at >90% precision (vendor-reported).
Related¶
- concepts/grey-failure — the failure mode anomaly detection is built to catch; static thresholds structurally miss it.
- concepts/observability — anomaly detection is an advanced requirement on the observability pipeline.
- concepts/alert-fatigue — unsupervised detection over many series is noisy without a filter/dedupe stage.
- concepts/service-level-indicator — the metrics anomaly detection runs on.
- patterns/alerts-as-code — deploying detection/alerting as reviewable artifacts.