Skip to content

CONCEPT Cited by 2 sources

Grey failure

Grey failure names a component that is not fully broken but not fully healthy — partially, intermittently, or sub-specification degraded. It is the failure mode that health-check booleans and up/down monitoring are structurally unable to catch: the node answers liveness probes, emits its usual metrics, reports its usual status — and silently drags user-visible behavior down.

Spelling / synonym. Written "grey failure" (British) or "gray failure" (American, and the spelling in the foundational paper and the Databricks RADAR post) — same concept. The academic root is Huang et al., Gray Failure: The Achilles' Heel of Cloud-Scale Systems (HotOS 2017), which names the underlying cause differential observability: the system's own failure detectors read healthy even while users clearly experience failure. (Source: sources/2026-09-19-databricks-radar-catch-gray-failures-with-anomaly-detection)

The term captures two facts:

  1. Binary failure monitoring is insufficient at scale. Most failures on a fleet of thousands of GPUs, drives, or pods are not "node crashed." They are "node is a little slower / lossier / hotter than nominal."
  2. Grey failures compound through fanout. A job that depends on N components tolerates a whole node going down (restart, reroute) far better than one node going slow — because slow propagates through synchronous dependencies and poisons the whole operation. This is the same math as concepts/tail-latency-at-scale.

Canonical examples

Detection strategies

  • High-cardinality correlation. Look for outliers in continuous metrics (per-GPU per-step time, per-NIC packet-drop rate, per-drive p99 IO latency) across a fleet and flag statistical deviations, not threshold crossings.
  • Streaming anomaly detection over per-slice series. Run unsupervised anomaly detection (e.g. SPOT / Extreme Value Theory) on high-cardinality series so a partial failure surfaces as a statistical outlier without hand-tuned thresholds. Count the right dimension: Databricks' RADAR tracks distinct users per error code per region, since "many distinct users, one error" is what distinguishes a gray failure from one client retrying. (Source: sources/2026-09-19-databricks-radar-catch-gray-failures-with-anomaly-detection)
  • Proactive alerting on degradation, not failure. Alert when throughput/latency trends wrong, not when a threshold breaks.
  • Structural isolation. Where individual grey-failing nodes cannot be reliably detected in time, the system is built so any given node's misbehavior is bounded in blast radius (S3's patterns/cell-based-architecture-for-blast-radius-reduction + redundancy-for-heat; partial-restart recovery at the training-job level — partial-restart-fault-recovery).

Why it matters for ML / GPU workloads specifically

A distributed training job often runs synchronous all-reduce / gradient exchanges; any slow rank stalls the collective. One grey-failing GPU throttles the entire job to its pace. The customary response — checkpoint and restart — wastes hours of expensive compute when the grey fault is actually "this one GPU is 10 °C hotter than nominal." The SageMaker HyperPod observability capability is explicitly built around detecting grey failures before they cascade. (Source: sources/2025-08-06-allthingsdistributed-removing-friction-sagemaker-ai-development)

Seen in

  • concepts/observability — grey-failure detection is an advanced requirement on the observability pipeline.
  • monitoring-paradox — if the monitoring layer itself can grey-fail (collectors throttled, agents disk-full), the whole scheme collapses.
  • concepts/tail-latency-at-scale — grey failures are the supply side of the tail-at-scale problem; one grey node × fanout = whole operation's latency.
  • concepts/anomaly-detection — the primitive for catching gray failures that static thresholds structurally miss.
  • concepts/alert-fatigue — unsupervised gray-failure detection over many series needs a filter/dedupe stage or it drowns on-call.
  • heat-management — S3's structural answer to grey-failing HDDs.
  • systems/aws-sagemaker-hyperpod

Merged aliases

  • partial-failure
  • silent-slowdown
Last updated · 766 distilled / 2,225 read