CONCEPT Cited by 1 source
Goodput¶
Definition¶
Goodput in large-scale GPU training is the proportion of GPU time spent on productive computation rather than waiting on inputs or recovering from failures. Databricks frames it as the single metric that determines training efficiency and, by extension, total GPU spend: "At scale, your training efficiency is determined by a single metric: goodput." (Source: sources/2026-08-28-databricks-fast-fault-tolerant-pytorch-training-on-ai-runtime)
It is distinct from raw GPU utilization and from Model FLOPs Utilization (MFU): a GPU can be 100% "busy" recomputing work lost to a crash, or idling in a synchronized step waiting for a slow peer — neither is goodput. Goodput nets out both idle time (starved by the data pipeline) and wasted time (recomputing work lost since the last checkpoint).
The two goodput destroyers¶
- Recovery waste — when a job fails, everything since the last valid checkpoint is lost and must be recomputed. Because failures are the expected case at scale (see gpu-training-failure-modes), the checkpointing subsystem sets how much goodput failures cost. Expected wasted work per failure ≈ half the checkpoint interval.
- Idle waste — when the data pipeline can't keep pace, accelerators sit idle waiting for the next batch (see gpu-stall-from-storage). This silently erodes goodput on every step, failure or not.
The arithmetic (Databricks' worked example)¶
Using Llama 3's cited ~8.6 interruptions/day and expected-waste ≈ half the interval:
- Checkpoint every 2 hours → expect to waste ~8.6 h/day recomputing → 64% goodput.
- Checkpoint every 30 minutes → waste ~2.15 h/day → 91% goodput.
Cutting the interval 10× cuts expected recovery time 10× — but only if saves are cheap enough to afford frequent checkpointing, which is why async saves and distributed checkpoint matter.
The unifying principle¶
"Frequent, inexpensive, complete checkpoints turn a hardware failure from a job-ending event into a rounding error, and an overlapped input pipeline keeps the accelerators busy in between."
Cheap (async) saves make frequency affordable; complete saves (model + data + RNG, see concepts/durable-execution) make recovery correct; overlapped dataloading eliminates idle. With all in place, effective training time approaches the ceiling the hardware allows regardless of how flaky the cluster is.
Seen in¶
- sources/2026-08-28-databricks-fast-fault-tolerant-pytorch-training-on-ai-runtime — defining source; goodput as the master metric bounding GPU spend, decomposed into checkpoint-recovery cost and dataloading idle.
Related¶
- checkpoint-frequency — the primary goodput lever on the recovery side.
- distributed-checkpoint — the format that makes frequent saves cheap.
- gpu-stall-from-storage — the idle-side destroyer of goodput.
- gpu-training-failure-modes — why recovery is the expected case.
- concepts/model-flops-utilization — a related-but-different efficiency metric.