Skip to content

CONCEPT Cited by 1 source

Zero-shot forecasting

Zero-shot forecasting is producing a forecast for a time series without fitting any model to that series — a pre-trained time-series foundation model takes a historical window as input (via in-context learning) and returns a multi-step forecast, typically as probabilistic quantiles. It is the time-series analogue of a zero-shot LLM: the model generalizes to a series it has never seen, so onboarding a new series requires only data, not a training job.

Why it matters

The operational value is the removal of the per-series ML pipeline. Classical methods (ARIMA, Holt-Winters, seasonal decomposition) and modern ML/DL approaches (LightGBM, DeepAR, Temporal Fusion Transformer) all require per-series model fitting — for a 10,000-SKU catalog, that is 10,000 models to train, validate, tune, and retrain, each with its own cold-start problem for new items. The operational burden scales linearly with catalog size. A zero-shot foundation model collapses that to a single hosted endpoint: onboarding a new SKU drops from weeks of training to minutes of data upload (Source: sources/2026-09-11-aws-from-zero-shot-forecast-to-purchase-order-with-agentcore).

Conditional, not just univariate

A useful zero-shot forecaster is conditional: it accepts covariates, distinguished by whether their future values are known —

  • Past-only covariates — features known only for past periods.
  • Known (future) covariates — features whose future values are supplied for the forecast horizon (a scheduled promotion, a planned price change).

Because covariates are explicit inputs, the same model supports what-if analysis — forecast with vs without a promotion, at current vs discounted price — before committing to a decision.

Probabilistic output is the point

Zero-shot demand forecasters emit quantiles (P10/P50/P90), not a point estimate. The spread is what downstream planning consumes: ordering to the P50 alone stocks out ~half the time, so a separate safety-stock buffer absorbs the P50→P90 uncertainty, tuned toward a target service level. A high P90/P50 ratio flags volatile or event-driven demand and should be surfaced, not hidden behind a single number.

Trade-off vs per-series-trained models

Zero-shot trades away catalog-specific structure a trained model could learn, in exchange for eliminating the pipeline. Where accuracy on a specific catalog matters, the two can be benchmarked head-to-head on the same traces — see the per-SKU-trained ZEOS Demand Forecaster (LightGBM quantile regression, weekly retrain) as the trained-model counterpoint to zero-shot Chronos2. This is a concrete instance of the training/serving boundary question: zero-shot forecasting pushes all "training" into pre-training done by the model provider, leaving the consumer with a pure serving concern.

Seen in

Last updated · 766 distilled / 2,225 read