Skip to content

SYSTEM Cited by 1 source

Meta Multi-Stage Sequence Model (LLaTTE)

Definition

Meta's multi-stage sequence model is the sequence-learning architecture behind Meta's ads ranking that decouples heavy offline / upstream user modeling from lightweight online / downstream ranking, and learns feature interactions directly from data via dense tokenization + target-aware multi-head attention. It is a core component of Meta's GEM foundation model and exhibits LLM-style scaling laws on production ads traffic. Companion paper: LLaTTE: Scaling Laws for Multi-Stage Sequence Modeling in Large-Scale Ads Recommendation (arXiv:2601.20083) (Source: sources/2026-08-05-meta-from-user-sequences-to-scaling-laws-a-multi-stage-architecture-for-metas-ads-ranking).

Architecture

Stage 1 — Offline (upstream) user model

  • Processes user-side features asynchronously with deep transformer models — several layers, sequence lengths in the thousands.
  • Produces user-level embeddings that are precomputed and cached (precomputed-user-embedding-cache).
  • Strictly separates user features from ad and context features so user embeddings stay independent of any particular ad candidate — the invariant that makes user-level caching valid.

Stage 2 — Online (downstream) ranking model

  • Combines the cached user embeddings with fresh real-time user signals and ad candidate information to produce the final ranking.
  • Optimized for speed under strict per-request latency budgets, while leveraging the deep representations computed offline.

Because heavy user modeling runs offline and only compact cached embeddings cross into the online path, model capacity can scale without a proportional increase in serving resources — the load-bearing decoupling (decouple-offline-user-modeling-from-online-ranking).

Learning-paradigm innovations

  • Dense tokenization — integrates sparse features + sequential behavioral data into a single dense vocabulary, so attention discovers cross-feature interactions from data rather than from hand-engineered sparse cross-features (dense-tokenization-learns-feature-interactions).
  • Target-aware multi-head attention — a memory-efficient attention that fuses tokenized sparse features + ad candidate info with the user behavior sequence and lets each layer weigh past behaviors against the specific ad being scored; stacked aligned blocks with stable attention distributions progressively distill long sequences into compact representations.

Scaling properties

  • On production ads traffic the model shows an LLM-style scaling law: performance (normalized entropy, NE) improves log-linearly with compute (FLOPs) across model depth, content/semantic enrichment, model width, and sequence length — extending scaling-laws-for-recommenders into Meta's ads-ranking domain.
  • Four scaling levers: balanced model shape (scaling synergy principle), multi-stage tunability (offline scales unbounded by latency; online gives steeper per-compute gains but is serving-bounded), sequence composition (diversity beats homogeneity), and semantic feature representation (helps cold-start).

Impact

  • Cumulative lift (with broader modeling innovations): +6% conversions on Instagram, +3% on Facebook, +3.5% ad clicks on Facebook.
  • Generalizes as a shared backbone: "the same multi-stage backbone and scaling properties can extend to any ads ranking task with minimal adaptation and overhead" — a core part of GEM.

Relation to other Meta ads systems

  • GEM — this multi-stage sequence model is a core component of Meta's Generative Ads Recommendation Model.
  • MARM — the LLM-scale ads-ranking serving stack; MARM's request-oriented computation sharing is the inference-side analogue of computing heavy user context once (here, offline).
  • Wukong — Meta's earlier ads-ranking feature- interaction architecture family.

Caveats

  • No exact layer counts, sequence-length numbers beyond "thousands", embedding dimensions, cache-refresh cadence, GPU count, or per-lever ablation magnitudes. Lift is cumulative with broader innovations. Deep internals deferred to the LLaTTE paper.

Seen in

Last updated · 766 distilled / 2,225 read