Skip to content

SYSTEM Cited by 2 sources

GEM (Generative Ads Recommendation Model)

Definition

GEM is Meta's central recommendations foundation model behind ads recommendations across Instagram and Facebook, first introduced in November 2025 and by 2026 trained at LLM scale on several thousand latest-generation GPUs. It has a hybrid architecture combining trillions of sparse embedding parameters with billions of dense parameters, trained on ad content and user-engagement data (Source: sources/2026-08-03-meta-gem-training-how-meta-doubled-the-efficiency-of-its-llm-scal).

Architecture

GEM ingests two categories of features, with customized attention mechanisms applied to each group independently while enabling cross-feature learning:

  • Sequence features — e.g. user activity history (can be tens of thousands of tokens, highly variable length → jagged inputs).
  • Non-sequence features — e.g. user location, ad creative representation.

The interaction of this hybrid architecture with rec-domain data properties is what makes GEM's training uniquely challenging — distinct from LLM training:

  • Jagged inputs — padding to max sequence length would waste up to 50% of compute.
  • Diverse, asymmetric attention shapes — self-attention (long sequence, short window), cross-attention (long query, short K/V), pooled multi-head attention / PMA (short query, long K/V). These asymmetric shapes defeat intra-kernel pipelining.
  • Memory-bound operations — small embedding dimensions for MLP + many normalizations leave compute units underused.
  • Numerical sensitivity — CTR/CVR (click/conversion) objectives are highly sensitive to precision, so naive low-precision training risks quality regression.

Sequence-learning core: the multi-stage sequence model

GEM's sequence-feature processing is anchored by Meta's multi-stage sequence model (LLaTTE) (sources/2026-08-05-meta-from-user-sequences-to-scaling-laws-a-multi-stage-architecture-for-metas-ads-ranking|2026-08-05), described as "a core component of Meta's GEM." It decouples offline / upstream user modeling (deep transformers over thousand-token histories → cached, candidate-independent user embeddings) from online / downstream ranking, and learns feature interactions from data via dense tokenization + target-aware attention rather than manual sparse cross-features. On production ads traffic it exhibits an LLM-style scaling law (log-linear NE vs FLOPs) across depth, width, sequence length, and semantic enrichment (scaling synergy principle), and contributed a cumulative lift of +6% IG / +3% FB conversions, +3.5% FB ad clicks. This is the sequence-architecture layer of GEM; the GEM training post covers the training-efficiency (kernels + parallelism + precision) layer.

Efficiency outcome

Over 12 months Meta doubled end-to-end training efficiency to 20–25% MFU while scaling total training FLOPs 4× (MFU), decomposed via E2E MFU = Local MFU × Scaling Ratio.

Compute-efficiency levers (per-GPU roofline):

Scaling-efficiency levers (single-GPU → multi-GPU gap):

  • Topology-aware 5D parallelism — concepts/tensor-parallelism|2D FSDP + Expert Parallelism for dense params; Fully Sharded 2D Model Parallelism for sparse params.
  • SM-free collectives via NCCLX.
  • Compiler-based activation checkpointing (AutoAC) + activation quantization.
  • Base Batch Shuffling for the recsys straggler problem.

Relation to other Meta ads systems

  • Meta Adaptive Ranking Model (MARM) — the LLM-scale ads-ranking serving/inference stack. GEM is the training counterpart; both share MFU accounting, selective FP8, and sparse-embedding sharding vocabulary.
  • Wukong / Wukong Turbo — the ads-ranking model architecture family MARM serves.
  • Andromeda — the ads retrieval model whose kernels KernelEvolve optimizes.

Caveats

  • Model architecture given only at the O(trillions) sparse / O(billions) dense level; layer topology, expert count (DHEN), and exact split undisclosed.
  • "Several thousand" GPUs of unnamed "latest-generation" hardware; no vendor/model, fleet size, or wall-clock training time.

Seen in

Last updated · 766 distilled / 2,225 read