Skip to content

SYSTEM Cited by 1 source

Generalized Dot-Product Attention (GDPA)

Definition

GDPA is a single unified GPU kernel that accelerates GEM's diverse attention-like interaction patterns — self-attention, pooled multi-head attention (PMA), and cross-attention — which share a common structure (two matrix multiplications with an element-wise activation between them) but replace softmax with activations like GELU or SiLU. GDPA is optimized for production recsys training workloads on latest-generation GPUs (Source: sources/2026-08-03-meta-gem-training-how-meta-doubled-the-efficiency-of-its-llm-scal).

Why FlashAttention breaks down here

Existing FlashAttention kernels are designed for LLM-style dense, long-sequence inputs. On real GEM production traffic Meta observed a 2.6× forward performance gap (up to 4× worst-case) between real-world workloads and synthetic benchmarks, driven by short/asymmetric K/V sequences, jagged inputs, and large batch sizes that break pipeline-occupancy assumptions.

Three optimizations

  1. Pipeline redesign for non-softmax activations — eliminating the softmax correction stage frees four warps and their registers. For short K/V, outer-loop software pipelining recovers ~10% lost to inner-loop pipelining when the inner loop runs only 1–2 iterations.
  2. Software-level tile scheduling for jagged tensors — precompute valid tiles on CPU, skip empty tiles entirely, and apply zigzag assignment across SMs — reducing workload skew from 6× to near-balanced.
  3. ALU-only activation approximation — replace GELU's SFU-bound tanh with a 6th-order Taylor expansion (ALU-only), accurate within the bounded input range enforced by QK-norm, eliminating SFU contention in forward and backward.

Results

  • 2× forward speedup — 1,145 BF16 TFLOPs, ~97% Tensor Core utilization.
  • 1.6× backward speedup over baseline.
  • Up to 3.5× forward speedup over FlashAttention 4 (FA4) under short-K/V production settings.
  • Applied across the full model, GDPA + sibling kernels deliver >30% end-to-end training throughput.

Seen in

Last updated · 766 distilled / 2,225 read