SYSTEM Cited by 1 source
Generalized Dot-Product Attention (GDPA)¶
Definition¶
GDPA is a single unified GPU kernel that accelerates GEM's diverse attention-like interaction patterns — self-attention, pooled multi-head attention (PMA), and cross-attention — which share a common structure (two matrix multiplications with an element-wise activation between them) but replace softmax with activations like GELU or SiLU. GDPA is optimized for production recsys training workloads on latest-generation GPUs (Source: sources/2026-08-03-meta-gem-training-how-meta-doubled-the-efficiency-of-its-llm-scal).
Why FlashAttention breaks down here¶
Existing FlashAttention kernels are designed for LLM-style dense, long-sequence inputs. On real GEM production traffic Meta observed a 2.6× forward performance gap (up to 4× worst-case) between real-world workloads and synthetic benchmarks, driven by short/asymmetric K/V sequences, jagged inputs, and large batch sizes that break pipeline-occupancy assumptions.
Three optimizations¶
- Pipeline redesign for non-softmax activations — eliminating the softmax correction stage frees four warps and their registers. For short K/V, outer-loop software pipelining recovers ~10% lost to inner-loop pipelining when the inner loop runs only 1–2 iterations.
- Software-level tile scheduling for jagged tensors — precompute valid tiles on CPU, skip empty tiles entirely, and apply zigzag assignment across SMs — reducing workload skew from 6× to near-balanced.
- ALU-only activation approximation — replace GELU's SFU-bound
tanhwith a 6th-order Taylor expansion (ALU-only), accurate within the bounded input range enforced by QK-norm, eliminating SFU contention in forward and backward.
Results¶
- 2× forward speedup — 1,145 BF16 TFLOPs, ~97% Tensor Core utilization.
- 1.6× backward speedup over baseline.
- Up to 3.5× forward speedup over FlashAttention 4 (FA4) under short-K/V production settings.
- Applied across the full model, GDPA + sibling kernels deliver >30% end-to-end training throughput.
Seen in¶
- sources/2026-08-03-meta-gem-training-how-meta-doubled-the-efficiency-of-its-llm-scal — GDPA's pipeline redesign and results.
Related¶
- systems/jagged-flash-attention — sibling GEM kernel for softmax attention on jagged inputs.
- systems/flash-attention — the LLM-oriented baseline GDPA beats on recsys shapes.
- systems/triton-lang — Meta's kernel authoring substrate (Triton / TLX).
- jagged-tensor — the variable-length inputs GDPA schedules around.
- systems/meta-gem — the model GDPA accelerates.