Skip to content

SYSTEM Cited by 3 sources

FlashAttention

FlashAttention is a family of IO-aware GPU attention kernels (Dao et al., 2022+) that compute softmax-attention in tiles directly against shared memory, avoiding materialisation of the full N × N attention matrix in HBM. FlashAttention-family varlen kernels also provide padding-removal / variable-length attention so concurrent sequences of different lengths share a batch without padding tokens wasting GPU cycles.

This page is the canonical wiki entry under the slug flash-attention that source pages reference as [systems/flash-attention](<./flash-attention.md>). The sibling flashattention slug (systems/flashattention) covers the same system from the Netflix post-training-framework angle.

Use at Pinterest — baseline for DCAT

Pinterest's DCAT (Deduplicated Cross-Attention Transformer) achieves "significant throughput gains over standard self-attention with FlashAttention" (Source: sources/2026-04-13-pinterest-scaling-recommendation-systems-with-request-level-deduplication). FlashAttention is the performance baseline Pinterest compares DCAT against — standard self-attention with FlashAttention was the prior ranking-attention path before DCAT's two-phase context/crossing split replaced it.

Why DCAT beats FlashAttention for this workload: FlashAttention is IO-aware self-attention — it reduces HBM traffic per attention call, but still computes one call per candidate in a batch. DCAT changes the computation shape — it factors the user-sequence context pass out of the per-candidate path, so the cost asymmetry is 1 user-sequence pass per request + B cross-attention calls vs FlashAttention's B self-attention calls. For recsys ranking with shared user history, the context factor-out is a bigger win than any per-call IO optimisation.

Use at Netflix

First wiki mention: sources/2026-02-13-netflix-scaling-llm-post-training-at-netflix as part of Netflix's internal optimised model definitions — along with systems/flex-attention, memory-efficient chunked cross-entropy, consistent MFU accounting, and uniform LoRA extensibility.

Use at Meta — GEM training (baseline generalized + displaced)

Meta's GEM ads foundation model shows two ways standard FlashAttention is inadequate for recsys training and how Meta responds (Source: sources/2026-08-03-meta-gem-training-how-meta-doubled-the-efficiency-of-its-llm-scal):

  • Jagged inputs — FlashAttention assumes uniform sequence length; recsys user sequences are jagged (hundreds to tens of thousands of tokens). Jagged Flash Attention (JFA) generalizes FlashAttention to operate directly on variable-length tensors, avoiding the up-to-50% padding waste.
  • Non-softmax, asymmetric attention — GEM's self/cross/PMA attention replaces softmax with GELU/SiLU and has short/asymmetric K/V shapes where FlashAttention shows a 2.6× (up to 4× worst-case) gap to hardware roofline. GDPA unifies these and beats FlashAttention 4 by up to 3.5× forward under short-K/V production settings.

Meta also extended the FA4 kernel with end-to-end MXFP8 block-scaled MMA (>1.3× fwd, >1.5× bwd on GEM shapes) — FlashAttention here is the substrate Meta modifies for low precision, not just a baseline.

Seen in

Last updated · 766 distilled / 2,225 read