Skip to content

SYSTEM Cited by 1 source

Jagged Flash Attention (JFA)

Definition

Jagged Flash Attention (JFA) is Meta's custom FlashAttention implementation that operates directly on variable-length jagged tensors, eliminating the padding overhead that standard FlashAttention incurs on recommendation workloads where user sequences vary from hundreds to tens of thousands of tokens. Padding to max length would waste up to 50% of compute; JFA avoids it while supporting rec-specific features such as custom attention biases, asymmetric query/key-value lengths, and efficient backward passes (Source: sources/2026-08-03-meta-gem-training-how-meta-doubled-the-efficiency-of-its-llm-scal).

Why standard FlashAttention falls short

Standard FlashAttention assumes uniform sequence lengths for efficient tiling and parallelization. With jagged inputs, naive approaches either pad (wasting compute) or leave SMs idle when short sequences finish early. JFA is purpose-built for the jagged regime.

Four generations

Meta evolved JFA through four generations, progressively closing the gap from slower-than-padded-SDPA to matching SOTA CUDA/Cutlass performance:

  1. Jagged masking via subtraction scheme — traditional 2D masking for jagged boundaries (marking invalid positions with −inf) consumes ~28% of executed instructions on non-Tensor-Core units. JFA replaces it by masking Query/Key with zeros (which the Tensor Memory Accelerator does for free) and subtracting the extra exponents — numerically equivalent, no masking overhead.
  2. Backward parallelization — FlashAttention's backward accumulates dQ across sequence tiles, typically via costly atomic adds. For rec workloads with high batch × heads, a non-seq-parallel scheme with split dQ computation delivers 21–40% backward speedup by eliminating both atomic writes and redundant recompute.
  3. Warp specialization + persistent kernels — upgrading to Triton Low-Level Extensions (TLX) enables explicit warp specialization, TMA use, and persistent kernel scheduling — 30–100% TFLOPS improvement.

Results

JFA v4 (TLX) achieves 40–140% TFLOPs improvement over JFA v2 under production jagged distributions (sparsity 0.5), contributing +18.5% relative local MFU and +12% QPS (MFU). As part of the full custom-kernel library (with GDPA and BlockAttention), it contributes to >30% end-to-end training throughput on GEM.

Seen in

Last updated · 766 distilled / 2,225 read