SYSTEM Cited by 1 source
MXFP8 (Micro-scaling FP8)¶
Definition¶
MXFP8 (micro-scaling FP8) is a block-scaled 8-bit floating-point format in which a shared scale factor is applied to a small block of values rather than to a whole tensor (per-tensor) or channel. Meta uses MXFP8 for both attention and MLP in GEM training to turn latest-generation GPUs' higher low-precision Tensor-Core throughput into real end-to-end speedups without regressing precision-sensitive CTR/CVR objectives (Source: sources/2026-08-03-meta-gem-training-how-meta-doubled-the-efficiency-of-its-llm-scal).
On the latest-generation GPU, FP8 delivers 2× peak FLOPS over FP16, and FP4 delivers 4× — and Meta expects low-precision peak FLOPS to grow faster than FP16 in future generations, making low-precision training increasingly attractive.
Low-precision Flash Attention (MXFP8 on FA4)¶
Meta extended the FA4 kernel with end-to-end MXFP8 block-scaled MMA for forward and backward. The challenge: it is "not just a datatype swap" — scale factors must be generated along each GEMM's K dimension, staged through SMEM/TMEM despite FA4's already-full TMEM footprint, and computed online for intermediates (softmax P, backward dS). Three kernel-level innovations:
- TMEM scale-factor placement — FA4 fully uses 512-column TMEM for accumulators. MXFP8 overlaps scale factors with temporarily-unused TMEM regions (e.g. placing S(i) scale factors in the S(1−i) accumulator region), needing only one lightweight barrier hidden behind existing GEMM latency.
- Online P-to-MXFP8 conversion — softmax output P is quantized to MXFP8 in
place within the softmax warp, reusing the row-max already computed for softmax
normalization; scale factors derived via optimized PTX bit-manipulation instead of
expensive
log2/round/clamp. - Block-wise quantization —
[32, 32]square quantization computes one scale factor per 32×32 block viaredux.sync.max.abs.f32warp-wide reduction, making quantization transpose-invariant so each tensor is quantized only once (useful for the backward pass, which needs transposed Q, K).
**Results on GEM representative shapes (power-capped latest-generation GPU):
1.3× forward, >1.5× backward.**
Neutralizing quantization overhead¶
Casting/scaling/data-movement can offset Tensor-Core speedups. Meta fuses it away:
- Weights — pre-all-gather shard quantization: quantize each rank's local FSDP shard before all-gather (amortizing cost across ranks) and communicate the low-precision payload to cut all-gather volume/latency.
- Activations — fuse quantization into the preceding kernel: PreNorm fusion for linear modules; fuse into the preceding projection for attention modules, so the attention kernel consumes low-precision activations directly.
Numerical stability¶
- Random Hadamard Transforms to spread outliers and smooth distributions before quantization.
- stochastic-rounding to eliminate deterministic rounding bias.
- Skip / higher-precision weight-gradient (WGrad) where activations/gradients exhibit severe outliers.
- Mixed precision — ultra-low precision on large GEMMs; BF16 fallback on later, more sensitive layers.
Seen in¶
- sources/2026-08-03-meta-gem-training-how-meta-doubled-the-efficiency-of-its-llm-scal — MXFP8 attention/MLP kernel and stability recipe.
Related¶
- selective-fp8-quantization — the layer-selective sibling strategy (MARM serving side).
- selective-fp8-quantization — the scaling-granularity axis MXFP8 sits on (block-level).
- random-hadamard-transform, stochastic-rounding — stability techniques.
- quantization-fused-into-preceding-kernel, pre-all-gather-shard-quantization — overhead-neutralization patterns.
- systems/meta-gem — the model MXFP8 trains.