Skip to content

SYSTEM Cited by 1 source

MXFP8 (Micro-scaling FP8)

Definition

MXFP8 (micro-scaling FP8) is a block-scaled 8-bit floating-point format in which a shared scale factor is applied to a small block of values rather than to a whole tensor (per-tensor) or channel. Meta uses MXFP8 for both attention and MLP in GEM training to turn latest-generation GPUs' higher low-precision Tensor-Core throughput into real end-to-end speedups without regressing precision-sensitive CTR/CVR objectives (Source: sources/2026-08-03-meta-gem-training-how-meta-doubled-the-efficiency-of-its-llm-scal).

On the latest-generation GPU, FP8 delivers 2× peak FLOPS over FP16, and FP4 delivers 4× — and Meta expects low-precision peak FLOPS to grow faster than FP16 in future generations, making low-precision training increasingly attractive.

Low-precision Flash Attention (MXFP8 on FA4)

Meta extended the FA4 kernel with end-to-end MXFP8 block-scaled MMA for forward and backward. The challenge: it is "not just a datatype swap" — scale factors must be generated along each GEMM's K dimension, staged through SMEM/TMEM despite FA4's already-full TMEM footprint, and computed online for intermediates (softmax P, backward dS). Three kernel-level innovations:

  1. TMEM scale-factor placement — FA4 fully uses 512-column TMEM for accumulators. MXFP8 overlaps scale factors with temporarily-unused TMEM regions (e.g. placing S(i) scale factors in the S(1−i) accumulator region), needing only one lightweight barrier hidden behind existing GEMM latency.
  2. Online P-to-MXFP8 conversion — softmax output P is quantized to MXFP8 in place within the softmax warp, reusing the row-max already computed for softmax normalization; scale factors derived via optimized PTX bit-manipulation instead of expensive log2/round/clamp.
  3. Block-wise quantization — [32, 32] square quantization computes one scale factor per 32×32 block via redux.sync.max.abs.f32 warp-wide reduction, making quantization transpose-invariant so each tensor is quantized only once (useful for the backward pass, which needs transposed Q, K).

**Results on GEM representative shapes (power-capped latest-generation GPU):

1.3× forward, >1.5× backward.**

Neutralizing quantization overhead

Casting/scaling/data-movement can offset Tensor-Core speedups. Meta fuses it away:

  • Weights — pre-all-gather shard quantization: quantize each rank's local FSDP shard before all-gather (amortizing cost across ranks) and communicate the low-precision payload to cut all-gather volume/latency.
  • Activations — fuse quantization into the preceding kernel: PreNorm fusion for linear modules; fuse into the preceding projection for attention modules, so the attention kernel consumes low-precision activations directly.

Numerical stability

  • Random Hadamard Transforms to spread outliers and smooth distributions before quantization.
  • stochastic-rounding to eliminate deterministic rounding bias.
  • Skip / higher-precision weight-gradient (WGrad) where activations/gradients exhibit severe outliers.
  • Mixed precision — ultra-low precision on large GEMMs; BF16 fallback on later, more sensitive layers.

Seen in

  • selective-fp8-quantization — the layer-selective sibling strategy (MARM serving side).
  • selective-fp8-quantization — the scaling-granularity axis MXFP8 sits on (block-level).
  • random-hadamard-transform, stochastic-rounding — stability techniques.
  • quantization-fused-into-preceding-kernel, pre-all-gather-shard-quantization — overhead-neutralization patterns.
  • systems/meta-gem — the model MXFP8 trains.
Last updated · 766 distilled / 2,225 read