Skip to content

SYSTEM Cited by 1 source

NCCLX

Definition

NCCLX is Meta's extension to NVIDIA's NCCL collective communications library, used in GEM training to make pure data-movement collectives copy-free and SM-free — offloading data transfer from streaming multiprocessors (SMs) to dedicated hardware engines (Source: sources/2026-08-03-meta-gem-training-how-meta-doubled-the-efficiency-of-its-llm-scal).

The problem: SM contention

With 5D parallelism, GEM hides most communication behind compute via pipelining. But communication collectives themselves occupy SMs — e.g. ~24 SMs for all-gather / reduce-scatter — that would otherwise run compute kernels in parallel, costing up to 15% efficiency. Worse, compute-kernel performance can drop more than the SM-occupancy loss suggests, because wave scheduling wastes additional cycles.

SM-free communication

For pure data-movement collectives (e.g. all-gather), NCCLX moves data without SM involvement:

  • Copy Engine (CE) handles intra-node NVLink transfers.
  • RDMA handles inter-node transfers.

This reduces all-gather SM usage from 24 to 1, reclaiming ~23 SMs for compute and yielding ~5% E2E QPS gain at full training scale.

For collectives requiring reduction (e.g. all-reduce), Meta uses NVLink SHARP with in-network reduction, offloading the reduction computation from SMs to the network switch hardware.

See sm-free-communication for the general principle.

Seen in

  • systems/nccl — the NVIDIA library NCCLX extends.
  • sm-free-communication — the general offload principle.
  • 5d-parallelism — the parallelism scheme whose collectives NCCLX carries.
  • systems/meta-gem — the training workload.
  • companies/meta
Last updated · 766 distilled / 2,225 read