SYSTEM Cited by 1 source
NCCLX¶
Definition¶
NCCLX is Meta's extension to NVIDIA's NCCL collective communications library, used in GEM training to make pure data-movement collectives copy-free and SM-free — offloading data transfer from streaming multiprocessors (SMs) to dedicated hardware engines (Source: sources/2026-08-03-meta-gem-training-how-meta-doubled-the-efficiency-of-its-llm-scal).
The problem: SM contention¶
With 5D parallelism, GEM hides most communication behind compute via pipelining. But communication collectives themselves occupy SMs — e.g. ~24 SMs for all-gather / reduce-scatter — that would otherwise run compute kernels in parallel, costing up to 15% efficiency. Worse, compute-kernel performance can drop more than the SM-occupancy loss suggests, because wave scheduling wastes additional cycles.
SM-free communication¶
For pure data-movement collectives (e.g. all-gather), NCCLX moves data without SM involvement:
- Copy Engine (CE) handles intra-node NVLink transfers.
- RDMA handles inter-node transfers.
This reduces all-gather SM usage from 24 to 1, reclaiming ~23 SMs for compute and yielding ~5% E2E QPS gain at full training scale.
For collectives requiring reduction (e.g. all-reduce), Meta uses NVLink SHARP with in-network reduction, offloading the reduction computation from SMs to the network switch hardware.
See sm-free-communication for the general principle.
Seen in¶
- sources/2026-08-03-meta-gem-training-how-meta-doubled-the-efficiency-of-its-llm-scal — NCCLX SM-free collectives for GEM.
Related¶
- systems/nccl — the NVIDIA library NCCLX extends.
- sm-free-communication — the general offload principle.
- 5d-parallelism — the parallelism scheme whose collectives NCCLX carries.
- systems/meta-gem — the training workload.
- companies/meta