Skip to content

SYSTEM Cited by 1 source

MetaRoCE

MetaRoCE is a clean-sheet RDMA transport protocol designed by Meta for AI workloads on commodity Ethernet at million-GPU scale, introduced 2026-08-24 and being contributed as an open specification, reference software implementation, and compliance suite through the Open Compute Project (OCP). It is the transport-layer successor to Meta's 2024 RoCE work and deliberately inverts the lossless/in-order fabric assumptions of standard RoCE.

Its organizing insight: "The fabric sees packets, but the NIC sees intent." Rather than centralizing intelligence in switches (enforcing losslessness, maintaining order), MetaRoCE pushes intelligence to the endpoint NIC, which decomposes each connection into many fine-grained logical paths, each with its own real-time telemetry — per-path RTT, ECN state, and utilization — and treats out-of-order arrival and packet loss as the normal case.

Core design pillars

  1. Native out-of-order delivery. Packets are sprayed across many paths and arrive out of order by design. Every packet carries its own destination, so data is written straight to its final memory location on arrival — no reorder buffer, no head-of-line blocking. Sends carry the match to a posted receive buffer, so a Send lands correctly even when prior messages have not arrived and without a round trip to learn placement.

  2. Native multipathing. Each connection gets first-class paths; the NIC sprays across them packet by packet. Each path uses a distinct UDP source port as its ECMP entropy, changeable at any time to route around a bad link. On multiplane fabrics plane selection is entirely the NIC's; each path keeps its own window and RTT estimate, letting the transport distinguish congestion from failure and rebalance explicitly.

  3. Loss tolerance by design. Treats Ethernet as lossy — no PFC, no pause frames. Because each path carries its own ordered sequence, a gap in its 256-bit selective-acknowledgment bitvector is evidence of loss (not reordering), triggering retransmission of exactly the missing packet, on the path that lost it, the moment the gap appears.

  4. Congestion control from both sides. Conventional ECN-based sender-driven AIMD combined with receiver-driven fair-share rate hints. Windows are kept per path as well as per connection, so a congestion mark trims the offending path and steers packets toward clear ones. Every ACK returns the receiver's allocated share of inbound bandwidth, so senders converge directly — incast resolves in one or two round trips.

  5. Topology independence. Asks the fabric for only two things every switch already has — ECN marking and ECMP. No packet trimming, in-network telemetry, credit-based flow control, or switch-side spraying required; runs over fat-tree, multiplane, deep-buffer, and shallow-buffer fabrics and over vendor clouds you don't control.

  6. Unified connections at scale. A traditional queue pair bundles an ordered stream and bandwidth, so scaling either means opening more QPs (dozens per node pair), each with independent CC state on the NIC. MetaRoCE separates them: one connection carries many independent ordered streams above and many paths below, under one congestion controller — so connection state stops growing with workload parallelism. Existing RDMA Verbs APIs and stacks work unmodified; extensions (e.g. multiplane) are exposed via extension APIs.

In practice (AMD Pensando validation)

To accelerate hardware validation Meta implemented MetaRoCE on AMD Pensando programmable NICs. On a 64-node AMD GPU cluster running RCCL collectives, MetaRoCE vs RoCEv2 across all-reduce and all-to-all delivered higher throughput and lower flow-completion times, held ~86% throughput at 1% packet loss (where RoCEv2 degrades) and useful bandwidth even at 10% loss — converging gracefully rather than collapsing. Multiplane validation across 4- and 8-plane topologies with up to 4,000 concurrent connections confirmed linear throughput scaling with plane count; simulated plane failures redistributed traffic autonomously, with no application or operator intervention. (design-for-loss-degrade-gracefully)

Open by design

MetaRoCE extends OCP's Ethernet Scalable Unified Network (ESUN) open, multi-vendor philosophy from the fabric into the transport layer:

  • Open specification via OCP — any vendor can implement interoperable hardware.
  • Multiple NIC implementations — designed for programmable and fixed-function NICs alike; proven on AMD Pensando, more underway.
  • Production compliance suite — vendors prove implementations match the spec.
  • libsoftmetaroce — a complete functional transport stack on commodity Linux over standard UDP sockets, the authoritative behavioral model for silicon development and the foundation of the compliance framework.

The full spec, a DPDK-optimized reference implementation, and the compliance framework ship at the 2026 OCP Global Summit (October). (open-spec-plus-compliance-suite-for-multivendor-hardware)

Relationship to Meta's 2024 RoCE stack

MetaRoCE inverts nearly every axis of the sources/2024-08-05-meta-a-roce-network-for-distributed-ai-training-at-scale|2024 SIGCOMM RoCE design, because that design leaned on the fabric being lossless/ordered while MetaRoCE pushes everything to the endpoint: out-of-order-by-design vs reorder-avoided; lossy vs PFC-lossless; native per-packet spraying vs the fat-flow + E-ECMP/QP-scaling workaround; one connection vs many QPs. See the source page for the full comparison table.

Caveats

Pre-production as of the post: validated on a 64-node AMD bring-up, not reported as running Meta production training traffic; spec + reference impl + compliance suite are promised for October 2026. AIMD parameters, the fair-share allocation algorithm, per-connection path count, and the reorder-free Send-matching wire format are not disclosed.

Seen in

  • systems/roce-rdma-over-converged-ethernet — the standard MetaRoCE supersedes.
  • systems/infiniband — the HPC-native alternative RoCE/MetaRoCE compete with on Ethernet.
  • systems/libsoftmetaroce — the software reference implementation.
  • systems/amd-pensando-nic — the programmable NIC used for validation.
  • systems/meta-genai-cluster-roce — Meta's prior RoCE training deployment.
  • endpoint-vs-fabric-intelligence · native-out-of-order-delivery · packet-spraying · per-path-congestion-control · selective-acknowledgment · receiver-driven-rate-hint · lossy-fabric-by-design · topology-independent-transport · incast
  • endpoint-intelligence-over-fabric-intelligence · per-path-multipath-transport · design-for-loss-degrade-gracefully · open-spec-plus-compliance-suite-for-multivendor-hardware
Last updated · 766 distilled / 2,225 read