Meta — MetaRoCE: A New RDMA Transport Built for AI-Scale Ethernet¶
Summary¶
Meta Engineering's 2026-08-24 Networking post introduces MetaRoCE, a clean-sheet RDMA transport protocol purpose-built for AI workloads on commodity Ethernet at million-GPU scale. It is the transport-layer sequel to Meta's 2024 SIGCOMM RoCE paper and inverts several of that paper's foundational choices. Where standard RoCE expects the fabric to deliver every frame in order — leaning on PFC and discouraging the packet spraying that gives multiplane fabrics their performance — MetaRoCE's core insight is that "the fabric sees packets, but the NIC sees intent." It moves intelligence to the endpoint (endpoint-over-fabric intelligence), decomposes each connection into many fine-grained logical paths each with its own real-time telemetry (per-path RTT, ECN state, utilization), and treats out-of-order arrival and packet loss as the normal case rather than as exceptions. Every packet carries its own destination, so data is written straight to its final memory location as it lands — no reorder buffer, no head-of-line blocking. Meta is releasing the specification, a reference software implementation (libsoftmetaroce), and a compliance test suite through the Open Compute Project (OCP), extending the same open, multi-vendor philosophy that OCP's ESUN initiative established for the fabric into the transport layer. Validated with AMD on Pensando programmable NICs: on a 64-node AMD GPU cluster running RCCL collectives, MetaRoCE beats RoCEv2 on all-reduce and all-to-all throughput and flow-completion time, holds ~86% throughput at 1% packet loss (where RoCEv2 degrades), keeps delivering useful bandwidth at 10% loss, and scales linearly across 4- and 8-plane topologies with up to 4,000 concurrent connections.
Key takeaways¶
-
The one-sentence thesis: "The fabric sees packets, but the NIC sees intent." Traditional architectures centralize intelligence in the fabric, relying on switches to enforce losslessness and maintain order. MetaRoCE moves that intelligence to the endpoint NIC (endpoint-vs-fabric-intelligence), which decomposes the network into many fine-grained logical paths, each with its own real-time telemetry — per-path RTT, ECN state, and utilization. This visibility "unlocks capabilities that are difficult to achieve with traditional RDMA." Canonical wiki instance of endpoint-intelligence-over-fabric-intelligence.
-
Native out-of-order delivery — no reorder buffer, no head-of-line blocking. MetaRoCE sprays packets across many paths, so they arrive out of order by design, and "the transport treats out-of-order arrival as the normal case." "Every packet carries its own destination, so data is written straight to its final memory location as it lands, with no reorder buffer and no head-of-line blocking." Writes carry their destination in every packet; Sends carry the match to a posted receive buffer, so a Send lands correctly even when the messages ahead of it have not arrived, and without a round trip to learn where the data goes. Collective libraries can use two-sided messaging where it suits them rather than reducing everything to Write. (native-out-of-order-delivery)
-
Native multipathing with NIC-owned path selection. Each connection gets first-class paths and the NIC sprays across them packet by packet. Each path carries a distinct UDP source port as its ECMP entropy, which the NIC can change at any time to move traffic off a bad route. On multiplane fabrics plane selection falls entirely to the NIC, and "the fabric is used only as well as the NIC sprays." Because each path keeps its own window and round-trip estimate, the transport can tell congestion from failure and rebalance explicitly — a hot or broken link slows one path instead of stalling the connection. (packet-spraying, per-path-multipath-transport)
-
Loss tolerance by design — no PFC, no pause frames. MetaRoCE "treats the Ethernet fabric as lossy and does not ask it to be otherwise — no PFC, no pause frames" (lossy-fabric-by-design). Because each path carries its own ordered sequence, a gap in its 256-bit selective-acknowledgment bitvector is evidence of loss rather than of reordering. Where a conventional SACK "mostly avoids resending data that already arrived," here it triggers retransmission of exactly the missing packet, on the path that lost it, the moment the gap appears. (selective-acknowledgment) This is the deliberate reversal of the PFC-based lossless posture of the 2024 SIGCOMM paper.
-
Two-sided congestion control: sender-driven AIMD + receiver-driven fair-share hints. MetaRoCE combines a conventional ECN-based, sender-driven AIMD congestion control with receiver-driven fair-share rate hints. Windows are kept per path as well as per connection (per-path-congestion-control), so a congestion mark trims the path that saw it and steers the next packets toward clear paths. In every acknowledgment the receiver returns the share of its inbound bandwidth it has allocated to that sender, so senders approach the right speed directly rather than searching for it (receiver-driven-rate-hint). Result: incast resolves in one or two round trips, with better fairness and lower tail latency.
-
Topology independence — asks the fabric only for ECN + ECMP. MetaRoCE requires only two things "every switch already has": ECN marking and ECMP. It does not require packet trimming, in-network telemetry, credit-based flow control, or switch-side spraying, and it does not break when a fabric offers them. The same transport runs over fat-tree, multiplane, deep-buffer, and shallow-buffer fabrics, and over vendor clouds whose configuration you don't control. "Nothing proprietary is involved, so the fabric stays free to optimize for cost and cabling." (topology-independent-transport)
-
Unified connections decouple ordered streams from bandwidth paths. A queue pair traditionally carries both an ordered stream of messages and bandwidth, so to get more of either you open more QPs (dozens per node pair), each with a congestion window blind to the rest and its own state on the NIC. MetaRoCE separates the two: a single connection carries many independent ordered streams above (one per communicator or collective) and many paths below, under one congestion controller. The consequence is that connection state stops growing with the parallelism of the workload — directly attacking the QP-scaling / NIC-state-explosion problem that E-ECMP + QP-scaling worked around in 2024. Existing RDMA Verbs APIs and software stacks work without modification; features like multiplane support are exposed via extension APIs.
-
Validated on AMD Pensando; degrades gracefully instead of collapsing. On a 64-node AMD GPU cluster running RCCL collectives, MetaRoCE beat RoCEv2 on all-reduce and all-to-all: higher throughput and lower flow-completion times. Under loss that would degrade RoCEv2, MetaRoCE maintains ~86% throughput at 1% packet loss and continues delivering useful bandwidth at an extreme 10% loss rate — "converging gracefully rather than collapsing." Multiplane validation across 4-plane and 8-plane topologies with up to 4,000 concurrent connections confirmed throughput scales linearly with plane count; during simulated plane failures, traffic redistributes without application involvement or operator intervention. (design-for-loss-degrade-gracefully)
-
Open by design — spec + reference impl + compliance suite via OCP. Meta is opening MetaRoCE the same way OCP's Ethernet Scalable Unified Network (ESUN) opened the fabric layer: (a) open specification via OCP (any vendor can implement interoperable hardware); (b) multiple NIC implementations — designed to run across programmable and fixed-function NICs, proven on AMD Pensando, more underway; (c) a production compliance suite so vendors can prove implementations match the spec; (d) libsoftmetaroce — a complete functional transport stack running on commodity Linux over standard UDP sockets with no specialized hardware, serving as the authoritative behavioral model for silicon development and the foundation of the compliance framework. Full spec + DPDK-optimized reference impl + compliance framework ship at the 2026 OCP Global Summit in October. (open-spec-plus-compliance-suite-for-multivendor-hardware)
Operational numbers¶
- ~86% throughput at 1% packet loss (RoCEv2 degrades here); useful bandwidth even at 10% loss — graceful convergence, not collapse.
- 256-bit selective-acknowledgment bitvector per path — a gap = loss (not reorder), triggering targeted retransmit on the losing path.
- 64-node AMD GPU cluster, RCCL collectives (all-reduce, all-to-all); MetaRoCE > RoCEv2 on throughput + flow-completion time.
- 4-plane and 8-plane multiplane topologies validated; up to 4,000 concurrent connections; throughput scales linearly with plane count.
- Incast resolves in 1–2 round trips via receiver-driven fair-share hints.
- Fabric requirements reduced to exactly two: ECN marking + ECMP.
- Target scale: million-GPU AI clusters spanning multiple data centers and regions.
The inversion vs. Meta's 2024 SIGCOMM RoCE stack¶
MetaRoCE is best read against Meta's own sources/2024-08-05-meta-a-roce-network-for-distributed-ai-training-at-scale|2024-08-05 RoCE paper. Nearly every axis is flipped, because the 2024 design leaned on the fabric being lossless/ordered while MetaRoCE pushes everything to the endpoint:
| Axis | 2024 SIGCOMM RoCE stack | 2026 MetaRoCE |
|---|---|---|
| Ordering | Fabric delivers in order; reorder avoided | Out-of-order by design; per-packet destination, no reorder buffer |
| Loss posture | Lossless via PFC | Lossy by design — no PFC, no pause frames |
| Multipath | Fat-flow problem; E-ECMP + QP scaling | Native per-packet spraying, NIC-owned path selection |
| Congestion control | DCQCN OFF; PFC + NCCL CTS admission | ECN AIMD + receiver fair-share hints, per-path windows |
| Loss detection | Avoid loss entirely | 256-bit SACK bitvector → targeted per-path retransmit |
| Connection model | Many QPs per node pair (state grows with parallelism) | One connection: many ordered streams above, many paths below; state bounded |
| Fabric dependency | Deep-buffer spine, tuned routing | Topology-independent — needs only ECN + ECMP |
| Intelligence locus | Fabric (switches enforce order/losslessness) | Endpoint (NIC sees intent) |
Note the receiver-driven idea persists across both eras but changes layer: in 2024 it was NCCL clear-to-send admission at the collective-library layer (receiver-driven-traffic-admission); in 2026 it is a receiver fair-share rate hint carried in every ACK at the transport layer (receiver-driven-rate-hint).
The road ahead (three latency regimes)¶
Meta frames future work by distance/latency regime:
- Scale-up (within a rack, small messages, every nanosecond matters): MetaRoCE already removes the two main latency sources — the reorder buffer and PFC; work continues on the fast signaling path for short memory operations issued processing-element to processing-element.
- Scale-across (a single job spanning buildings thousands of km apart; round-trips in milliseconds): treating paths as first-class is what lets MetaRoCE prefer uncongested paths and seek fairness at every level; open work is fairly sharing contended long-haul links.
- Storage / KV-cache (a single read fans out to many servers → incast): receiver-driven rate hints let whichever side is receiving moderate the inbound rate whether the request hit ten servers or a thousand; the new dimension is keeping that rate accurate across networks of varying speed and requests of varying size.
Caveats¶
Architecture-and-results voice. No absolute throughput/latency numbers beyond the relative RoCEv2 comparison and the loss-tolerance percentages; the 64-node AMD validation is a pre-production hardware bring-up, not a fleet deployment — MetaRoCE is not reported as running Meta production training traffic yet. The spec, DPDK reference impl, and compliance suite are promised for the October 2026 OCP Global Summit (this post pre-announces them). No disclosure of the exact AIMD parameters, the fair-share allocation algorithm, path-count per connection, the reorder-free Send-matching wire format, or how MetaRoCE composes with NCCL/RCCL admission control now that admission has moved into the transport.
Source¶
- Original: https://engineering.fb.com/2026/08/24/networking-traffic/metaroce-rdma-transport-ai-ethernet/
- Raw markdown:
raw/meta/2026-08-24-metaroce-a-new-rdma-transport-built-for-ai-scale-ethernet-fb8e515a.md
Related¶
- systems/metaroce — the transport protocol.
- systems/libsoftmetaroce — the software reference implementation.
- systems/amd-pensando-nic — the programmable NIC used for hardware validation.
- systems/roce-rdma-over-converged-ethernet — the standard MetaRoCE supersedes.
- endpoint-vs-fabric-intelligence · native-out-of-order-delivery · packet-spraying · per-path-congestion-control · selective-acknowledgment · receiver-driven-rate-hint · lossy-fabric-by-design · topology-independent-transport · incast
- endpoint-intelligence-over-fabric-intelligence · per-path-multipath-transport · design-for-loss-degrade-gracefully · open-spec-plus-compliance-suite-for-multivendor-hardware
- companies/meta