BIER Extension or New Protocol for AI P2MP
draft-mcbride-mcast4ai-bier-or-new-protocol-00
| Document | Type |
Replaced Internet-Draft
(individual)
Expired & archived
|
|
|---|---|---|---|
| Authors | Mike McBride , Yisong Liu , Li Zhang | ||
| Last updated | 2026-07-01 | ||
| Replaced by | draft-mcbride-mcast4ai-p2mp-mechanism-evaluation | ||
| RFC stream | (None) | ||
| Intended RFC status | (None) | ||
| Formats | |||
| Stream | Stream state | (No stream defined) | |
| Consensus boilerplate | Unknown | ||
| RFC Editor Note | (None) | ||
| IESG | IESG state | Replaced by draft-mcbride-mcast4ai-p2mp-mechanism-evaluation | |
| Telechat date | (None) | ||
| Responsible AD | (None) | ||
| Send notices to | (None) |
This Internet-Draft is no longer active. A copy of the expired Internet-Draft is available in these formats:
Abstract
AI workloads in data centers exhibit inherently point-to-multipoint (P2MP) communication patterns, particularly during collective operations such as AllReduce, AllGather and broadcast in distributed training. Unicast replication of these flows does not scale to large GPU clusters. This document analyzes two architectural approaches to addressing this problem: extending BIER (Bit Index Explicit Replication) to support AI P2MP requirements, or defining a new purpose-built protocol. The tradeoffs of each approach are discussed, including considerations around ACK aggregation, congestion control, RoCE/RDMA compatibility and operational complexity. This document does not define a protocol but is intended instead to help the mcast4ai community evaluate this problem space.
Authors
Mike McBride
Yisong Liu
Li Zhang
(Note: The e-mail addresses provided for the authors of this Internet-Draft may no longer be valid.)