<?xml version="1.0" encoding="UTF-8"?>
<reference anchor="I-D.calabria-bmwg-ai-fabric-inference-bench" target="https://datatracker.ietf.org/doc/html/draft-calabria-bmwg-ai-fabric-inference-bench-03">
   <front>
      <title>Benchmarking Methodology for AI Inference Serving Network Fabrics</title>
      <author initials="F." surname="Calabria" fullname="Fernando Calabria">
         <organization>Cisco</organization>
      </author>
      <author initials="C." surname="Pignataro" fullname="Carlos Pignataro">
         <organization>Blue Fern Consulting</organization>
      </author>
      <author initials="Q." surname="Wu" fullname="Qin Wu">
         <organization>Huawei</organization>
      </author>
      <author initials="G." surname="Fioccola" fullname="Giuseppe Fioccola">
         <organization>Huawei</organization>
      </author>
      <author initials="S." surname="Reddy" fullname="Sowjanya Reddy">
         <organization>Apple</organization>
      </author>
      <date month="July" day="6" year="2026" />
      <abstract>
	 <t>   This document defines benchmarking terminology, methodologies, and
   Key Performance Indicators (KPIs) for evaluating Ethernet-based AI
   inference serving network fabrics.  As Large Language Model (LLM)
   inference deployments scale to disaggregated prefill/decode
   architectures spanning hundreds or thousands of accelerators (GPUs/
   XPUs), the interconnect fabric determines Time to First Token (TTFT),
   Inter-Token Latency (ITL), and aggregate throughput in tokens per
   second (TPS).  This document establishes vendor-independent,
   reproducible test procedures for benchmarking fabric-level
   performance under realistic AI inference workloads.

   Coverage includes RDMA-based KV cache transfer between disaggregated
   prefill and decode workers, Mixture-of-Experts (MoE) expert
   parallelism AllToAll communication, request routing and load
   balancing for inference serving, congestion management under bursty
   inference traffic patterns, and scale/soak testing.  The methodology
   enables direct comparison across NIC transport stacks (RoCEv2 and
   UET) and fabric architectures.

   This document is a companion to the AI training fabric benchmarking
   methodology, which addresses training workloads.

	 </t>
      </abstract>
   </front>
   <seriesInfo name="Internet-Draft" value="draft-calabria-bmwg-ai-fabric-inference-bench-03" />
   
</reference>
