<?xml version="1.0" encoding="UTF-8"?>
<reference anchor="I-D.fu-nmop-tokenops-probelem-statement" target="https://datatracker.ietf.org/doc/html/draft-fu-nmop-tokenops-probelem-statement-00">
   <front>
      <title>Token Operation Problem Statement</title>
      <author initials="Y." surname="Fu" fullname="Yu Fu">
         <organization>China Telecom</organization>
      </author>
      <author initials="S." surname="Qiong" fullname="Sun Qiong">
         <organization>China Telecom</organization>
      </author>
      <author initials="X." surname="Song" fullname="Xin Song">
         <organization>China Telecom</organization>
      </author>
      <author initials="C." surname="Xie" fullname="Chongfeng Xie">
         <organization>China Telecom</organization>
      </author>
      <date month="July" day="6" year="2026" />
      <abstract>
	 <t>   Distributed LLM inference relies heavily on high-performance
   networking to synchronize states across accelerators (e.g., GPUs) and
   nodes.  Unlike traditional web services, inference workloads
   particularly those involving Mixture-of-Experts (MoE) models and
   long-context windows exhibit unique traffic patterns characterized by
   massive east-west traffic and strict latency constraints.  Current
   network infrastructures and scheduling methods often treat compute
   resources and network paths independently, leading to suboptimal
   performance and degraded Quality of Experience (QoE).  This document
   elaborates on these issues to guide potential protocol enhancements
   within the IETF.

	 </t>
      </abstract>
   </front>
   <seriesInfo name="Internet-Draft" value="draft-fu-nmop-tokenops-probelem-statement-00" />
   
</reference>
