Skip to main content

KVCache over MoQT
draft-shi-moq-kvcache-01

Document Type Expired Internet-Draft (individual)
Expired & archived
Authors Hang Shi , Shengnan Yue
Last updated 2025-10-11 (Latest revision 2025-04-09)
RFC stream (None)
Intended RFC status (None)
Formats
Additional resources GitHub Repository
Stream Stream state (No stream defined)
Consensus boilerplate Unknown
RFC Editor Note (None)
IESG IESG state Expired
Telechat date (None)
Responsible AD (None)
Send notices to (None)

This Internet-Draft is no longer active. A copy of the expired Internet-Draft is available in these formats:

Abstract

Large language model (LLM) inference involves two stages: prefill and decode. The prefill phase processes the prompt in parallel, generating the KVCache, which is then used by the decode phase to produce tokens sequentially. KVCache can be reused if the model and prompt is the same, reducing computing cost of the prefill. However, its large size makes efficient transfer challenging. Delivering these over architectures enabled by publish/subscribe transport like MoQT, allows local nodes to cache the KVCache to be later retrieved via new subscriptions, saving the bandwidth. This document specifies the transmission of KVCache over MoQT.

Authors

Hang Shi
Shengnan Yue

(Note: The e-mail addresses provided for the authors of this Internet-Draft may no longer be valid.)