Skip to main content

State, Storage, and Compute Affinity for AI Agent Service Selection in Computing-Aware Traffic Steering
draft-mo-cats-agent-state-affinity-00

Document Type Active Internet-Draft (individual)
Authors Yijun Mo , Dewei Yang , 周诚宇
Last updated 2026-10-05
RFC stream (None)
Intended RFC status (None)
Formats
Stream Stream state (No stream defined)
Consensus boilerplate Unknown
RFC Editor Note (None)
IESG IESG state I-D Exists
Telechat date (None)
Responsible AD (None)
Send notices to (None)
draft-mo-cats-agent-state-affinity-00
CATS                                                               Y. Mo
Internet-Draft                                                   D. Yang
Intended status: Informational                                   C. Zhou
Expires: 6 April 2027      Huazhong University of Science and Technology
                                                       06 October 2026

       State, Storage, and Compute Affinity for AI Agent Service
             Selection in Computing-Aware Traffic Steering

                 draft-mo-cats-agent-state-affinity-00

Abstract

   AI agent services are stateful and long-running: the speed with which
   a step can be served depends on whether the session's context,
   retrieved data, tool results, and model-side state such as a
   key-value (KV) cache are already available at, or near, the selected
   service contact instance, and on whether the computation that the
   step needs is ready there. Computing-Aware Traffic Steering (CATS)
   exposes computing and network metrics and defines service contact
   instance affinity, but it does not expose the availability of
   reusable state, does not distinguish a hard locality constraint on
   state from a preference for reusing it, does not provide a way to
   compare the cost of moving state with the cost of recomputing it, and
   does not represent the readiness of a specific computation. This
   document describes affinity in three coupled dimensions -- state,
   storage, and compute -- for agent service selection. It states the
   motivation and the goals of introducing affinity, the requirements
   and constraints that affinity places on the mapping of an agent
   workload onto storage and compute resources, the metrics and
   measurement methods by which affinity is observed, the mechanisms and
   the procedure by which affinity is assured, and two cases in detail:
   a long-horizon session that passes through several stages, and a
   group of similar agents or of tenants served by shared reusable
   state. This document defines no wire protocol, no encoding, and no
   data model, and it does not define the transfer mechanisms
   themselves.

Status of This Memo

   This Internet-Draft is submitted in full conformance with the
   provisions of BCP 78 and BCP 79.

   Internet-Drafts are working documents of the Internet Engineering
   Task Force (IETF).  Note that other groups may also distribute
   working documents as Internet-Drafts.  The list of current Internet-
   Drafts is at https://datatracker.ietf.org/drafts/current/.

Mo, et al.               Expires 3 April 2027                   [Page 1]
Internet-Draft           Agent Affinity                   September 2026

   Internet-Drafts are draft documents valid for a maximum of six months
   and may be updated, replaced, or obsoleted by other documents at any
   time.  It is inappropriate to use Internet-Drafts as reference
   material or to cite them other than as "work in progress."

   This Internet-Draft will expire on 3 April 2027.

Copyright Notice

   Copyright (c) 2026 IETF Trust and the persons identified as the
   document authors.  All rights reserved.

   This document is subject to BCP 78 and the IETF Trust's Legal
   Provisions Relating to IETF Documents (https://trustee.ietf.org/
   license-info) in effect on the date of publication of this document.
   Please review these documents carefully, as they describe your rights
   and restrictions with respect to this document.  Code Components
   extracted from this document must include Revised BSD License text as
   described in Section 4.e of the Trust Legal Provisions and are
   provided without warranty as described in the Revised BSD License.

Discussion Venues

   Discussion of this document takes place on the CATS Working Group
   mailing list (cats@ietf.org), which is archived at
   https://mailarchive.ietf.org/arch/browse/cats/.

Table of Contents

Mo, et al.               Expires 3 April 2027                   [Page 2]
Internet-Draft           Agent Affinity                   September 2026

   1.  Introduction  . . . . . . . . . . . . . . . . . . . . . . . . . 4
   2.  Conventions and Definitions  . . . . . . . . . . . . . . . . . .6
   3.  Terminology  . . . . . . . . . . . . . . . . . . . . . . . . . .6
   4.  Motivation and Goals  . . . . . . . . . . . . . . . . . . . . . 8
      4.1.  Why Agent State, Storage, and Compute Are Coupled  . . . . 8
      4.2.  Failure Modes without an Affinity View  . . . . . . . . . .9
      4.3.  Goals  . . . . . . . . . . . . . . . . . . . . . . . . . .10
      4.4.  Non-Goals  . . . . . . . . . . . . . . . . . . . . . . . .11
   5.  Problem Statement  . . . . . . . . . . . . . . . . . . . . . . 11
      5.1.  Storage Is Represented Only as Capacity  . . . . . . . . .11
      5.2.  Affinity Is Instance-Level  . . . . . . . . . . . . . . . 11
      5.3.  Transfer and Re-computation Are Not Comparable  . . . . . .11
      5.4.  Compute Affinity Is Not Represented  . . . . . . . . . . .11
      5.5.  Reuse across Similar Agents and Tenants Is Not
         Represented  . . . . . . . . . . . . . . . . . . . . . . . . 12
      5.6.  The Long Horizon Amplifies Each of These Gaps  . . . . . .12
   6.  Working Set Elements Relevant to Agent Services  . . . . . . . 13
      6.1.  Element Classes  . . . . . . . . . . . . . . . . . . . . .13
      6.2.  Compute-Side Readiness  . . . . . . . . . . . . . . . . . 14
      6.3.  Why the Properties Are Pairwise  . . . . . . . . . . . . .14
   7.  State, Storage, and Compute Affinity  . . . . . . . . . . . . .14
      7.1.  Affinity, Preference, and Locality  . . . . . . . . . . . 14
      7.2.  How the Three Affinities Interact  . . . . . . . . . . . .15
      7.3.  The Granularity of an Affinity Decision  . . . . . . . . .16
   8.  Requirements and Constraints on the Mapping  . . . . . . . . . 17
      8.1.  The Mapping Relation  . . . . . . . . . . . . . . . . . . 17
      8.2.  Constraint Families  . . . . . . . . . . . . . . . . . . .17
      8.3.  Mapping Requirements  . . . . . . . . . . . . . . . . . . 18
   9.  Affinity Information Requirements  . . . . . . . . . . . . . . 20
      9.1.  State Availability  . . . . . . . . . . . . . . . . . . . 20
      9.2.  Residency Tiers and Retrieval Cost  . . . . . . . . . . . 20
      9.3.  Constraints, Consistency, and Sharing  . . . . . . . . . .20
      9.4.  Compute Affinity and Reuse Scope  . . . . . . . . . . . . 21
   10.  State Handles and Their Relationship to CATS Identifiers  . . 21
   11.  Measuring Affinity: Metrics and Methods  . . . . . . . . . . .22
      11.1.  What Has to Be Measured  . . . . . . . . . . . . . . . . 23
      11.2.  Affinity Metric Catalogue  . . . . . . . . . . . . . . . 23
      11.3.  Metric Semantics and Reporting Standards  . . . . . . . .24
      11.4.  Measurement Methods  . . . . . . . . . . . . . . . . . . 25
      11.5.  Measurement Requirements  . . . . . . . . . . . . . . . .26
   12.  Affinity Assurance: Mechanisms and Procedures  . . . . . . . .27
      12.1.  Mechanism Catalogue  . . . . . . . . . . . . . . . . . . 27
      12.2.  Assurance Procedure  . . . . . . . . . . . . . . . . . . 28
      12.3.  Triggers for Re-evaluation  . . . . . . . . . . . . . . .29
      12.4.  Abort, Failure, and Fallback  . . . . . . . . . . . . . .30
      12.5.  Assurance Requirements  . . . . . . . . . . . . . . . . .31
   13.  Affinity Assurance for Long-Horizon and Multi-Stage
      Sessions  . . . . . . . . . . . . . . . . . . . . . . . . . . . 32
      13.1.  Stage Model  . . . . . . . . . . . . . . . . . . . . . . 32
      13.2.  Phase-Differentiated Assurance  . . . . . . . . . . . . .32
      13.3.  Pinning, Retention, and Tier Budget  . . . . . . . . . . 33
Mo, et al.               Expires 3 April 2027                   [Page 3]
Internet-Draft           Agent Affinity                   September 2026

      13.4.  Checkpoint and Resume after a Pause  . . . . . . . . . . 33
      13.5.  Stage Transitions  . . . . . . . . . . . . . . . . . . . 34
      13.6.  Drift, Decay, and Re-evaluation  . . . . . . . . . . . . 34
      13.7.  Budget Pacing across Stages  . . . . . . . . . . . . . . 35
      13.8.  Failure and Recovery  . . . . . . . . . . . . . . . . . .35
      13.9.  Multi-Agent Stages  . . . . . . . . . . . . . . . . . . .36
      13.10.  Long-Horizon Requirements  . . . . . . . . . . . . . . .36
   14.  Affinity for Similar Agents and for Tenants  . . . . . . . . .37
      14.1.  Similarity and Reuse Groups  . . . . . . . . . . . . . . 37
      14.2.  Elements That Can Be Shared  . . . . . . . . . . . . . . 37
      14.3.  Conditions for Safe Reuse  . . . . . . . . . . . . . . . 38
      14.4.  Tenant Affinity Domains and Isolation  . . . . . . . . . 39
      14.5.  Fairness, Herding, and Quota  . . . . . . . . . . . . . .39
      14.6.  Measurement and Observability across Tenants  . . . . . .40
      14.7.  Requirements for Similar Agents and Tenants  . . . . . . 41
   15.  Interaction with Existing Work  . . . . . . . . . . . . . . . 42
      15.1.  Traceability to the Agent Service Requirements  . . . . .43
   16.  Operational Considerations  . . . . . . . . . . . . . . . . . 43
   17.  Security Considerations  . . . . . . . . . . . . . . . . . . .45
   18.  Privacy Considerations  . . . . . . . . . . . . . . . . . . . 45
   19.  IANA Considerations  . . . . . . . . . . . . . . . . . . . . .46
   20.  Normative References  . . . . . . . . . . . . . . . . . . . . 46
   21.  Informative References  . . . . . . . . . . . . . . . . . . . 46
   Acknowledgments  . . . . . . . . . . . . . . . . . . . . . . . . . 47
   Authors' Addresses  . . . . . . . . . . . . . . . . . . . . . . . .48

1.  Introduction

   The CATS framework selects a service contact instance using computing
   and network metrics [I-D.ietf-cats-framework]. Agent services add a
   consideration that is not represented in that metric set: how much of
   what a step needs is already in place at the candidate. A step of an
   agent session frequently can be served in two ways at a candidate
   instance -- by using state that the instance already holds and by
   using the computation that is already ready there, or by obtaining
   that state and that readiness again, either by transferring them from
   where they are held or by reconstructing them. The difference between
   the two paths is often larger than the difference between two
   candidate instances on any currently defined metric.

   Three dimensions of that consideration are distinguished in this
   document. State affinity is the preference for a candidate because it
   already holds the working set elements that a step needs. Storage
   affinity is the preference for a candidate because the tier at which
   those elements are held is usable for the step and because the
   durable store that can serve them is close to it. Compute affinity is
   the preference for a candidate because the model revision, precision,
   adapter, runtime, and accelerator that the step needs are resident
   and ready there, so that no artifact load, quantization change, or
   capability substitution is required.

Mo, et al.               Expires 3 April 2027                   [Page 4]
Internet-Draft           Agent Affinity                   September 2026

   The three are coupled, and that coupling is the reason this document
   treats them together. State can be reused only by a computation that
   can consume it, so a state that is present at an instance whose
   runtime cannot consume it has no reuse value at that instance.
   Changing the instance to gain compute affinity pays for the change in
   state: the accumulated state of a session is either moved, which
   costs transfer, or rebuilt, which costs accelerator time. Storage
   affinity is the bridge between the two, because it determines at
   which tier an element is held and how cheaply it can be made usable
   by the computation that needs it.

   This document states what a CATS system needs to know in order to
   take the three affinities into account, and what it has to be able to
   do about them. The document is organized around six questions.
   Section 4 states the motivation and the goals of introducing
   affinity. Section 8 states the requirements and the constraints that
   affinity places on the mapping of an agent workload onto storage and
   compute resources. Section 11 states how affinity is measured, and
   which metric properties make two measurements comparable. Section 12
   states the mechanisms by which affinity is assured and the procedure
   that applies them. Section 13 applies that procedure to a
   long-horizon session that passes through several stages. Section 14
   applies it to a group of similar agents and to tenants that share
   reusable state.

   The gap that motivates the document has been described for one
   element of the working set: the KV cache of a large language model,
   for which it has been observed that the existing CATS metrics do not
   expose cache state and that the distribution framework does not
   describe how cached content is distributed or synchronized across
   instances [I-D.li-cats-kv-cache-distribution]. The same reasoning
   applies to the other elements of an agent working set, and to the
   compute side of the same question.

   This document assumes the single-domain deployment model of the CATS
   framework. The intended standing of the document is informational 
   groundwork in the sense of the CATS charter [CATS-CHARTER]: it states 
   what a CATS system has to be able to express, so that the work can be
   taken up by the metric definition [I-D.ietf-cats-metric-definition]
   and by the data model [I-D.ietf-cats-data-model] rather than by a 
   protocol extension.

   Three distinctions are introduced here that the current documents do
   not make.

   *  A distinction between state availability, which is a quantity, and
      state locality, which is a constraint.

Mo, et al.               Expires 3 April 2027                   [Page 5]
Internet-Draft           Agent Affinity                   September 2026

   *  A distinction between instance affinity, which is about which
      service contact instance serves a session, and state affinity,
      which is about where the session's state and its computation are
      held.

   *  A distinction between the cost of obtaining state by transfer and
      the cost of obtaining it by re-computation or re-retrieval, and, on
      the compute side, between load, which is a property of an
      instance, and readiness, which is a property of the pair of a step
      requirement and an instance.

2.  Conventions and Definitions

   The key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT",
   "SHOULD", "SHOULD NOT", "RECOMMENDED", "NOT RECOMMENDED", "MAY", and
   "OPTIONAL" in this document are to be interpreted as described in BCP
   14 [RFC2119] [RFC8174] when, and only when, they appear in all
   capitals, as shown here.

   In this document these key words are used to state requirements on
   the design of a CATS system that supports agent services. They do not
   describe protocol behaviour, and this document defines no protocol,
   message format, or data model.

3.  Terminology

   This document uses the terms of [I-D.ietf-cats-framework] and
   [I-D.mo-cats-agent-service-characteristics]. The following additional
   terms are used.

   State Handle: An identifier under which a working set element can be
   addressed. A state handle is not a network address and does not by
   itself authorize access to the state that it names.

   State Availability: The presence of a working set element at, or
   within a reachable tier of, a candidate instance.

Mo, et al.               Expires 3 April 2027                   [Page 6]
Internet-Draft           Agent Affinity                   September 2026

   State Affinity: The preference for a candidate instance because it
   already holds, or is able to reach cheaply, the working set elements
   that a step needs.

   Storage Affinity: The preference for a candidate instance because the
   tier at which a working set element is held there is usable for the
   step, and because the durable store that can serve the element is
   cheap to reach from it.

   Compute Affinity: The preference for a candidate instance because the
   computation that a step requires is ready there: the model revision,
   the precision, the adapter, the runtime, and the accelerator class
   that the step requires are resident, so that the step can start
   without a load, a conversion, or a substitution.

   State Locality Constraint: A rule that forbids a working set element
   from being placed, transferred, or reused in a given location or
   scope.

   Affinity Domain: The set of instances among which a working set
   element may be reused, or among which the sessions of a tenant may be
   placed, under a stated scope.

   Affinity Gain: The reduction in the cost of serving a step that is
   attributable to reuse and to readiness, measured against the cost
   that the same step would have had if the state and the computation
   had been re-established.

   Affinity Decay: The reduction of affinity gain over time, caused by
   eviction, expiry, replacement of a model or adapter, drift of the
   session's requirement, or growth of the working set beyond the
   capacity that the holder is willing to commit.

   Affinity Budget: The share of the cost, latency, or capacity budget
   of a session that the session may spend on establishing, maintaining,
   or changing its affinity.

Mo, et al.               Expires 3 April 2027                   [Page 7]
Internet-Draft           Agent Affinity                   September 2026

   Retention Commitment: A statement by a holder that it will keep a
   working set element for a stated period or until a stated event,
   within a stated capacity limit and subject to stated preemption
   rules.

   Reuse Group: The set of sessions, steps, agents, or tenants for which
   a given working set element may legitimately be reused.

   Similar Agent: An agent whose steps require the same computation or
   the same reusable elements as another agent, as determined by the
   conditions of reuse and not by the identity of the agent.

   Stage Transition: The boundary at which the requirement profile of
   the next stage becomes known, and at which the placement and the
   affinity of the session may be re-established.

4.  Motivation and Goals

4.1.  Why Agent State, Storage, and Compute Are Coupled

   An agent session consumes three kinds of readiness, and each of them
   is produced at a cost that the selection decision can either pay or
   avoid.

   First, the state of the session. A step that can use the accumulated
   context of a session avoids re-reading the transcript, re-retrieving
   the documents, re-invoking the tools, and re-running the pre-fill. The
   alternative to reuse is not a constant: it depends on the length of
   the context that has to be re-established, on the throughput of the
   candidate at that operation, and on whether the material that has to
   be re-retrieved is available at all.

   Second, the tier at which that state is held. The same element has
   different costs depending on whether it is in accelerator memory, in
   host memory, in a node-local store, or only in a shared store that is
   reached over the network. The tier determines whether the element is
   usable as it stands or whether it has to be copied, paged, or
   deserialized before a step can use it, and it determines whether
   holding the element competes for the same accelerator capacity that
   the step itself needs.

   Third, the computation. A step of an agent session is not served by
   an arbitrary processor: it needs a model revision, a precision,
   sometimes an adapter, a runtime that can consume the state that is
   available, and an accelerator class that can execute it. When those
   are absent, the step can be served only after an artifact is loaded,
   a model is converted, or the step is redirected to a different
   capability tier, and each of those is a cost that is paid before the
   step begins.

Mo, et al.               Expires 3 April 2027                   [Page 8]
Internet-Draft           Agent Affinity                   September 2026

   The three are coupled because a state is reusable only by a
   computation that can consume it, and because a computation that is
   ready is worth little without the state to run on. A candidate that
   holds the KV state of a session but runs a different model revision
   cannot reuse it. A candidate that has the accelerator free but not
   the artifact must load it first, and the load costs more than the
   difference between many pairs of candidates. This is why an affinity
   decision that considers only one of the three dimensions produces a
   placement that is worse than one made without considering affinity at
   all: it moves the session to gain a dimension and pays for the gain
   in the other two.

4.2.  Failure Modes without an Affinity View

   Six failure modes are observed when affinity is not represented in
   the selection input.

   *  Repeated reconstruction. The same context is prefilled, the same
      documents are retrieved, and the same tools are invoked again
      because the selection did not know where the results were already
      held.

   *  Migration thrash. Because the cost of moving state is not
      comparable with the benefit of the move, sessions are moved for
      gains that do not survive the move, and the traffic that the moves
      generate degrades the conditions that motivated them.

   *  Concentration and its correction. If affinity is applied without a
      rule that overrides it, the instances that hold popular state absorb
      the demand, and if affinity is not applied at all, the state is
      re-established everywhere. Both extremes waste capacity in opposite
      directions.

   *  Isolation that is assumed rather than enforced. A reuse key that
      is treated as a hint rather than as a scoped identifier can cause
      state of one tenant to be reused by another, or a locality rule that
      was expressed as a cost to be traded to be violated silently.

   *  Loss of accumulated work. A long session that pauses and resumes
      on an instance that no longer holds its state, or that holds it at
      a tier the step cannot use, restarts from a checkpoint that may be
      much older than the session itself.

Mo, et al.               Expires 3 April 2027                   [Page 9]
Internet-Draft           Agent Affinity                   September 2026

   *  Decisions that are not reviewable. When the reason for a placement
      is not recorded, an operator cannot distinguish a session that was
      correctly pinned from one that was pinned by the absence of a
      competitive alternative.

4.3.  Goals

   The goals of introducing state, storage, and compute affinity into
   CATS are the following. They are stated as goals rather than as
   requirements; the requirements appear in Sections 8, 9, 11, 12, 13,
   and 14.

   *  G1. Make readiness visible. A selection function is able to learn,
      for a candidate and for the elements that a step needs, whether the
      element is present, at which tier, for how long, and whether the
      computation that would consume it is ready.

   *  G2. Make the alternatives comparable. The cost of reusing what is
      present is comparable with the cost of transferring it and with the
      cost of reconstructing it, on a common basis and for the same step.

   *  G3. Keep constraints out of the ranking. Locality, tenancy,
      capability, and consistency conditions are evaluated as constraints
      before any affinity preference is valued, so that a preference is
      never satisfied by paying for it in a dimension in which the step
      has a constraint.

   *  G4. Assure affinity rather than assume it. Affinity is established,
      maintained, verified, and released by stated mechanisms, with a
      procedure that reserves, commits, and can abort.

   *  G5. Sustain affinity over a horizon. Affinity is maintained across
      the steps of a long session, across pauses and stage transitions,
      and is re-evaluated when the assumption behind it changes.

   *  G6. Share without leaking. Reuse across similar agents and across
      the sessions of one tenant is enabled where it is permitted, and is
      bounded by scope, isolation, quota, and fairness rules where it is
      not.

Mo, et al.               Expires 3 April 2027                  [Page 10]
Internet-Draft           Agent Affinity                   September 2026

5.  Problem Statement

5.1.  Storage Is Represented Only as Capacity

   The CATS metric definition includes storage among the raw metrics
   that may be collected, in the form of available capacity, read
   throughput, and write throughput [I-D.ietf-cats-metric-definition].
   Those raw metrics describe an instance's storage as a resource that a
   workload consumes, in the same way that processor utilization or
   bandwidth describe other resources. They do not describe whether a
   particular piece of state is available, and they cannot be used to
   answer the question that matters for agent service selection.

5.2.  Affinity Is Instance-Level

   The framework defines service contact instance affinity, which keeps
   the traffic of a session on the same instance
   [I-D.ietf-cats-framework]. That concept is binary with respect to the
   instance: traffic either stays on the instance or does not. For agent
   services, affinity has a structure: a session may be able to continue
   on a new instance without penalty if one element of its working set
   is transferable and cheap, while it must stay where it is if the
   element is not transferable or its transfer is prohibited. An
   instance-level affinity flag cannot express that distinction.

5.3.  Transfer and Re-computation Are Not Comparable

   A CATS system that can see storage capacity but not which state is
   present cannot compare the two ways of obtaining missing state.
   Transferring a KV cache of a given size over a path with a given
   capacity and latency has a cost; recomputing it from the request has
   a cost that depends on the capability of the instance and the length
   of the context. The cheaper option determines which candidate is
   preferable, and the current metric set expresses neither.

5.4.  Compute Affinity Is Not Represented
Mo, et al.               Expires 3 April 2027                  [Page 11]
Internet-Draft           Agent Affinity                   September 2026

   The computing metrics of the framework describe an instance: its
   capability, its utilization, its queue. A step, however, does not
   need an instance; it needs a specific computation to be ready.
   Nothing in the current metric set states whether the artifact of the
   model revision that the step requires is resident at the candidate,
   whether the adapter that the session uses is loaded, whether the
   precision that the step needs is available on the accelerator of that
   candidate, or whether the runtime at the candidate can consume the
   state that is available there.

   The consequence is that a capability requirement is expressed as a
   filter on an instance rather than as a property of the pair of a
   requirement and a candidate: a candidate that is capable in general
   may still be unable to serve the step without a load, and a load that
   takes longer than the step itself is not visible to a selection that
   compares only utilization and path quality. The same applies to the
   state: a candidate whose runtime cannot consume the KV layout in
   which a prefix is stored holds a state that has no reuse value for
   that step.

5.5.  Reuse across Similar Agents and Tenants Is Not Represented

   An agent service usually serves many sessions whose steps require the
   same computation and the same reusable material: the same system
   prompt prefix, the same tool schemas, the same retrieval corpus, the
   same model artifact, the same adapter. Because the current
   information is per instance capacity and load, the fact that a prefix
   is already resident is invisible to the other sessions that could use
   it, and the pre-fill of a shared prefix is paid once per session
   instead of once per reuse group. Prefix-level steering for agent
   traffic has been proposed on the forwarding side
   [I-D.zhang-cats-token-aware-ts]; the corresponding reuse side is not
   represented.

   The tenant dimension has the same shape. A tenant has a placement
   domain, a capacity expectation, and an isolation requirement, and its
   sessions share state that may not be shared with another tenant.
   Neither the placement domain of a tenant nor the scope of a reusable
   element is expressible in the current selection input, and neither is
   the effect of affinity on the fairness between tenants.

5.6.  The Long Horizon Amplifies Each of These Gaps

Mo, et al.               Expires 3 April 2027                  [Page 12]
Internet-Draft           Agent Affinity                   September 2026

   An agent session can run for a long time, with pauses that exceed the
   lifetime of a cache entry, with stages that have different
   requirements, and with a working set that grows while it runs. Every
   one of the gaps above becomes more expensive over a longer horizon:
   state that is lost at a pause costs a resume from an older
   checkpoint, a placement that is re-evaluated at every step costs a
   transfer per step, and a stage transition that is not anticipated
   costs a cold start in the middle of a task. A selection input that is
   valid for a single request is therefore not sufficient, and the
   assurance procedure for a long session is not the same as the one for
   a single step.

6.  Working Set Elements Relevant to Agent Services

6.1.  Element Classes

   The elements below are the ones that agent service selection is
   expected to take into account. They differ in size, in volatility, in
   whether they may be shared, and in whether they may move.

      Element            | Volatility  | Reuse scope   | Locality
      -------------------+-------------+---------------+-----------
      Session context    | grows per   | per session   | jurisdiction
                         | turn        |               |
      Retrieved data set | per query   | per tenant    | policy
                         |             |               | dependent
      Tool results       | short-lived | per session   | endpoint
                         |             |               | bound
      KV cache or prefix | grows per   | per session   | accelerator
      state              | token       | or per prefix | bound
      Plan and scratch   | per step    | per session   | jurisdiction
      state              |             |               |
      Long-term memory   | appended    | per tenant or | policy
      or index           |             | per user      | dependent
      Model artifact     | versioned   | per site or   | hardware
                         |             | per tier      | bound
      Adapter or         | versioned   | per model or  | license
      quantization       |             | per tenant    | bound
      artifact           |             |               |
      Training           | per         | per job       | jurisdiction
      checkpoint         | iteration   |               |
      Runtime and kernel | versioned   | per site      | hardware
      image              |             |               | bound

   The last four elements arise in long-running jobs, such as the
   distributed training use case of [I-D.ietf-cats-usecases-requirements], 
   and are listed here because a long-running session raises the same
   selection question as an agent session.

Mo, et al.               Expires 3 April 2027                  [Page 13]
Internet-Draft           Agent Affinity                   September 2026

   The list is not exhaustive, and it is expected that deployments will
   add elements. What matters for this document is that the properties
   in the columns are the ones that selection needs, and that they are
   not properties of the instance alone but of the pair of the element
   and the instance.

   The reuse scope column is the basis of Section 14: an element whose
   reuse scope is wider than a session is material that several
   sessions, several agents, or one tenant can share, and it is the
   element class for which the cost of re-establishing the element is
   paid most often for no reason.

6.2.  Compute-Side Readiness

   The compute side has its own working set, and it is as expensive to
   re-establish as the state side. It consists of the artifacts and the
   runtime properties that a step requires: the weights of the model
   revision that the step needs, the adapter or the quantization
   artifact that the session uses, the runtime and kernel images that
   the accelerator requires, and the compiled or cached forms of the
   operations that the step executes.

   These elements differ from the state elements in one respect that
   matters to the selection: they are shared by many sessions, they are
   versioned rather than volatile, and their re-establishment cost is a
   load rather than a re-computation. A selection that compares two
   candidates on the basis of free accelerator capacity alone will
   prefer the candidate that has the memory free and will then pay the
   load, which is exactly the cost that a compute affinity view makes
   visible.

6.3.  Why the Properties Are Pairwise

   Each property in the table above is a property of a pair: the element
   and the instance. The tuple of the model revision, the tokenizer, and
   the precision that produced a KV state is part of the identity of
   that state for the purpose of reuse; the runtime of a candidate
   determines whether the state that the candidate holds is usable by a
   step; and the free capacity of a candidate determines whether the
   element fits. A summary of an instance that is not expressed against
   the element that a step needs therefore cannot answer the question
   that selection asks.

7.  State, Storage, and Compute Affinity

7.1.  Affinity, Preference, and Locality

Mo, et al.               Expires 3 April 2027                  [Page 14]
Internet-Draft           Agent Affinity                   September 2026

   Affinity is a preference, not a constraint. It states that a
   candidate is better because the material that a step needs is already
   in place there, and it may be traded against path quality, load,
   cost, and budget. A locality constraint, by contrast, is a rule that
   makes a candidate ineligible when the element is not where the rule
   permits it to be. The two are frequently confused in designs that
   express a locality rule as a very large weight in a score, which
   produces a system that admits a violation of the rule when the score
   happens to favor it.

   The three affinities are defined over different objects and have
   different lifetimes.

   *  State affinity is defined over a working set element and a
      candidate. It lasts as long as the element is present at that 
      candidate, and it is invalidated by eviction, expiry, or a change
      of the content of the element.

   *  Storage affinity is defined over a residency tier, a durable
      store, and a candidate. It lasts as long as the element is held
      at a tier that the step can use, or as long as the store that can
      serve the element is reachable at the cost that the decision、
      assumed.

   *  Compute affinity is defined over a step requirement and a
      candidate. It lasts as long as the artifacts and the runtime
      properties that the requirement names are resident, and it is
      invalidated by a change of the model revision, the adapter, 
      the precision, or the runtime.

7.2.  How the Three Affinities Interact

   The three dimensions are not additive, and the reason is that reuse
   requires a compatible consumer. The relations below are stated as the
   pairwise conditions that a selection has to test.

      Pairwise conditions at a candidate C for a step of session S:

Mo, et al.               Expires 3 April 2027                  [Page 15]
Internet-Draft           Agent Affinity                   September 2026

      state_reusable   | holds the element AND C can consume it:
                       | same model revision, tokenizer, and runtime
      state_usable     | element is at a tier that the step can use, or
                       | can be promoted into one within the step budget
      compute_ready    | the artifacts, precision, adapter, and
                       | accelerator of the step are resident at C
      compute_fit      | the accelerator has capacity for the state and
                       | for the working set of the step together
      affinity_gain    | cost_without_reuse - cost_with_reuse, where the
                       | two costs are measured for the same step

   A candidate that satisfies state affinity but not compute affinity
   holds a state it cannot use. A candidate that satisfies compute
   affinity but not state affinity must either receive the state or
   reconstruct it, and the cheaper of those two is the cost that the
   decision has to compare against the gain of the move. A candidate
   that satisfies storage affinity without either is close to the
   material but not ready to use it, and its value is the difference
   between the transfer time from the store and the transfer time from a
   distant holder.

   The interaction also produces an ordering that the rest of this
   document follows: constraints are tested first (Sections 8 and 9),
   the cost of not holding the state and of not being compute-ready is
   derived next (Section 11), and only then is affinity valued, assured,
   and maintained (Sections 12, 13, and 14). This is the same ordering,
   and the same division between constraints and quantities, as the
   selection mapping of [I-D.mo-cats-agent-selection-mapping].

7.3.  The Granularity of an Affinity Decision

   Affinity is decided per element and per step, not per instance and
   not per session. The practical consequences are as follows.

   *  A decision may hold for one element and not for another, so the
      unit of a decision is the element that a step needs.

   *  A decision may be revisited at a step boundary without moving the
      session, because one element may be re-established where it stands
      while the rest of the working set stays in place.

   *  A decision that covers several steps has a longer lifetime than a
     decision that covers one, and the longer lifetime is what makes a
     retention commitment worth its capacity.

   *  A decision that applies to a group of sessions (Section 14) has a

Mo, et al.               Expires 3 April 2027                  [Page 16]
Internet-Draft           Agent Affinity                   September 2026

     lifetime that is bounded by the version of the element, not by the
     lifetime of any one session.

8.  Requirements and Constraints on the Mapping

8.1.  The Mapping Relation

   The mapping that this section constrains takes a step of an agent
   session and produces, for each candidate instance, the constraints
   that the step imposes and the quantities by which the candidates
   differ. It is the storage-and-compute part of the mapping described
   in [I-D.mo-cats-agent-selection-mapping], stated here at the level of
   detail that affinity requires.

   The demand side of the mapping is derived from the step and from the
   session it belongs to. The supply side is observed at the candidate.
   The result is a pairwise valuation.

      Mapping element      | Item             | Content
      ---------------------+------------------+----------------------
      Demand of a step     | needs_state      | keys, tier, locality,
                           |                  | age
                           | needs_storage    | capacity, tier, store
                           |                  | locality
                           | needs_compute    | revision, precision,
                           |                  | adapter, accelerator,
                           |                  | rate
      Supply at a          | holds_state      | key, tier, retention
      candidate            |                  |
                           | offers_storage   | capacity, tier,
                           |                  | locality
                           | offers_compute   | readiness, free
                           |                  | capacity, queue
      Pairwise result      | state_ready      | holds the element and
                           |                  | can consume it
                           | state_cost       | min(transfer time,
                           |                  | recompute time)
                           | affinity_gain    | cost without reuse,
                           |                  | minus the cost with
                           |                  | reuse

   Three properties of this mapping are required for affinity to be
   usable. The mapping is per step, because the requirement of a step is
   what a candidate is evaluated against. The mapping distinguishes
   constraints from quantities, because a constraint cannot be traded.
   And the mapping is recomputable from information that the candidate
   can report without disclosing the content of the state, because
   otherwise affinity would be expressible only inside a single trust
   domain.

8.2.  Constraint Families
Mo, et al.               Expires 3 April 2027                  [Page 17]
Internet-Draft           Agent Affinity                   September 2026

   The constraints that affinity imposes on the mapping belong to seven
   families. Each of them is evaluated before candidates are ranked, and
   each of them can make a candidate ineligible.

   *  Capacity: the element and the working set of the step fit in the
      tier that the step requires at the candidate.

   *  Capability: the model revision, precision, adapter, and
      accelerator class that the step requires are available at the
      candidate.

   *  Compatibility: the candidate can consume the state that it holds
      or receives, which requires a matching model, tokenizer, runtime, 
      and layout.

   *  Locality: the element and the execution remain within the
      jurisdiction, tenant domain, or device boundary that the element
      carries.

   *  Consistency: the element is current with respect to the version
      that the step assumes, and is not under an eviction or a replacement
      that would make it unusable during the step.

   *  Isolation: reuse does not cross a session, user, or tenant
      boundary that policy protects, and the timing of a hit does not
      disclose the presence of another session's state.

   *  Temporal: the element remains held, and the computation remains
      ready,

   for the duration that the decision assumes, which for a long-horizon
   session is a commitment that exceeds the step.

   A budget condition is not a constraint of this kind: exceeding a
   budget reduces the value of a candidate rather than making it
   ineligible, and it is therefore expressed as a quantity, except where
   a session has a hard budget and a step that cannot be served within
   it must be failed rather than degraded.

8.3.  Mapping Requirements
Mo, et al.               Expires 3 April 2027                  [Page 18]
Internet-Draft           Agent Affinity                   September 2026

   The requirements below are stated using the conventions of BCP 14
   [RFC2119] [RFC8174]. They are requirements on the information and on
   the mapping that a CATS system needs in order to consider affinity,
   and are not protocol requirements.

   WM1. A CATS system SHOULD be able to express the demand of a step as
   a requirement on each of the state, storage, and compute dimensions,
   rather than as a single scalar.

   WM2. A CATS system MUST distinguish, in that expression, a constraint
   that makes a candidate ineligible from a preference that may be
   traded against other quantities.

   WM3. A CATS system SHOULD be able to express a state requirement as a
   set of required elements, each named by a state handle, with an
   indication of whether the element is required to be present at a
   stated tier or may be made present at a stated cost.

   WM4. A CATS system SHOULD be able to express a storage requirement as
   a capacity and a tier requirement, together with a locality
   constraint on the durable store that holds the element.

   WM5. A CATS system SHOULD be able to express a compute requirement as
   a capability requirement, comprising the model revision, the
   precision, the accelerator class, the runtime, and any adapter,
   together with the rate at which the step needs to consume.

   WM6. A CATS system MUST evaluate capacity, capability, compatibility,
   locality, consistency, isolation, and temporal constraints before it
   values any affinity preference, and MUST NOT satisfy a constraint by
   paying for it in another dimension.

   WM7. A CATS system SHOULD be able to express the mapping at the
   granularity of a step, and SHOULD be able to carry the part of a
   session's mapping that is unchanged from one step to the next without
   re-deriving it.

   WM8. A CATS system SHOULD be able to express the residual requirement
   of a step that cannot be served in full, so that a partial reuse, a
   partial result, or a degraded execution is expressed as a reduction
   of the requirement rather than as a silent failure.

   WM9. A CATS system SHOULD be able to state, for each element of the
   mapping, the unit in which it is expressed, the component that
   observes it, and the bound within which it remains valid.

   WM10. A CATS system SHOULD be able to compare the mapping across
   candidates of different capability tiers without assuming that the
   tiers are interchangeable, and SHOULD preserve the reason when a
   candidate is excluded by a constraint.
Mo, et al.               Expires 3 April 2027                  [Page 19]
Internet-Draft           Agent Affinity                   September 2026

9.  Affinity Information Requirements

   The requirements below are stated using the conventions of BCP 14
   [RFC2119] [RFC8174]. They are requirements on the information that a
   CATS system needs in order to consider state, storage, and compute
   affinity, and are not protocol requirements.

9.1.  State Availability

   S1. A CATS system SHOULD be able to determine, for a candidate
   instance, whether a working set element is available, and under which
   state handle it can be addressed.

   S2. State availability information SHOULD be expressed in a form that
   can be aggregated, so that an instance is not required to advertise
   every element that it holds.

   S3. State availability information SHOULD carry an indication of its
   freshness, so that a selection function can distinguish a recent
   observation from a stale one.

   S4. State availability SHOULD be expressed at the granularity at
   which reuse is meaningful, which for a KV cache means at the level of
   a reusable prefix or block rather than at the level of an entire
   instance.

9.2.  Residency Tiers and Retrieval Cost

   S5. A CATS system SHOULD be able to distinguish the tier at which a
   working set element resides, at least between accelerator memory,
   host memory, node-local storage, and a shared store.

   S6. A CATS system SHOULD be able to represent the cost of making an
   element usable at a candidate instance, including both the cost of
   transferring it and the cost of recomputing or re-retrieving it.

   S7. A CATS system SHOULD be able to identify the cheaper of transfer
   and re-computation for a given element and candidate, so that the
   selection function can prefer the candidate that yields the lower
   total cost.

9.3.  Constraints, Consistency, and Sharing

   S8. A CATS system MUST be able to express state locality constraints
   as constraints that are evaluated before candidates are ranked.

   S9. A CATS system SHOULD be able to distinguish state that may be
   reused across sessions, across tenants, or not at all, and SHOULD NOT
   require a candidate to disclose state that may not be shared.

Mo, et al.               Expires 3 April 2027                  [Page 20]
Internet-Draft           Agent Affinity                   September 2026

   S10. A CATS system SHOULD be able to represent the consistency
   conditions under which a stored element may be reused, including at
   least whether it is current and whether an eviction is pending.

   S11. A CATS system SHOULD be able to represent the cost and the
   conditions of changing the instance that serves a session, so that
   affinity is applied only when the migration is actually cheaper than
   the alternative.

   S12. A CATS system SHOULD NOT require the network to learn the
   content of a working set element in order to select an instance for
   it.

9.4.  Compute Affinity and Reuse Scope

   S13. A CATS system SHOULD be able to determine, for a candidate,
   whether the computation that a step requires is ready there without
   being re-established: the model revision, the precision, the adapter,
   the runtime, and the accelerator class that the step needs.

   S14. A CATS system SHOULD be able to distinguish compute readiness,
   which is a property of the pair of a step requirement and a
   candidate, from computing load, which is a property of the candidate
   alone.

   S15. A CATS system SHOULD be able to represent the affinity that a
   session, an agent, or a tenant has accumulated with an instance in a
   form that survives the end of a turn and the interval between turns.

   S16. A CATS system SHOULD be able to represent the reuse scope of an
   element, that is, the set of sessions, agents, or tenants for which
   the element may be reused, without requiring the content of the
   element to be disclosed.

10.  State Handles and Their Relationship to CATS Identifiers

   Selection requires that a candidate can state which elements it
   holds, and that a selection function can compare that statement with
   what a step needs. This requires an addressing scheme for working set
   elements.

   The following properties are proposed for such a handle.

   *  A state handle identifies a working set element, not a location.
      It does not replace the CATS Service Identifier or the CATS Service
      Contact Instance ID, and it is not a routable address.

   *  A state handle is scoped. It is meaningful between the parties
      that use it, and it carries or implies the scope within which it
      may be used, such as a session, a tenant, a reuse group, or a
      site.
Mo, et al.               Expires 3 April 2027                  [Page 21]
Internet-Draft           Agent Affinity                   September 2026

   *  A state handle is subject to authorization. Knowledge of a handle
      is not authorization to read, transfer, or reuse the state that it
      names. Authorization is expected to be provided by the identity and
      authorization mechanisms of the environment in which the agent
      service runs.

   *  A state handle is revocable. Eviction, expiry, a change of the
      reuse scope, or a policy change can invalidate a handle, and the
      selection function is expected to tolerate a handle that has become
      invalid.

   *  A state handle is comparable only within a defined equivalence.
      Two handles denote the same reusable state only if the deployment
      defines the equivalence, for example equality of a content hash over
      a defined model, tokenizer, precision, and prefix.

   The last property is the one that the compute and sharing dimensions
   make sharper. For the reuse of a model-side state, the equivalence
   has to cover the model revision, the tokenizer, the precision, and
   the runtime layout, because a KV state that was produced by a
   different combination is not reusable by the step even when its
   content is identical. This is the same condition that a cache applies
   when it decides whether a stored representation may be reused for a
   request [RFC9111]: the stored material is reusable only if the key
   and the validators of the current request match. The parallel is
   stated here because it shows that the condition is a property of the
   pair of the material and the consumer, and not a property of the
   material alone.

   The granularity of a state handle is the granularity at which reuse
   is meaningful, as required by S4, and it is the key under which the
   storage dimension of agent service selection reports availability
   [I-D.mo-cats-agent-selection-mapping].

   The framework does not define the syntax of a state handle, and this
   document does not propose one. A syntax is needed only if a protocol
   is later defined to exchange state availability, at which point the
   syntax becomes the subject of the document that defines that
   exchange.

11.  Measuring Affinity: Metrics and Methods

Mo, et al.               Expires 3 April 2027                  [Page 22]
Internet-Draft           Agent Affinity                   September 2026

11.1.  What Has to Be Measured

   Affinity is an expectation about the cost of the next step, and it
   can be verified only by measuring what the step actually cost and
   what it would have cost without reuse. Four questions therefore have
   to be answerable from measurement.

   *  Is the material there? Presence and residency have to be
      observable per reuse key, and not only as an aggregate capacity,
     because the value of a candidate depends on the specific element
     that a step needs.

   *  Is the computation ready? Readiness has to be observable as a
      property of the pair of a requirement and a candidate, which means
     that a measurement of utilization or of installed capability is not
     a substitute for it.

   *  What did the reuse save? Affinity gain has to be measured as a
     difference between two costs of the same step, with the basis of the
     comparison stated, because a gain measured against a different
     baseline is not comparable with a gain measured elsewhere.

   *  What did affinity cost? The transfer that a decision caused, the
     capacity that a retention commitment withheld, and the degradation
     that the concentration of sessions caused are costs of affinity and
     need to be attributed to the decision that incurred them.

11.2.  Affinity Metric Catalogue

   The metrics below are the ones that the four questions require. They
   are stated at the level of the quantity and its observation point,
   not as an encoding, and the levels follow the raw and derived
   distinction of [I-D.ietf-cats-metric-definition].

Mo, et al.               Expires 3 April 2027                  [Page 23]
Internet-Draft           Agent Affinity                   September 2026

      Affinity metric        | Unit      | Level  | Observed at
      -----------------------+-----------+--------+----------------
      Element present under  | boolean   | raw    | C-SMA, per key
      a reuse key            |           |        |
      Residency tier of an   | tier      | raw    | C-SMA, per key
      element                |           |        |
      Retention remaining    | seconds   | raw    | C-SMA, per key
      Compute readiness of a | boolean   | raw    | C-SMA, per pair
      step requirement       |           |        |
      Reuse hit ratio        | fraction  | derived | C-PS, per key
                             |           |        | and window
      Reuse distance of a    | steps,    | derived | C-PS, per
      session                | seconds   |        | session
      Transfer volume caused | bytes     | raw    | C-NMA, per
      by a decision          |           |        | decision
      Transfer time caused   | seconds   | derived | C-NMA, per
      by a decision          |           |        | decision
      Recompute volume       | tokens,   | raw    | C-SMA, per step
      (re-prefill, reload)   | bytes     |        |
      State cost of a step   | seconds,  | derived | C-PS, per pair
      at a candidate         | cost      |        |
      Affinity gain of a     | seconds,  | derived | C-PS, per step
      step                   | cost      |        |
      Selections changed for | count,    | derived | C-PS, per
      lack of state          | reason    |        | window
      Affinity concentration | fraction  | derived | C-PS, per site
      on an instance         |           |        |
      Tenant affinity        | bytes,    | derived | C-SMA, per
      footprint              | fraction  |        | tenant
      Reuse refused by scope | count,    | derived | C-SMA, per
      or policy              | reason    |        | window

   The metrics are grouped by the question they answer. Presence,
   residency, retention, and readiness answer the first two questions.
   Reuse hit ratio, reuse distance, transfer volume and time, recompute
   volume, and state cost answer the third. Affinity gain answers the
   third and the fourth together, because it states the saving rather
   than the cost. Concentration, tenant footprint, and refused reuse
   answer the fourth question and are also the inputs of the fairness
   rules of Section 14.

11.3.  Metric Semantics and Reporting Standards

   A quantity that two components report under the same name is
   comparable only if the properties below are fixed. They are the
   reporting standards that this document asks of a metric definition,
   and they follow the treatment of freshness and of unknown values in
   [I-D.zhu-cats-metric-semantics].

   *  Unit and basis. Each metric names the unit in which it is
      expressed and the population over which it is computed. A hit
      ratio without its window and its population is not comparable
      with another hit ratio.
Mo, et al.               Expires 3 April 2027                  [Page 24]
Internet-Draft           Agent Affinity                   September 2026

   *  Level. Each metric is stated as raw, observed at a component, or
      as derived, computed from raw values. Affinity gain, state cost, and
      reuse distance are derived; presence, residency, retention,
      readiness, and transfer volume are raw.

   *  Observation point. Each metric names the component that observes
      it and the granularity at which it is observed, which for the
      affinity metrics is the reuse key, the pair of a requirement and a
      candidate, the session, or the site.

   *  Freshness. Each reported value carries the time at which it was
      observed and the bound within which it may be used, and a value
      outside that bound is treated as unknown.

   *  Percentile and tail. Quantities that describe a distribution are

   reported at a stated percentile over a stated window, and a quantity
   that is a mean is identified as a mean, so that a tail objective is
   not evaluated against an average.

   *  Unknown. A value that is absent, stale, or of unknown provenance
      is reported as unknown rather than as a default, and an unknown value
      does not satisfy a constraint.

   *  Aggregation. Affinity metrics are aggregated per class of element,
      per key, per site, and per tenant. Aggregation reduces cardinality but
      must not merge populations whose objectives differ, and it must not
      turn a per-tenant quantity into a value that discloses another
      tenant.

11.4.  Measurement Methods

   The quantities above are obtainable by four methods, and the choice
   among them is a deployment decision.

   *  Observation at the holder: the instance that holds an element
      reports presence, tier, retention, and the outcome of the reuse of the
Mo, et al.               Expires 3 April 2027                  [Page 25]
Internet-Draft           Agent Affinity                   September 2026

        element. This is the most accurate method for the state side and the
        only one that can see eviction.

   *  Observation at the decision point: the component that selects
        records the cost of the step for which a decision was taken, the
        alternative cost that the same step faced at the candidates that
        were rejected, and the reason for the choice. This is the only method
        that can produce affinity gain, because the gain is a difference between
        two costs of the same step.

   *  Probing: a periodic or on-demand test of the presence of a key,
      used where the holder does not report, and used after a pause to
        re-verify the state on which a session depends. Probing has to be rate
        limited, because a probe is an observable event and because a probe storm
        competes with the traffic it measures.

   *  Accounting of transfers: the transfer volume and time that a
      decision caused, separated from background traffic, so that the cost of
        affinity maintenance is visible. This is required for the operational
        rule that bounds transfer as a fraction of decisions.

   Measurement itself is subject to the constraints of the framework: a
   measurement that can say which session produced a byte has to be
   protected as session information, and multi-tenant measurement has to
   be aggregated so that it does not become an information channel
   (Section 18). The OAM functions of [I-D.ietf-cats-oam-fw] are the
   natural carrier of these quantities.

11.5.  Measurement Requirements

   MA1. A CATS system SHOULD define each affinity metric with its unit,
   its observation point, and its level, so that two components that
   report the same metric report a comparable quantity.

   MA2. A CATS system SHOULD measure the presence and the residency of a
   working set element per reuse key, and SHOULD NOT require an instance
   to enumerate its keys to the network.

   MA3. A CATS system SHOULD measure the hit ratio of state reuse per
   key and per session over a stated window, and SHOULD report it
   together with the window and the population over which it was
   computed.

Mo, et al.               Expires 3 April 2027                  [Page 26]
Internet-Draft           Agent Affinity                   September 2026

   MA4. A CATS system SHOULD measure the transfer that a decision causes
   and the recomputation that a decision causes separately, so that the
   two ways of obtaining missing state can be compared.

   MA5. A CATS system SHOULD measure affinity gain as the difference
   between the cost of a step that reused state or ready compute and the
   cost that the same step would have had without that reuse, and SHOULD
   state the basis of the comparison.

   MA6. A CATS system SHOULD report latency and cost quantities at a
   stated percentile over a stated window, and MUST NOT report a
   quantity measured over a population as if it were a property of a
   single session.

   MA7. A CATS system SHOULD carry the observation time and the validity
   bound of each measurement, and SHOULD NOT use a measurement outside
   that bound to satisfy a constraint.

   MA8. A CATS system SHOULD aggregate affinity measurements per class
   of element, per site, and per tenant, and MUST NOT expose a
   per-session affinity measurement to a party that is not authorized to
   observe that session.

   MA9. A CATS system SHOULD distinguish, in its measurements, a miss
   that is caused by eviction from a miss that is caused by a scope or
   policy prohibition, because the two have different remedies.

   MA10. A CATS system SHOULD measure the effect of affinity on the
   distribution of load, including the concentration of sessions on the
   instances that hold popular state and the effect of an
   affinity-driven selection on the tail of other sessions.

12.  Affinity Assurance: Mechanisms and Procedures

12.1.  Mechanism Catalogue

   Affinity is assured by a small set of mechanisms. Each mechanism has
   an effect and a cost, and the choice among them is what turns the
   information of the previous sections into a placement that is stable
   and fair.

Mo, et al.               Expires 3 April 2027                  [Page 27]
Internet-Draft           Agent Affinity                   September 2026

      Mechanism                | Effect              | Cost or risk
      -------------------------+---------------------+--------------
      Placement pinning        | keeps a session     | load
                               | with the holder of  | concentration
                               | its state           |
      Retention lease          | a holder commits to | capacity
                               | keep an element     | withheld
      Replication or warm copy | shortens the path   | refresh cost
                               | to the state        |
      Prefetch before a stage  | state is ready      | transfer if
                               | before the first    | unused
                               | step                |
      Recompute in place       | avoids a transfer   | accelerator
                               | entirely            | time
      Handoff with reserve and | moves a session     | two-instance
      commit                   | without a gap       | overlap
      Eviction protection      | keeps a hot or      | capacity for
                               | shared element      | others
                               | resident            |
      Admission control        | bounds transfers in | queued
                               | flight              | sessions
      Quota or reservation     | bounds per-session  | under-use
                               | or per-tenant use   |
      Load-aware override      | prevents herding on | affinity gain
                               | one holder          | lost
      Degradation ladder       | keeps a session     | worse
                               | alive under         | objective
                               | pressure            |
      Release and eviction     | reclaims state that | early loss of
      policy                   | has no horizon      | reuse

   The mechanisms are complementary and are combined rather than chosen
   between. A retention lease without a release policy exhausts the
   capacity of the holder; a release policy without a lease makes the
   state of a long session disappear during a pause; an override without
   a measurement of concentration cannot be applied at the right moment.

12.2.  Assurance Procedure

   The procedure below is the loop that applies the mechanisms. It is
   stated as a sequence of steps with the information that each step
   consumes and the decision that it produces, and it is applied at each
   decision point of a session.

Mo, et al.               Expires 3 April 2027                  [Page 28]
Internet-Draft           Agent Affinity                   September 2026

      Step   | Input                    | Output
      -------+--------------------------+---------------------------
      A1     | presence, tier,          | per-candidate supply view
             | retention, readiness,    |
             | load                     |
      A2     | supply view, constraints | eligible candidate set
      A3     | eligible set, size,      | state cost per candidate
             | rate, locality           |
      A4     | gain, path, load,        | preferred candidate
             | budget, concentration    |
      A5     | preferred candidate,     | commitment or reservation
             | next steps               |
      A6     | missing elements         | transfer, prefetch, or
             |                          | reconstruction
      A7     | committed state and      | verified placement and
             | compute                  | record
      A8     | observed gain, decay,    | next decision point
             | triggers                 |

   Step A2 applies the constraints of Section 8.2 and excludes the
   candidates that break one of them, before any affinity is valued.
   Step A3 derives the cost of making the working set usable, which is
   the cheaper of the transfer and the reconstruction of each missing
   element. Step A4 values affinity as the reduction of that cost and
   trades it against path quality, load, budget, and the concentration
   that the choice would create. Steps A5 and A6 act, and step A7
   verifies rather than assumes that the material is in place, because a
   commitment can expire between the decision and the step. Step A8
   closes the loop and is the reason the procedure is not a one-shot
   placement: the observed gain and the decay of the affinity are the
   inputs of the next decision.

12.3.  Triggers for Re-evaluation

   The procedure is re-entered when one of the following events is
   observed.

   *  Session admission, at which the first decision is taken.

   *  A step boundary, at which the requirement of the next step is
      known and may differ from that of the previous one.

   *  A stage transition of a long-horizon session (Section 13).

   *  A pause that exceeds the freshness bound of the presence
      information that the current placement relies on.

Mo, et al.               Expires 3 April 2027                  [Page 29]
Internet-Draft           Agent Affinity                   September 2026

   *  An eviction notice, an expiry of a retention commitment, or a
        replacement of a model revision, an adapter, or a runtime that the
        session depends on.

   *  A capacity pressure or a rejection of an admission request, which
        indicates that the capacity that the decision assumed is no longer
        available.

   *  A measured breach of the objective of the session, or a measured
        concentration of sessions on an instance that degrades the objective
        of other sessions.

   *  A change of the locality, tenancy, or reuse-scope policy that
      governs an element of the working set.

12.4.  Abort, Failure, and Fallback

   A move that has started may fail at any of its steps, and the failure
   semantics matter more for affinity than they do for a stateless
   decision, because a partially moved working set is worse than either
   endpoint. The following properties are expected of the procedure.

   *  The source of a move remains usable until the target has verified
      that it holds the material that the session needs, so that a failure
        leaves the session where it was rather than in neither location.

   *  The verification that a step performs before it runs is the point
      at which a failed commitment is detected, and the detection leads to the
        degradation classes of the step rather than to an unbounded retry.

   *  A move that cannot be completed within the affinity budget of the
        session is abandoned and the session continues where it is, with the
        element re-established locally if that is cheaper than the move.

   *  When no candidate satisfies the constraints, the session degrades
      in the order of the fallback ladder of the selection mapping rather than
        failing silently, which is consistent with the notion of a fallback
        decision [I-D.pang-cats-fallback-decision-framework].

Mo, et al.               Expires 3 April 2027                  [Page 30]
Internet-Draft           Agent Affinity                   September 2026

12.5.  Assurance Requirements

   AM1. A CATS system SHOULD apply affinity as a preference that is
   valued after the constraints and alongside load, with a stated weight
   or order, and SHOULD NOT allow it to override a constraint.

   AM2. A CATS system SHOULD be able to hold a retention commitment for
   a working set element for a stated period or until a stated event,
   and SHOULD be able to release that commitment before its end.

   AM3. A CATS system SHOULD be able to establish the state and the
   compute readiness that the next steps need before those steps start,
   rather than only at the moment at which a step is served.

   AM4. A CATS system SHOULD be able to choose reconstruction in place
   as an alternative to a transfer, when the transfer would be larger,
   slower, or prohibited.

   AM5. A CATS system SHOULD verify, before a step is served, that the
   state and the compute readiness that the decision assumed are still
   present and usable, and SHOULD fall back when they are not.

   AM6. A CATS system SHOULD bound the fraction of decisions that may
   trigger a transfer, and SHOULD attribute the transfer that a decision
   causes to the session that the decision serves.

   AM7. A CATS system SHOULD override affinity when the concentration of
   sessions on the instances that hold popular state would degrade the
   objective of other sessions, and SHOULD make that override visible in
   the selection context.

   AM8. A CATS system SHOULD release state that is no longer eligible
   for reuse, and SHOULD NOT hold state for a session whose horizon has
   ended.

   AM9. A CATS system SHOULD re-evaluate an affinity decision when an
   assumption behind it changes, including an eviction, an expiry, a
   change of model revision or adapter, a change of policy, or a
   sustained deviation of the observed gain from the expected gain.

   AM10. A CATS system SHOULD be able to abandon a move after it has
   started, and SHOULD leave both the source and the target in a usable
   state when it does so.

   AM11. A CATS system SHOULD apply the same assurance procedure to a
   change of residency tier within an instance as to a change of
   instance, because both change the cost of the next step.

   AM12. A CATS system SHOULD retain the decision and its outcome for
   the steps that follow, so that the affinity of a session is not
   re-derived from scratch at every step.
Mo, et al.               Expires 3 April 2027                  [Page 31]
Internet-Draft           Agent Affinity                   September 2026

13.  Affinity Assurance for Long-Horizon and Multi-Stage Sessions

13.1.  Stage Model

   A long-horizon agent session does not have one requirement profile;
   it has a sequence of them. A session that plans, then retrieves, then
   analyzes, then calls a tool, and then reports has stages whose
   working sets overlap in part, whose compute requirements differ, and
   whose tolerable latency differs. The stage is the unit at which the
   assurance procedure of Section 12 can be applied without paying for a
   change at every step.

      Stage          | Working set that grows   | Compute profile
      ---------------+--------------------------+-------------------
      Plan           | task state, plan         | capable model
      Retrieve       | retrieved set, index     | embedding, search
      Analyze        | context, intermediate    | long context
      Act            | tool results, arguments  | low-latency calls
      Report         | assembled result         | capable model

   The model is a description, not a specification: a deployment defines
   its own stages. What matters for affinity is that the boundary
   between two stages is the point at which the requirement profile
   becomes known in advance, and therefore the point at which affinity
   can be re-established proactively rather than reactively.

13.2.  Phase-Differentiated Assurance

   The assurance procedure has a different emphasis in each phase of a
   long-horizon session.

   *  Admission: the session is placed on the basis of its first stage,
      and the elements of the later stages that are already known are
        prefetched where that is cheap.

   *  Steady state within a stage: the placement is kept stable, the
      retention commitments are renewed as they approach their end, 
        and the affinity gain is measured rather than assumed.

   *  Pause: the retention commitment is what protects the session, and
      its length is chosen by comparing the cost of holding the state with
      the cost of re-establishing it at resume.

Mo, et al.               Expires 3 April 2027                  [Page 32]
Internet-Draft           Agent Affinity                   September 2026

   *  Resume: the presence of the elements on which the session depends
      is re-verified, and an element that is no longer present is treated
        as absent, so that the resume decision is taken on observed rather than
        on remembered state.

   *  Stage transition: the affinity of the next stage is established
        proactively, and the placement of the session is changed at that
        boundary if the next stage is better served elsewhere.

   *  Completion: the state is released, and the elements whose reuse
      scope is wider than the session (Section 14) are kept where the reuse
        group benefits from them.

13.3.  Pinning, Retention, and Tier Budget

   The capacity that a session may keep warm is finite, and the decision
   of what to keep is a comparison of the cost of holding against the
   cost of re-establishing. Four rules make that decision tractable.

   *  Keep the elements whose reconstruction is expensive and whose
      reuse is certain, which for a long session is the accumulated context
        and the model-side prefix state.

   *  Keep the elements whose reconstruction is impossible, such as the
      result of a tool invocation that cannot be repeated safely or a
        retrieval that is no longer reachable.

   *  Do not keep an element whose reconstruction is cheaper than its
        retention, which for a small derived artifact is usually the case.

   *  Place the elements that are kept at the cheapest tier from which
      the step can use them, and promote an element to a faster tier only
        for the stage that needs it, because promotion competes with the state
        of the steps that are running.

13.4.  Checkpoint and Resume after a Pause

Mo, et al.               Expires 3 April 2027                  [Page 33]
Internet-Draft           Agent Affinity                   September 2026

   A pause longer than the retention commitment of the holder, a failure
   of the holder, or an administrative move makes the state of a session
   unavailable. Three properties make the session resumable.

   *  A checkpoint of the durable part of the working set, taken at a
      stage boundary, bounds the work that a resume can lose.

   *  The part of the working set that is sufficient to continue is
      stated, so that a resume does not attempt to re-establish the whole
        set before the session can make progress.

   *  The resume is a decision point like any other: the presence of the
        elements is verified, the cost of the resume is compared across
        candidates, and the placement that was in force before the pause is
        not assumed to be valid.

13.5.  Stage Transitions

   A stage transition is the natural point at which a placement may
   change, and a change elsewhere is what the stability rule of the
   selection mapping is intended to prevent. Three properties apply.

   *  The requirement profile of the next stage is known at the
      boundary, so

   the constraint test of the next stage can be performed before the
   stage starts, and a candidate that will fail a constraint can be
   excluded without waiting for the failure.

   *  The elements that the next stage needs can be prefetched or

   reconstituted during the last part of the current stage, in parallel
   with work that is still running, so that the transition does not
   begin with a serial load.

   *  The cost of the transition is attributed to the stage that
      benefits from it, and a transition that serves several remaining
        stages is charged to the session rather than to one step.

13.6.  Drift, Decay, and Re-evaluation

Mo, et al.               Expires 3 April 2027                  [Page 34]
Internet-Draft           Agent Affinity                   September 2026

   Affinity decays, and the decay is gradual rather than binary. The
   working set grows with the session, so the element that fits at
   admission may not fit later. The context that the session uses may
   become less similar to the prefix that is cached. The load of the
   holder may grow until the affinity gain no longer compensates for it.
   A session whose placement was correct at step ten can be wrong at
   step thirty without any single event having changed.

   Two quantities make the decay observable: the observed affinity gain,
   compared with the gain that the placement was expected to deliver,
   and the reuse hit ratio of the session over a recent window. A
   sustained reduction of either, beyond the tolerance of the session,
   is the trigger for re-evaluation. Re-evaluation is not the same as
   movement: it may conclude that re-establishing one element at the
   current instance is cheaper than moving the session, which for a
   large working set is usually the case.

13.7.  Budget Pacing across Stages

   A long-horizon session has a budget, and affinity maintenance spends
   it. The cost of a stage transition, the capacity that a retention
   commitment withholds, and the traffic of a prefetch are all
   affordable in isolation and not affordable in a loop. The procedure
   therefore paces affinity maintenance over the remaining horizon: at
   each stage boundary, the budget that the remaining stages need is
   reserved before the affinity of the current stage is extended, and
   the transfer that a decision causes is charged to the session budget
   that the decision serves.

13.8.  Failure and Recovery

   The failure modes that matter for affinity over a horizon are the
   loss of the holder, the loss of the material, and the loss of the
   placement's value. Each has a recovery path, and the path that was
   taken is part of the decision record.

   *  Loss of the holder: the session is resumed from a checkpoint, on
      an instance chosen by the same procedure, with the elements that
        survived reused where they are reachable.

   *  Loss of the material: the element is reconstructed if that is
      cheaper than re-establishing it elsewhere, and the reconstruction
        is measured as a recompute event so that the failure is visible in the
        accounting.

   *  Loss of the value: the placement no longer delivers a gain, and
      the session is moved at the next stage boundary rather than immediately,
Mo, et al.               Expires 3 April 2027                  [Page 35]
Internet-Draft           Agent Affinity                   September 2026

        unless a constraint makes the current placement ineligible.

13.9.  Multi-Agent Stages

   A stage may be served by several agents that share a working set: a
   planner and a set of workers, or a set of agents that read the same
   retrieved material. Affinity for such a stage is a property of the
   shared element as well as of the session, and the placement of the
   agents of one stage is decided together where the shared element is
   large, because duplicating it across instances costs as much as the
   transfer that the shared placement avoids.

13.10.  Long-Horizon Requirements

   LH1. A CATS system SHOULD represent a long-horizon session as a
   sequence of stages, and SHOULD be able to state, for each stage, the
   elements and the compute capabilities that the stage needs.

   LH2. A CATS system SHOULD keep the placement of a session stable
   across the steps of a stage, and SHOULD change it at a stage boundary
   rather than within a stage, unless a failure or a violated constraint
   forces the change.

   LH3. A CATS system SHOULD maintain the affinity of a session across a
   pause, and SHOULD carry a retention commitment that covers the
   expected idle interval where the cost of losing the state exceeds the
   cost of holding it.

   LH4. A CATS system MUST re-verify the presence of the elements that a
   session depends on when the session resumes, and MUST treat an
   element that is no longer present as absent.

   LH5. A CATS system SHOULD be able to checkpoint the durable part of a
   working set, so that a session can be resumed on another instance
   after a failure, an eviction, or an administrative move.

   LH6. A CATS system SHOULD state the part of a working set that is
   sufficient to continue a session, so that a resume reconstructs only
   the material that the session needs.

   LH7. A CATS system SHOULD be able to establish the affinity of a
   stage before the stage starts, including the prefetch of the
   artifacts and the elements that the next stage needs.

   LH8. A CATS system SHOULD detect affinity decay, that is, a sustained
   reduction of the observed gain of the current placement, and SHOULD
   re-evaluate the placement when that reduction exceeds the tolerance
   of the session.

Mo, et al.               Expires 3 April 2027                  [Page 36]
Internet-Draft           Agent Affinity                   September 2026

   LH9. A CATS system SHOULD pace the cost of affinity maintenance over
   the remaining horizon of a session, and SHOULD NOT spend the budget
   of the remaining stages to preserve the state of the current one.

   LH10. A CATS system SHOULD attribute the cost of a stage transition
   to the stage that benefits from it, and SHOULD attribute a transfer
   that serves several remaining stages to the session rather than to a
   single step.

   LH11. A CATS system SHOULD be able to recover the affinity of a
   session after the failure of the instance that holds its state, by
   falling back to a checkpoint or to a reconstruction path, and SHOULD
   record which path was taken.

   LH12. A CATS system SHOULD decide the placement of a multi-agent
   stage that shares a working set as one decision, and SHOULD account
   for the duplication when the agents of a stage are placed on
   different instances.

14.  Affinity for Similar Agents and for Tenants

14.1.  Similarity and Reuse Groups

   Two agents are similar for the purpose of this document when their
   steps require the same reusable elements, and not when they are the
   same software or serve the same user. The reusable elements of
   Section 6 are frequently shared: a system prompt prefix is common to
   every session of an agent service, a tool schema is common to every
   agent that uses the tool, a retrieval corpus is common to a tenant,
   and a model artifact is common to every step of every session that
   runs on it.

   Affinity for a group has a different economics from affinity for a
   single session. The cost of establishing an element is paid once and
   is amortized over the reuse group, so an element whose
   re-establishment is expensive is worth keeping even when the
   individual session that caused it has ended. The risk is also
   different: the group is what makes a position on one instance
   popular, and it is what makes the isolation and fairness rules of the
   following subsections necessary.

14.2.  Elements That Can Be Shared

Mo, et al.               Expires 3 April 2027                  [Page 37]
Internet-Draft           Agent Affinity                   September 2026

      Reusable element       | Reuse group         | Precondition
      -----------------------+---------------------+--------------
      System prompt prefix   | agents of one agent | same model
      state                  | service             | and tokenizer
      Tool schema and        | agents of one tool  | same template
      templates              | set                 | version
      Retrieved corpus and   | tenant or user      | same rights
      index                  |                     | and version
      Embeddings of shared   | tenant              | same
      documents              |                     | embedding
                             |                     | model
      Model weights and      | site or capability  | same
      adapter                | tier                | precision and
                             |                     | revision
      Policy and evaluation  | agent service       | same policy
      state                  |                     | version
      Tenant memory and plan | tenant only         | no
      state                  |                     | cross-tenant
                             |                     | scope

   The conditions of reuse in the third column are what make a group a
   reuse group: an element may be reused by the members of the group
   only when the conditions hold at the candidate that holds it. A
   prefix state produced by one model revision is not reusable by a step
   that runs another, even when the prompt text is identical, because
   the state is not a copy of the prompt.

14.3.  Conditions for Safe Reuse

   Shared reuse is safe when the following conditions are met, and each
   of them is a condition that the holder evaluates rather than one that
   the requester asserts.

   *  Identity of the computation: the model revision, tokenizer,
      precision, runtime, and adapter that produced the element match those
        that the reusing step requires.

   *  Identity of the material: the element is the same version of the
      same material, established by a defined equivalence (Section 10) rather
        than by a name that two producers may use for different content.

   *  Authorization: the reuse group of the element includes the session
      that requests the reuse, and the holder enforces that membership.

Mo, et al.               Expires 3 April 2027                  [Page 38]
Internet-Draft           Agent Affinity                   September 2026

   *  Isolation: the reuse does not disclose the content of one session
      to another, and does not disclose the presence of one session's
        state to another through the timing of a hit.

   *  Currency: the element has not been superseded by a version that
      the reusing step assumes, and is not under a pending eviction or
        replacement that would make it unusable during the step.

   The fourth condition is the one that is most easily lost in design,
   because sharing is implemented for efficiency and its observability
   consequences are considered later. A shared hit that is faster than a
   miss is an observable signal, and in a multi-tenant system it is a
   signal about another tenant's activity if the sharing is not bounded.

14.4.  Tenant Affinity Domains and Isolation

   A tenant is the unit at which placement policy, capacity expectation,
   and isolation are usually expressed. Three properties apply to the
   affinity of a tenant.

   *  A tenant has an affinity domain, that is, the set of instances at
      which its sessions and its state may be placed. The domain is a
        constraint: a candidate outside the domain is ineligible for that
        tenant's sessions and for the state that belongs to the tenant, 
        regardless of the affinity gain that it would produce.

   *  The scope of a reusable element is part of the element, not a
      property of the requester. An element that a tenant holds is reusable
        by the sessions of that tenant within the domain, and it is not reusable
        outside it, even when the content would be identical for both
        tenants.

   *  The affinity of one tenant is not visible to another. A tenant may
        observe its own affinity, and the operator may observe the aggregate,
        but neither may observe the presence or the reuse of another tenant's
        state.

14.5.  Fairness, Herding, and Quota

   Affinity concentrates demand, and concentration has to be bounded by
   rules that are stated in advance rather than applied when the
   concentration has already degraded the service.
Mo, et al.               Expires 3 April 2027                  [Page 39]
Internet-Draft           Agent Affinity                   September 2026

   *  Herding. Popular reusable elements attract sessions, and the
      attraction is self-reinforcing, because the sessions that are steered
        to the holder make the element more valuable there. A selection that
        applies affinity without a bound therefore produces a distribution that
        is worse than the one it started from in the tail, even when the mean
        improves.

   *  Duplication. When the concentration of a reuse group on one
      instance degrades the objective of the sessions that share the element, 
        the remedy is to admit a second copy of the element at another instance
        and to split the group, at the cost of establishing the element
        twice. The decision to duplicate is a comparison of the two costs,
        not a default.

   *  Quota. A tenant with a large state footprint can occupy the
      capacity

   that other tenants need. The affinity of a tenant is therefore
   bounded by a quota on the capacity that its state may occupy and by a
   share of the reuse capacity of a shared instance, and the binding of
   the quota is observable rather than silent.

   *  Fairness of measurement. The affinity gain of one tenant is not
      allowed to be achieved by degrading the tail of another, which requires
        that the objective of a session be evaluated per tenant and per
        communication mode rather than over the aggregate.

14.6.  Measurement and Observability across Tenants

   The affinity metrics of Section 11 are reported per tenant as well as
   per key and per site, and the following restrictions apply to them.

   *  A tenant may see its own footprint, its own hit ratio, and the
      binding of its own quota.

   *  The operator may see the aggregate footprint, the concentration on
      an instance, and the refused reuse, because those are the quantities
        that the fairness rules bind to.

Mo, et al.               Expires 3 April 2027                  [Page 40]
Internet-Draft           Agent Affinity                   September 2026

   *  Neither may see the per-session affinity of another tenant, and
      neither may infer it from a difference in latency, so the counters
        that are exposed are aggregated over the population of the tenant
        rather than over a key that another tenant could probe.

14.7.  Requirements for Similar Agents and Tenants

   ST1. A CATS system SHOULD be able to identify the group of sessions,
   agents, or tenants that may reuse a given working set element, and
   SHOULD be able to express that group as part of the state
   information.

   ST2. A CATS system SHOULD be able to determine whether two agents may
   share an element by comparing the conditions of reuse, including the
   model revision, the tokenizer, the runtime, the precision, the
   template version, and the policy version, rather than by comparing
   their content or their identity.

   ST3. A CATS system SHOULD treat a shared element as available for a
   step only when the conditions of reuse are satisfied at the candidate
   that holds it.

   ST4. A CATS system SHOULD be able to express the affinity domain of a
   tenant, that is, the set of instances at which the sessions and the
   state of the tenant may be placed, and MUST enforce that domain as a
   constraint.

   ST5. A CATS system MUST NOT reuse a working set element across
   tenants, or across sessions of different users within a tenant,
   unless the policy of the holder of the element permits that reuse.

   ST6. A CATS system SHOULD account for the cost of a shared element
   once for the reuse group that establishes it, and SHOULD distinguish
   a shared reuse from a private one in its measurements.

   ST7. A CATS system SHOULD bound the capacity that the state of one
   tenant may occupy at a shared instance, and SHOULD make the binding
   of that bound observable.

   ST8. A CATS system SHOULD detect the concentration of a reuse group
   on a small number of instances, and SHOULD be able to admit an
   additional copy of an element when that concentration degrades the
   objective of the sessions that share the element.

   ST9. A CATS system SHOULD be able to state, per tenant, whether
   affinity may be traded against load, distance, or cost, and SHOULD
   apply the affinity policy of the tenant to the sessions of that
   tenant.

Mo, et al.               Expires 3 April 2027                  [Page 41]
Internet-Draft           Agent Affinity                   September 2026

   ST10. A CATS system SHOULD measure the affinity of a tenant without
   disclosing the affinity of an individual session to another tenant,
   and SHOULD NOT allow a tenant to observe the presence of another
   tenant's state.

   ST11. A CATS system SHOULD apply the isolation requirements of the
   tenant boundary to a shared element as well as to a private one,
   including to the timing of a hit and a miss.

   ST12. A CATS system SHOULD be able to revoke the reuse scope of an
   element when a tenant or a policy changes, without requiring the
   content of the element to be disclosed.

15.  Interaction with Existing Work

   Metrics: The state, storage, and compute information described here
   is intended to appear as an additional category or categories in the
   metric framework of [I-D.ietf-cats-metric-definition], alongside the
   existing computing, communication, and service categories. In the
   terms used there, state availability, residency tier, compute
   readiness, and retention are naturally raw (Level 0) metrics, while
   the comparison of transfer against recomputation, affinity gain, and
   the concentration of affinity are derived (Level 1) quantities.

   Data model: If affinity information is to be configured or monitored,
   the data model [I-D.ietf-cats-data-model] needs to accommodate it.
   The minimum set of objects is a state handle or key space, a
   residency tier, a readiness descriptor for a computation, a reuse
   scope, a retention commitment, and a policy that states the permitted
   scope of reuse.

   OAM: The operational indicators for affinity-based selection are the
   hit ratio of state reuse, the volume and the latency of the state
   that a decision transfers, the affinity gain, the concentration of
   sessions on the instances that hold popular state, and the number of
   selections changed because state was not available where it was
   expected. These are candidates for the monitoring functions of
   [I-D.ietf-cats-oam-fw].

   Selection mapping: The mapping described in
   [I-D.mo-cats-agent-selection-mapping] provides the framework in which
   the requirements of Sections 8 to 14 are applied: the descriptor of a
   step, the resource view of a candidate, the decision points, and the
   stability requirement are defined there, and this document adds the
   state, storage, and compute affinity content of each.

Mo, et al.               Expires 3 April 2027                  [Page 42]
Internet-Draft           Agent Affinity                   September 2026

   Distributed cache and storage protocols: Moving working set elements
   between instances is a data transfer problem that is expected to use
   existing transport building blocks. This document does not define a
   transfer protocol, and multiple such protocols may be used in one
   deployment. The reuse condition of Section 10 is the same condition
   that a cache applies when it decides whether a stored representation
   may be reused [RFC9111], and cache-control semantics are a useful
   model for the retention commitments of Section 12.

15.1.  Traceability to the Agent Service Requirements

   The requirements of [I-D.mo-cats-agent-service-characteristics] are
   addressed as follows. The last two rows map the characteristics of
   that document that this document develops further.

   *  R1, R2: WM1, WM2, WM6.

   *  R8, R9, R10: S13, S14, WM5.

   *  R11: S1, S2, S4.

   *  R12: S6, S7, WM3.

   *  R13: S8, S9, WM6.

   *  R14: S15, S16, AM1, ST1, ST3.

   *  R15, R16: AM6, LH9, LH10.

   *  R17: MA6.

   *  R19: AM3, AM4, LH5, LH6.

   *  R20: AM4, WM8.

   *  R21: MA4, MA5, MA9.

   *  R22: MA8, MA10.

   *  Long-horizon state and memory persistence (Section 5.8 of that
      document): LH1 to LH12, AM2, AM8.

   *  Locality, governance, and tenancy (Section 5.9 of that document):
      S8, S9, ST4, ST5, ST11, ST12.

16.  Operational Considerations

Mo, et al.               Expires 3 April 2027                  [Page 43]
Internet-Draft           Agent Affinity                   September 2026

   State-based selection changes the shape of the traffic that an
   operator sees. State transfer is additional traffic between service
   sites, it can be large, and it can be triggered by a steering
   decision, which means that a selection that ignores transfer cost can
   create the congestion that then degrades the next selection.
   Operators are therefore expected to bound the fraction of decisions
   that are allowed to trigger a transfer, and to treat the transfer
   network as a resource that the selection function must account for.

   Affinity can also concentrate load and can increase the impact of a
   failure: the instance that holds the state of many sessions is a
   higher-value target and a higher-impact failure than an instance that
   holds none. Deployments are expected to decide how much state is
   worth keeping warm, and to have a fallback path when the preferred
   instance is unavailable, which is consistent with the notion of a
   fallback decision [I-D.pang-cats-fallback-decision-framework].

   Five operational properties follow from the mechanisms of Section 12.

   *  The retention policy of an instance is a capacity decision, not
      only a performance decision, because a commitment that is not released
        withholds capacity from the steps that are running.

   *  The override rule has to be testable in advance. An operator is
      expected to know at which measured concentration the override will
        be applied, rather than discovering it during an incident.

   *  The failure of the holder of a large amount of state is a
      common-mode event for every session that it holds. A deployment is
        expected to bound the amount of state that one instance holds for
        one tenant, and to keep the recovery path (checkpoint or reconstruction)
        exercised.

   *  Affinity metrics have to be attributable to a decision, because a
      hit ratio that cannot be attributed to a placement cannot be used to
        correct the placement.

   *  The duplication decision of Section 14 has a cost that appears in
      the storage metrics and a benefit that appears in the tail of the
        objectives. An operator is expected to state which of the two is
        bounded, so that the decision can be taken consistently.

Mo, et al.               Expires 3 April 2027                  [Page 44]
Internet-Draft           Agent Affinity                   September 2026

17.  Security Considerations

   State is more sensitive than capacity. Exposing which elements an
   instance holds reveals patterns of use even when the content is not
   exposed, and the ability to request a transfer is the ability to move
   data between locations. The following considerations follow.

   Content protection: State that is transferred between instances
   should be protected in the same way as the session data from which it
   is derived, including at rest where the tier is persistent.

   Handle confidentiality and unforgeability: A state handle that can be
   guessed or forged is a handle that can be used to probe for the
   existence of state, and possession of a handle MUST NOT grant access
   to the state. Handles should be unguessable, scoped, and validated
   against authorization before use. A handle that names a shared
   element is scoped to the reuse group of that element, so that
   knowledge of it does not disclose the existence of state outside the
   group.

   Cross-tenant leakage: Reuse of state across sessions or tenants
   creates a direct path to information disclosure, including through
   the timing of a hit or a miss. Sharing scope should be enforced by
   the entity that holds the state, not only by the entity that requests
   it.

   Poisoning: A participant that can cause incorrect state to be
   associated with a valid handle can influence the output of the
   sessions that reuse it. Integrity of the association between a handle
   and its content should be established by the holder of the state.

   Denial through eviction: Because state availability is a resource
   that can be exhausted, a participant that can cause eviction can
   degrade other sessions. Implementations should not allow one session
   to evict the state of another without authorization. Affinity widens
   this surface, because an element that is shared by a reuse group is a
   single object whose eviction degrades the whole group, and because
   the retention commitments of one session compete with the state of
   the others.

   Compute substitution: A selection that treats capability tiers as
   interchangeable can place a step on a candidate that is cheaper but
   less capable, which changes the result of the step rather than only
   its cost. The constraint families of Section 8.2 keep capability out
   of the affinity valuation for this reason.

18.  Privacy Considerations

Mo, et al.               Expires 3 April 2027                  [Page 45]
Internet-Draft           Agent Affinity                   September 2026

   Working set elements are derived from user inputs and therefore
   inherit their sensitivity. Even aggregated state availability
   information can reveal activity, and transfer events reveal which
   sessions are active between which sites. Where such information is
   exposed to the network it should be minimized, aggregated where
   possible, and retained only as long as it serves the selection
   function.

   Locality constraints are frequently imposed because of legal or
   regulatory requirements, and where a constraint is expressed in the
   information exchanged, the accuracy of that expression determines
   whether the requirement is met. Implementations SHOULD treat an
   unknown constraint as a prohibition rather than as permission.

   Shared reuse adds two considerations. First, a reuse group is an
   inference surface: a party that can observe the hit ratio of a shared
   element learns how many sessions use it and when, even when it learns
   nothing about their content. Second, the revocation of a reuse scope
   is a privacy control, and it is expected to be effective for the
   state that is already held and not only for the state that is
   subsequently created. The metrics of Section 11 are therefore
   reported per tenant and aggregated over a population, and not per key
   where a key could be probed by another tenant.

19.  IANA Considerations

   This document has no IANA actions.

20.  Normative References

   [I-D.ietf-cats-framework] Li, C., Du, Z., Boucadair, M., Contreras,
   L. M., et al., "A Framework for Computing-Aware Traffic Steering
   (CATS)", Work in Progress, Internet-Draft,
   draft-ietf-cats-framework-24, September 2026.

   [I-D.ietf-cats-metric-definition] Yao, K., et al., "CATS Metrics
   Definition", Work in Progress, Internet-Draft,
   draft-ietf-cats-metric-definition-12, September 2026.

   [I-D.mo-cats-agent-service-characteristics] Mo, Y., Yang, D., Zhou,
   C., "AI Agent Service Characteristics and Their Implications for
   Computing-Aware Traffic Steering", Work in Progress, Internet-Draft,
   draft-mo-cats-agent-service-characteristics-00, September 2026.

21.  Informative References

   [CATS-CHARTER] IETF, "Computing-Aware Traffic Steering (CATS) Working
   Group Charter",
   <https://datatracker.ietf.org/doc/charter-ietf-cats/>.

Mo, et al.               Expires 3 April 2027                  [Page 46]
Internet-Draft           Agent Affinity                   September 2026

   [I-D.ietf-cats-usecases-requirements] Yao, K., et al.,
   "Computing-Aware Traffic Steering (CATS) Problem Statement, Use
   Cases, and Requirements", Work in Progress, Internet-Draft,
   draft-ietf-cats-usecases-requirements-14, September 2026.

   [I-D.mo-cats-agent-selection-mapping] Mo, Y., Yang, D., Zhou, C., "A
   Selection Mapping Framework for AI Agent Services in Computing-Aware
   Traffic Steering", Work in Progress, Internet-Draft,
   draft-mo-cats-agent-selection-mapping-00, September 2026.

   [I-D.ietf-cats-data-model] Yao, H., Lin, C., et al., "Data Model for
   Computing-Aware Traffic Steering (CATS)", Work in Progress,
   draft-ietf-cats-data-model-00, September 2026.

   [I-D.ietf-cats-oam-fw] Fu, H., Xiong, Q., Du, Z., et al., "Computing-
   Aware Traffic Steering (CATS) Operations, Administration, and
   Maintenance (OAM) Framework", Work in Progress, Internet-Draft,
   draft-ietf-cats-oam-fw-01, July 2026.

   [I-D.li-cats-kv-cache-distribution] Li, Z., et al., "KV Cache
   Distribution for Distributed LLM Inference: Use Case and
   Requirements", Work in Progress,
   draft-li-cats-kv-cache-distribution-00, July 2026.

   [I-D.pang-cats-fallback-decision-framework] Pang, R., Ed., Han, M.,
   Ed., Huang, T., Ed., "CATS Fallback Decision Framework", Work in
   Progress, Internet-Draft,
   draft-pang-cats-fallback-decision-framework-00, July 2026.

   [I-D.zhang-cats-token-aware-ts] Zhang, N., Ed., Han, M., Ed., Yi, X.,
   Ed., "A token-aware traffic steering solution for agent service",
   Work in Progress, draft-zhang-cats-token-aware-ts-00, March 2026.

   [I-D.zhu-cats-metric-semantics] Zhu, M., "Operational Semantics for
   CATS Metric Consumption", Work in Progress,
   draft-zhu-cats-metric-semantics-01, August 2026.

   [RFC2119] Bradner, S., "Key words for use in RFCs to Indicate
   Requirement Levels", BCP 14, RFC 2119, DOI 10.17487/RFC2119, March
   1997, <https://www.rfc-editor.org/info/rfc2119>.

   [RFC8174] Leiba, B., "Ambiguity of Uppercase vs Lowercase in RFC 2119
   Key Words", BCP 14, RFC 8174, DOI 10.17487/RFC8174, May 2017,
   <https://www.rfc-editor.org/info/rfc8174>.

   [RFC9111] Fielding, R., Nottingham, M., Reschke, J., "HTTP Caching",
   STD 98, RFC 9111, DOI 10.17487/RFC9111, June 2022,
   <https://www.rfc-editor.org/info/rfc9111>.
Acknowledgments

Mo, et al.               Expires 3 April 2027                  [Page 47]
Internet-Draft           Agent Affinity                   September 2026

   The authors would like to thank the participants of the CATS working
   group for the discussions that shaped this document.

Authors' Addresses

   Y. Mo
   Huazhong University of Science and Technology
   Email: moyj@hust.edu.cn

   D. Yang
   Huazhong University of Science and Technology
   Email: d202581903@hust.edu.cn

   C. Zhou
   Huazhong University of Science and Technology
   Email: m202474228@hust.edu.cn

Mo, et al.               Expires 3 April 2027                  [Page 48]