Skip to main content

A Selection Mapping Framework for AI Agent Services in Computing-Aware Traffic Steering
draft-mo-cats-agent-selection-mapping-00

Document Type Active Internet-Draft (individual)
Authors Yijun Mo , Dewei Yang , 周诚宇
Last updated 2026-09-29
RFC stream (None)
Intended RFC status (None)
Formats
Stream Stream state (No stream defined)
Consensus boilerplate Unknown
RFC Editor Note (None)
IESG IESG state I-D Exists
Telechat date (None)
Responsible AD (None)
Send notices to (None)
draft-mo-cats-agent-selection-mapping-00
CATS                                                               Y. Mo
Internet-Draft                                                   D. Yang
Intended status: Informational                                   C. Zhou
Expires: 3 April 2027      Huazhong University of Science and Technology
                                                       30 September 2026

         A Selection Mapping Framework for AI Agent Services in
                    Computing-Aware Traffic Steering

                draft-mo-cats-agent-selection-mapping-00

Abstract

   The Computing-Aware Traffic Steering (CATS) framework selects a
   service contact instance for a service request by combining computing
   and network metrics that are distributed by CATS Service Metric
   Agents and CATS Network Metric Agents. For AI agent services, the
   request that arrives at the network is not a single unit of work: it
   is a session that expands into multiple steps, each of which may
   require a different capability, a different state, and a different
   path. This document describes a mapping framework that turns the
   characteristics of an agent step into (1) a set of hard constraints
   and (2) a per-dimension valuation over a three-dimensional resource
   view composed of forwarding, computing, and storage, and that feeds
   the resulting suitability of each candidate service contact instance
   into the existing CATS selection function. The framework introduces
   no new functional component.

Status of This Memo

   This Internet-Draft is submitted in full conformance with the
   provisions of BCP 78 and BCP 79.

   Internet-Drafts are working documents of the Internet Engineering
   Task Force (IETF).  Note that other groups may also distribute
   working documents as Internet-Drafts.  The list of current Internet-
   Drafts is at https://datatracker.ietf.org/drafts/current/.

   Internet-Drafts are draft documents valid for a maximum of six months
   and may be updated, replaced, or obsoleted by other documents at any
   time.  It is inappropriate to use Internet-Drafts as reference
   material or to cite them other than as "work in progress."

   This Internet-Draft will expire on 3 April 2027.

Copyright Notice

Mo, et al.               Expires 3 April 2027                   [Page 1]
Internet-Draft           Agent Selection Mapping          September 2026

   Copyright (c) 2026 IETF Trust and the persons identified as the
   document authors.  All rights reserved.

   This document is subject to BCP 78 and the IETF Trust's Legal
   Provisions Relating to IETF Documents (https://trustee.ietf.org/
   license-info) in effect on the date of publication of this document.
   Please review these documents carefully, as they describe your rights
   and restrictions with respect to this document.  Code Components
   extracted from this document must include Revised BSD License text as
   described in Section 4.e of the Trust Legal Provisions and are
   provided without warranty as described in the Revised BSD License.

Discussion Venues

   Discussion of this document takes place on the CATS Working Group
   mailing list (cats@ietf.org), which is archived at
   https://mailarchive.ietf.org/arch/browse/cats/.

Table of Contents

Mo, et al.               Expires 3 April 2027                   [Page 2]
Internet-Draft           Agent Selection Mapping          September 2026

   1.  Introduction  . . . . . . . . . . . . . . . . . . . . . . . . . 4
   2.  Conventions and Definitions  . . . . . . . . . . . . . . . . . .5
   3.  Terminology  . . . . . . . . . . . . . . . . . . . . . . . . . .5
   4.  Problem Statement  . . . . . . . . . . . . . . . . . . . . . . .7
   5.  Agent Service Properties and Their Selection Implications  . . .8
      5.1.  Long-Horizon Sessions  . . . . . . . . . . . . . . . . . . 8
      5.2.  Stateful Execution  . . . . . . . . . . . . . . . . . . . .9
      5.3.  Discrete Progress  . . . . . . . . . . . . . . . . . . . . 9
      5.4.  Heavy-Tailed Consumption  . . . . . . . . . . . . . . . . .9
   6.  Selection Mapping Framework  . . . . . . . . . . . . . . . . . .9
      6.1.  Framework Entities and Information Flow  . . . . . . . . . 9
      6.2.  Layer 1: Agent Service Descriptor  . . . . . . . . . . . .10
      6.3.  Layer 2: Resource View  . . . . . . . . . . . . . . . . . 12
      6.4.  Layer 3: Selection Context  . . . . . . . . . . . . . . . 13
      6.5.  Decision Artifacts  . . . . . . . . . . . . . . . . . . . 14
      6.6.  Characteristic-to-Resource Mapping  . . . . . . . . . . . 14
      6.7.  Characteristic-to-Resource Routing  . . . . . . . . . . . 18
      6.8.  Resource-Level Routing Recipe  . . . . . . . . . . . . . .21
   7.  Information Elements and Their Abstract Syntax  . . . . . . . .22
      7.1.  Descriptor Elements  . . . . . . . . . . . . . . . . . . .23
      7.2.  Resource View Elements  . . . . . . . . . . . . . . . . . 24
      7.3.  Accounting Elements  . . . . . . . . . . . . . . . . . . .25
      7.4.  Exchange Metadata Elements  . . . . . . . . . . . . . . . 25
      7.5.  Semantics Rules Common to the Elements  . . . . . . . . . 26
   8.  Selection Procedure  . . . . . . . . . . . . . . . . . . . . . 27
      8.1.  Phases  . . . . . . . . . . . . . . . . . . . . . . . . . 27
      8.2.  Decision Points  . . . . . . . . . . . . . . . . . . . . .29
      8.3.  Re-Evaluation Triggers and Stability  . . . . . . . . . . 29
      8.4.  Fallback Ladder  . . . . . . . . . . . . . . . . . . . . .30
   9.  Message Semantics  . . . . . . . . . . . . . . . . . . . . . . 32
      9.1.  Exchange Catalogue  . . . . . . . . . . . . . . . . . . . 32
      9.2.  Exchange Sequences  . . . . . . . . . . . . . . . . . . . 34
      9.3.  Rules Common to the Exchanges  . . . . . . . . . . . . . .35
   10.  Aggregation, Normalization, and Freshness  . . . . . . . . . .36
   11.  Requirements  . . . . . . . . . . . . . . . . . . . . . . . . 36
      11.1.  Framework and Descriptor Requirements  . . . . . . . . . 37
      11.2.  State and Storage Requirements  . . . . . . . . . . . . .38
      11.3.  Discrete Decision Requirements  . . . . . . . . . . . . .39
      11.4.  Tail and Long-Horizon Requirements  . . . . . . . . . . .39
      11.5.  Exchange Requirements  . . . . . . . . . . . . . . . . . 40
      11.6.  Mapping Requirements  . . . . . . . . . . . . . . . . . .40
   12.  Relationship to the CATS Framework  . . . . . . . . . . . . . 41
      12.1.  Traceability to the Agent Service Requirements  . . . . .41
   13.  Operational Considerations  . . . . . . . . . . . . . . . . . 42
   14.  Security Considerations  . . . . . . . . . . . . . . . . . . .43
   15.  Privacy Considerations  . . . . . . . . . . . . . . . . . . . 43
   16.  IANA Considerations  . . . . . . . . . . . . . . . . . . . . .44
   17.  Normative References  . . . . . . . . . . . . . . . . . . . . 44
   18.  Informative References  . . . . . . . . . . . . . . . . . . . 44
   Acknowledgments  . . . . . . . . . . . . . . . . . . . . . . . . . 45
   Authors' Addresses  . . . . . . . . . . . . . . . . . . . . . . . .45
Mo, et al.               Expires 3 April 2027                   [Page 3]
Internet-Draft           Agent Selection Mapping          September 2026

1.  Introduction

   A CATS system classifies the traffic of a service request, selects a
   service contact instance, and steers the traffic of the request
   towards the selected instance [I-D.ietf-cats-framework]. The
   selection is made by the CATS Path Selector (C-PS) using computing
   metrics collected by a CATS Service Metric Agent (C-SMA) and network
   metrics collected by a CATS Network Metric Agent (C-NMA). The metrics
   that are exchanged for that purpose are specified in
   [I-D.ietf-cats-metric-definition] at three levels of abstraction.

   The use cases and requirements document already anticipates
   distributed AI training and inference as a use case of CATS, and
   notes that the resources that matter for inference include processor
   cores and the memory used for cache [I-D.ietf-cats-usecases-requirements].
   For agent services, two properties of that model need to be extended
   without changing it. First, the unit that arrives at the network is a
   session or a step within a session, and the requirements of a step
   are richer than a request for a service: they include a capability
   constraint, a state affinity, a budget, and a locality constraint.
   Second, the state that determines how quickly a step can be served is
   not part of the current metric view, as has been observed for the 
   specific case of a key-value (KV) cache
   [I-D.li-cats-kv-cache-distribution].

   This document describes how the characteristics of an agent step
   [I-D.mo-cats-agent-service-characteristics] are mapped onto a
   selection decision over three dimensions -- forwarding, computing,
   and storage -- using the existing CATS components. The mapping
   framework is intended to be used as follows:

   *  as input to a future extension of the metric definition, by
      identifying the information that the mapping needs;

   *  as input to the data model work [I-D.ietf-cats-data-model], by
      identifying the objects that configuration and monitoring must cover;

   *  as a description of a selection procedure that an implementation
      can follow today using metrics that are already defined.

Mo, et al.               Expires 3 April 2027                   [Page 4]
Internet-Draft           Agent Selection Mapping          September 2026

   The mapping is written for four properties of agent services. an 
   agent session is long-horizon, stateful, discrete in the way that
   it makes progress, and heavy-tailed in what it consumes. Section 5
   states what each property requires of a selection procedure. Sections
   6, 7, 8, and 9 give the framework, the information elements and 
   their abstract syntax, the procedure with its decision points, and
   the exchanges through which the information reaches the selection
   function. Section 11 collects the requirements.

   The intended standing of this document is informational groundwork in
   the sense of the CATS charter [CATS-CHARTER]: it states what a CATS
   system has to be able to express, so that the work can be taken up by
   the metric definition and the data model rather than by a protocol
   extension. The resource inventory of Section 6.7 is what a metric
   framework would have to expose, and the elements of Section 7 are
   what a data model would have to accommodate. This document defines no
   encoding and no protocol.

2.  Conventions and Definitions

   The key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT",
   "SHOULD", "SHOULD NOT", "RECOMMENDED", "NOT RECOMMENDED", "MAY", and
   "OPTIONAL" in this document are to be interpreted as described in BCP
   14 [RFC2119] [RFC8174] when, and only when, they appear in all
   capitals, as shown here.

   In this document these key words are used to state requirements on
   the design of a CATS system that supports agent services. They do not
   describe protocol behavior, and this document defines no protocol,
   message format, or data model.

3.  Terminology

   This document uses the terms of [I-D.ietf-cats-framework] and
   [I-D.mo-cats-agent-service-characteristics]. The following additional
   terms are used.

   Agent Service Descriptor: The representation of what a step of an
   agent session needs, expressed as a set of constraints and
   preferences.

Mo, et al.               Expires 3 April 2027                   [Page 5]
Internet-Draft           Agent Selection Mapping          September 2026

   Resource View: The set of properties of a candidate instance and of
   the paths towards it, grouped into the forwarding, computing, and
   storage dimensions, that are used to evaluate a step.

   Affinity: A preference for a candidate instance on the grounds that
   it already holds session state, model artifacts, or other reusable
   material.

   Selection Point: The point in the execution of a session at which a
   selection is made, such as the start of a session, the start of a
   turn, the start of a step, or a re-evaluation triggered by a change
   in the resource view.

   Decision Point: A point in the life of a session at which the mapping
   is required to produce a selection, such as session admission, the
   start of a turn, a step boundary, the detection of a tail event, or a
   change of the resource view that invalidates the current selection.

   Decision Record: The artifact that records the descriptor, the
   resource view, the trigger, and the outcome of one application of the
   procedure, for audit, for accounting, and as input to later decisions
   of the same session.

   Working Set Class: A category of reusable state with a distinct cost
   of reconstruction, such as model-side context state, retrieved
   material, tool results, plan and scratch state, and longer-term
   memory.

   State Handle: An opaque reference to a working set element that
   allows a component to test presence, request a transfer, or attach
   affinity to the element without learning its content.

   Reuse Key: The identifier under which a working set element can be
   addressed for the purpose of presence testing, transfer, or reuse.

   Residency Tier: The memory tier at which a working set element
   currently resides at an instance, for example accelerator memory,
   host memory, local storage, or a remote store.

Mo, et al.               Expires 3 April 2027                   [Page 6]
Internet-Draft           Agent Selection Mapping          September 2026

   Stability Window: The minimum interval, or the minimum improvement of
   suitability, that must hold before a different instance may be
   selected for a session that is already placed.

   Commitment: The binding of a step, or of a range of steps, to a
   selected instance, together with the state and the budget that the
   binding assumes.

   Handoff: The planned move of a session, or of part of its working
   set, from one instance to another, including reservation, transfer,
   and commit or abort.

   Tail Event: An execution whose cost or duration lies in the upper
   tail of the distribution for comparable steps, for example a tool
   call that takes much longer than the same tool normally takes.

   Fallback Ladder: The ordered set of behaviors to which the procedure
   degrades when no candidate satisfies all constraints, when the
   selected instance fails, or when a tail event exhausts the budget of
   a step.

4.  Problem Statement

   The mapping exists because the current selection input and the
   current demand do not align.

   First, the mapping from demand to metrics is implicit. Today the C-PS
   selects on metrics that describe instances and paths, but the
   relationship between what a step needs and those metrics is not
   modeled: for a step that must reason over a 200 kB context, must not
   send that context outside a region, and must complete within two
   seconds, it is not obvious which metric at which level expresses the
   need, or which part of the need is a constraint rather than a
   preference.

   Second, a single scalar cannot represent the decision. A normalized
   score ranks candidates, but it cannot express that an instance is
   ineligible because it cannot serve the required capability, nor that
   two eligible instances differ in whether they already hold the
   session state. Mixing a hard constraint into a score either violates
   the constraint or distorts the ranking.

Mo, et al.               Expires 3 April 2027                   [Page 7]
Internet-Draft           Agent Selection Mapping          September 2026

   Third, the decision is discrete and its lifetime is not uniform,
   while the quantities that inform it are continuous and heavy-tailed.
   Placement changes slowly, because it may involve loading a model or
   migrating state, whereas traffic allocation and per-step routing
   change quickly [I-D.luan-cats-catpts], so a single validity rule for
   the whole decision is wrong. A session places a step, executes it,
   and then places the next one, so a mapping that reacts to every
   change of every quantity moves a long autonomous loop for gains that
   do not survive the move. At the same time the objective that matters
   is usually a percentile: an objective stated as a mean hides exactly
   the sessions that dominate cost and user-visible latency
   [I-D.mo-cats-agent-service-characteristics].

5.  Agent Service Properties and Their Selection Implications

   The four properties below are taken from the agent service
   characteristics [I-D.mo-cats-agent-service-characteristics]. Each
   constrains the mapping in a way that the current selection function
   does not face.

      Property     | Required of the mapping        | Where it is
                   |                                | handled
      -------------+--------------------------------+-------------------
      Long-horizon | Decide for a session of many   | 6.6 mapping; 6.4
                   | steps rather than for a single | commitment policy;
                   | request, and keep decisions    | 8.3 stability
                   | stable while it runs           |
      Stateful     | Treat reusable state as both a | 6.3 and 6.6; 8.1
                   | constraint and a quantity, and | and 9 handoff
                   | plan the move of that state    |
                   | before committing to it        |
      Discrete     | Decide at step boundaries on   | 6.6 mapping; 8.2
                   | banded values, and change an   | decision points;
                   | instance only when a           | 8.3 stability
                   | constraint requires it or the  |
                   | gain is durable                |
      Heavy-tailed | Evaluate objectives at high    | 6.6 mapping; 8.1
                   | percentiles, and keep the      | valuation; 8.4
                   | slowest call of a turn from    | fallback
                   | dictating the objective        |

5.1.  Long-Horizon Sessions

Mo, et al.               Expires 3 April 2027                   [Page 8]
Internet-Draft           Agent Selection Mapping          September 2026

   A session runs for tens or hundreds of steps and pauses for minutes
   while a tool runs, so the decision outlives the snapshot that
   produced it without being valid for the whole session: establishing a
   placement that loads a model artifact costs seconds to minutes, while
   a per-step routing decision is revised in milliseconds
   [I-D.mo-cats-agent-service-characteristics]. The descriptor therefore
   carries the session objective, the remaining budget, the remaining
   horizon, and the freshness need of each element, and a wake-up after
   a pause re-checks whether the state that the session depends on is
   still there.

5.2.  Stateful Execution

   State is a first-class input rather than an optimization. Its classes
   differ by orders of magnitude in size and in reconstruction cost, so
   evaluation is per element and uses the accumulated working set rather
   than the working set of the first step. Where a candidate does not
   report whether an element is present, the presence is unknown, and an
   unknown value does not satisfy a constraint.

5.3.  Discrete Progress

   The steps of a session are not known when it starts, and one step can
   fan out to several endpoints. The mapping therefore faces discrete
   decisions rather than a continuous allocation: it acts at step
   boundaries, compares banded values, and changes an instance only when
   a constraint requires it or the gain survives both the cost of the
   change and the stability requirement.

5.4.  Heavy-Tailed Consumption

   Work is skewed across sessions and the objective is usually a
   percentile, so the mapping needs tail-aware valuation, that is,
   comparison at a stated percentile over a stated window, and tail
   protection: detecting a running step that exceeds its expected cost
   and degrading it, or failing it with a report, in the order of the
   fallback ladder [I-D.pang-cats-fallback-decision-framework].

6.  Selection Mapping Framework

   The mapping framework has three layers.

6.1.  Framework Entities and Information Flow

   The mapping is performed by the components that already exist in the
   CATS framework. No new functional component is introduced; what the
   mapping adds is the information that flows between them and the times
   at which they act.

Mo, et al.               Expires 3 April 2027                   [Page 9]
Internet-Draft           Agent Selection Mapping          September 2026

      agent gateway
        |  descriptor: capability, working set, constraints, objective,
        |  budget, tail target; step boundary; wake-up after a pause
        v
         C-TC -- classify and attribute to session / turn / step-------+
        |                                                             |
        v                                                             |
         C-PS - derive descriptor -> apply constraints -> value dims---+
        |           ^                ^                   ^            |
        |           |                |                   |            |
        |      C-SMA: capability  C-SMA: state       C-NMA: paths     |
        |      load, tail         presence, tier     cost, prefixes   |
        v                                                             |
         forwarders <---- steering instruction-------------------------+
        |
        v
      decision record -> budget accounting -> next decision point

   *  Agent gateway: terminates the session and derives the descriptor
      where the network cannot; it supplies the descriptor, the step
      boundary, the wake-up after a pause, and the budget.

   *  C-TC: classifies session traffic, attributes it to the session,
      the turn, and the step, and carries the attribution and the
      descriptor keys.

   *  C-PS: applies constraints, values the dimensions, combines them,
      commits, and produces the selection and the decision record.

   *  C-SMA: reports capability, load, and state presence for its site,
      which is the computing and state view of a candidate.

   *  C-NMA: reports paths and the per-prefix cost towards endpoints,
      which is the forwarding view of a candidate.

   *  Forwarders: execute the steering decision for a step, and report
      the progress that indicates a tail event.

6.2.  Layer 1: Agent Service Descriptor

   The descriptor is derived from the step to be served and, where
   unavailable from the network, is provided by the agent gateway that
   terminates the session. It contains:

   *  the identity of the service, expressed as the CATS Service
      Identifier (CS-ID), and, where known, the session and step
      identities used for attribution;

   *  the capability required by the step, expressed in terms that can
      be matched against the capability offered by an instance, such as
      the model capability tier, the required context length, and the
      required tool or data reachability;
Mo, et al.               Expires 3 April 2027                  [Page 10]
Internet-Draft           Agent Selection Mapping          September 2026

   *  the working set that the step needs, expressed as the set of state
      elements that, if present, would allow the step to be served without
      re-computation, together with the keys under which those elements can
      be addressed;

   *  the tolerance of the step, expressed as whether its traffic may be
      buffered, delayed, or carried at reduced precision when the network
      is congested, in the terms used by the forwarding dimension;

   *  the constraints of the step, including locality and governance
      constraints that apply to the input data and to the working set, the
      tenant to which the session belongs, and any hard ceiling on cost or
      energy;

   *  the preferences of the step, including the latency objective, the
      willingness to trade precision or delay, and the relative importance
      of state reuse against load balancing.

   Beyond the elements above, the descriptor carries the quantities that
   let the mapping work over a long session rather than over a single
   request:

   *  the objective of the session and the horizon over which it is
      evaluated, where the gateway can estimate that horizon;

   *  the remaining budget of the session, and whether the budget is a
      ceiling that may not be exceeded or a target that may be missed
      at a stated cost;

   *  the tail objective of the step, expressed as a percentile, the
      window over which the percentile is taken, and the maximum
      probability of exceeding the objective that the session accepts;

   *  the degradation class of the step, that is, which reductions the
      step tolerates: lower precision, a lower capability tier, a truncated
      working set, or additional delay;

Mo, et al.               Expires 3 April 2027                  [Page 11]
Internet-Draft           Agent Selection Mapping          September 2026

   *  the freshness tolerance of the step, that is, how old a value may
      be for a decision on that step to remain valid.

   The descriptor is a local representation. Whether and how any element
   of it is carried in packets or signaled between components is
   explicitly out of scope for this document.

6.3.  Layer 2: Resource View

   The resource view is the three-dimensional description of a candidate
   instance and of the paths towards it. Each dimension is composed of
   constraints and of quantities.

      Dimension    | Constraint elements            | Quantity elements
      -------------+--------------------------------+-------------------
      Forwarding   | path policy, determinism       | latency, jitter,
                   | class, reachability of tools   | loss, bandwidth,
                   | and data sources, permitted    | path diversity,
                   | transfer window                | per-prefix cost to
                   |                                | tool and data
                   |                                | endpoints
      Computing    | capability tier, model and     | queueing delay,
                   | adapter versions, precision    | free accelerator
                   | support, maximum context       | capacity, memory
                   | length                         | headroom,
                   |                                | comparable-step
                   |                                | latency at median
                   |                                | and p95
      Storage      | permitted location per         | presence per reuse
                   | element, tenant isolation,     | key, residency
                   | sharing scope, retention       | tier, retrieval
                   | requirement                    | cost, recompute
                   |                                | cost, freshness,
                   |                                | element size

   The storage dimension is evaluated per working set element rather
   than per instance, because the elements of one working set differ in
   class, in size, and in the cost of reconstructing them. The classes
   that the mapping distinguishes are the following.

   *  Model-side context state: the state that allows earlier context to
      be reused without reprocessing it, for example a key-value cache or a
      cached prefix. It is large, it grows with the context, and it may be
      evicted during a pause.

Mo, et al.               Expires 3 April 2027                  [Page 12]
Internet-Draft           Agent Selection Mapping          September 2026

   *  Retrieved material: the items that a step has retrieved, for
      example documents or passages. It is small or moderate in size and
      cheap to re-retrieve while the retrieval endpoint is reachable.

   *  Tool results: the outputs of tool invocations. They are small, and
      re-obtaining one may be impossible because the tool is not idempotent,
      which makes their presence a hard requirement rather than a preference.

   *  Plan and scratch state: the intermediate reasoning and bookkeeping
      of the current task. It is small, but it is the most expensive class
      to reconstruct, because reconstruction requires the earlier steps to
      be re-executed.

   *  Longer-term memory: state that outlives the session, for example a

   profile or a set of preferences. It is small, it is read-mostly, and
   it is frequently subject to a location constraint.

   The forward, computing, and storage dimensions are evaluated for the
   same candidate instance. A candidate is a pair of a service contact
   instance (identified by the CATS Service Contact Instance ID,
   CSCI-ID) and the path (or set of paths) towards it.

6.4.  Layer 3: Selection Context

   The selection context records when and how often the mapping is
   applied. It contains the Selection Point, the objective of the
   session as a whole, the remaining budget of the session, and the
   instruction that determines whether a change of instance is
   permitted. This layer exists because the same descriptor evaluated at
   the start of a session and in the middle of a turn can produce
   different results, and because moving a session between instances has
   a cost that must be accounted for.

   The selection context is also where the discrete and the long-horizon
   properties meet the per-step decision. It therefore carries:

   *  the commitment policy, that is, which parts of the decision may
      change within a step, between the steps of a turn, and between
      turns, and which parts are held for the life of the session;

   *  the stability requirement in force, that is, the minimum interval
      and the minimum improvement of suitability that must hold before
      a different instance may be selected;

Mo, et al.               Expires 3 April 2027                  [Page 13]
Internet-Draft           Agent Selection Mapping          September 2026

   *  the banding policy, that is, the coarseness at which continuous
      quantities are compared, which bounds the rate at which decisions
      can change;

   *  the look-ahead horizon, that is, the number of future steps over
      which

   the objective is considered when a choice made now would be expensive
   to reverse;

   *  the budget pacing rule, that is, whether the session may spend its
      budget unevenly across steps, and how much of the remaining budget
      is reserved for the steps that are expected to be the most expensive.

6.5.  Decision Artifacts

   Five artifacts are produced or consumed by one application of the
   procedure. Their separation matters because they have different
   lifetimes, and mixing them is what makes a selection oscillate. The
   decision record is also the artifact through which an operations and
   management view of CATS can observe the mapping
   [I-D.ietf-cats-oam-fw].

   *  Descriptor: produced by the agent gateway or by C-TC; it lives for
      the step, and it is the input to one or more evaluations.

   *  Resource view snapshot: produced by C-SMA and C-NMA; it lives
      until its freshness bound expires, and it is the input to an
      evaluation.

   *  Decision record: produced by C-PS; it lives for the session, and
      it serves audit, accounting, and the later decisions of the same
      session.

   *  Commitment: produced by C-PS; it lives until it is revised at a
      decision point, and it binds a step to an instance.

   *  Handoff record: produced by C-PS; it lives from the reservation to
      the commit or the abort, and it drives the movement of state.

6.6.  Characteristic-to-Resource Mapping

   This subsection is the core of the mapping. It states, for each
   characteristic of an agent service, how the characteristic appears in
   the three dimensions at selection time and what the selection
   function has to do about it. The dimensions are not symmetric at
   selection time.

   *  Transfer is paid once per change of instance, and it is paid in
      bulk: what matters is the bytes that have to be moved before the
Mo, et al.               Expires 3 April 2027                  [Page 14]
Internet-Draft           Agent Selection Mapping          September 2026

      step can run, not the bytes of the step itself. Transfer cost
      appears in the forwarding dimension, which is where the paths
      and their cost are evaluated.

   *  Storage is paid per element and per step: what matters is which
      working set elements are present, where they reside, and what
      it costs to make them usable.

   *  Compute is paid per step: what matters is the capability that the
      step needs and the capacity that is free for it when it arrives.
  
   Table 1 maps each characteristic onto the signal that each dimension
   observes. Table 2 maps it onto the decision, and states whether the
   characteristic acts as a constraint or as a preference. Table 3 gives
   the order of magnitude that the dimensions have to compare; the
   values are measured in [I-D.mo-cats-agent-service-characteristics],
   and they are indicative rather than normative.

Mo, et al.               Expires 3 April 2027                  [Page 15]
Internet-Draft           Agent Selection Mapping          September 2026

      Characteristic   | Transfer      | Storage       | Compute
      -----------------+---------------+---------------+----------------
      Session state    | bytes moved   | presence per  | recompute cost
                       | per change;   | reuse key;    | as the
                       | per-prefix    | tier;         | alternative to
                       | cost to the   | retrieval     | transfer
                       | state source  | cost          |
      Multi-step and   | per-endpoint  | tool results  | tier needed by
      tools            | path cost;    | retained per  | this step
                       | fan-out to    | key           |
                       | several       |               |
                       | prefixes      |               |
      Long-horizon and | transfer      | retention     | re-prefill cost
      pauses           | window after  | against       | after eviction
                       | a wake-up     | eviction in a |
                       |               | pause         |
      Context growth   | bulk prefix   | context state | prefill for the
                       | transfer;     | grows with    | retained prefix
                       | compression   | tokens        |
                       | cuts bytes    |               |
      Capability tiers | not           | artifact      | tier match, as
                       | applicable    | residency per | a constraint
                       |               | tier          |
      Communication    | flow shape:   | state updated | concurrency the
      mode             | streaming,    | per response  | step needs
                       | fan-out,      |               |
                       | burst         |               |
      Memory           | reachability  | long-lived    | lookup cost per
      persistence      | of the memory | elements;     | step
                       | store         | sharing scope |
      Locality and     | permitted     | permitted     | permitted tier
      tenancy          | paths and     | location per  | per element
                       | regions       | element       |
      Budget           | bytes charged | storage time  | compute time
                       | to the        | charged       | charged
                       | session       |               |
      Turn and task    | p95 path      | p95 retrieval | p95 queue and
      objective        | delay over    | delay         | step latency
                       | the window    |               |
      Model artifact   | bulk load     | residency and | load time
      size             | transfer,     | headroom for  | before the
                       | seconds to    | weights       | first token
                       | minutes       |               |

Mo, et al.               Expires 3 April 2027                  [Page 16]
Internet-Draft           Agent Selection Mapping          September 2026

      Characteristic  | Decision rule at         | Role       | Req.
                      | selection time           |            |
      ----------------+--------------------------+------------+---------
      Session state   | prefer the instance      | preference | M4, M5,
                      | holding the largest      |            | M23, M27
                      | share of the working     |            |
                      | set, unless transfer or  |            |
                      | recompute costs more     |            |
                      | than the gain            |            |
      Multi-step and  | value the path to each   | both       | M3, M34,
      tools           | endpoint and bound the   |            | M40
                      | turn by its slowest call |            |
      Long-horizon    | re-verify state presence | preference | M25,
      and pauses      | after a pause, and keep  |            | M31, M37
                      | the commitment stable    |            |
                      | across steps             |            |
      Context growth  | require memory headroom  | constraint | M22, M24
                      | for the accumulated      |            |
                      | context state, not for   |            |
                      | the prompt               |            |
      Capability      | exclude instances below  | constraint | M6, M19
      tiers           | the required tier; do    |            |
                      | not trade the tier for   |            |
                      | distance                 |            |
      Communication   | keep a streaming         | both       | M29, M42
      mode            | response on one          |            |
                      | instance; let fan-out    |            |
                      | paths differ             |            |
      Memory          | prefer proximity to the  | preference | M23, M24
      persistence     | memory store             |            |
      Locality and    | filter before ranking;   | constraint | M6, M20,
      tenancy         | an element that may not  |            | M28
                      | move stays where it is   |            |
      Budget          | charge the cost of the   | both       | M15,
                      | change to the session    |            | M32, M38
                      | budget before moving     |            |
      Turn and task   | compare at the stated    | preference | M17,
      objective       | percentile over the      |            | M33, M34
                      | stated window            |            |
      Model artifact  | exclude instances that   | both       | M2, M24,
      size            | cannot hold weights and  |            | M32
                      | state, then account load |            |
                      | time in the change cost  |            |

Mo, et al.               Expires 3 April 2027                  [Page 17]
Internet-Draft           Agent Selection Mapping          September 2026

      Anchor quantity  | Value and unit             | Source
      -----------------+----------------------------+----------------
      Context state    | 55 MiB per 1k tokens at    | Section 5.4
      size             | 7B; 313 MiB at 72B, half   |
                       | precision                  |
      Model artifacts  | 0.88 to 688 GB over        | Section 6.1
                       | fourteen models; 15.23 GB  |
                       | to 5.57 GB at 4-bit        |
      Bulk transfer    | seconds to minutes per     | Section 6.1
      cost             | change of instance, not    |
                       | milliseconds               |
      Tool call        | heterogeneous and          | Section 5.2
      latency          | heavy-tailed; the slowest  |
                       | call bounds the turn       |
      Objective        | a percentile over a        | Section 5.3,
      statement        | window; a mean is not a    | R17
                       | substitute                 |

   Three rules follow from the tables and apply to every characteristic.

   *  A characteristic that cannot be observed in a dimension leaves
      that dimension unknown for the decision. Unknown is not a favorable
      value, and a constraint that rests on it is not satisfied.

   *  A characteristic that appears as a constraint is filtered before
      ranking; the same characteristic may appear as a preference in
      another session, and the descriptor states which it is.

   *  A cost that is paid once per change is compared with the remaining
      budget and the remaining horizon, not with the gain of the single
      step that triggers it.

   Section 6.7 takes the same characteristics to the level of the
   concrete resources that have to be compared.

6.7.  Characteristic-to-Resource Routing

   Section 6.6 maps a characteristic onto a dimension. This subsection
   takes the same characteristics to the level of the concrete resources
   that a decision has to compare, because that is the level at which a
   CATS system can route: a dimension is not observable, whereas
   accelerator memory free, path egress, and state residency are.

Mo, et al.               Expires 3 April 2027                  [Page 18]
Internet-Draft           Agent Selection Mapping          September 2026

   The resource list below is also the list that a metric framework
   would have to expose for this mapping to be operational. Stating it
   at the level of the physical resource and its unit is what allows the
   requirements of Section 11 to be met by a metric definition and a
   data model rather than by an implementation.

   Table 4 is the resource inventory. T, S, and C denote the transfer,
   storage, and compute dimensions.

      Resource          | Dim. | Observed as                | Observed
                        |      |                            | by
      ------------------+------+----------------------------+-----------
      Path egress       | T    | available bytes/s towards  | C-NMA
      capacity          |      | the candidate              |
      Path delay and    | T    | ms, median and p95         | C-NMA
      jitter            |      |                            |
      Endpoint cost     | T    | per-prefix ms to a tool or | C-NMA
                        |      | data endpoint              |
      Transfer window   | T    | class and permitted start  | operator
                        |      | window                     |
      Accelerator       | C/S  | bytes free for weights and | C-SMA
      memory free       |      | state                      |
      Host memory free  | S    | bytes free                 | C-SMA
      Local persistent  | S    | bytes free and read MB/s   | C-SMA
      store             |      |                            |
      Remote state      | S    | RTT and throughput         | C-SMA
      store             |      |                            |
      State residency   | S    | tier per reuse key         | C-SMA
      State retention   | S    | seconds the instance will  | C-SMA
                        |      | keep it                    |
      Accelerator       | C    | type, count, precision set | C-SMA
      capability        |      |                            |
      Waiting work      | C    | queue depth and p95 wait   | C-SMA
      Prefill and       | C    | available tokens/s         | C-SMA
      decode rate       |      |                            |
      Context window    | C    | tokens the instance can    | C-SMA
                        |      | hold                       |
      Artifact locality | C/S  | registry RTT and load MB/s | C-SMA

   Table 5 states, for each characteristic, which of those resources
   carries it and what the routing step does about it.

Mo, et al.               Expires 3 April 2027                  [Page 19]
Internet-Draft           Agent Selection Mapping          September 2026

      Characteristic  | Resource that     | Routing action
                      | carries it        |
      ----------------+-------------------+--------------------------
      Session state   | state residency,  | route to a holder when
                      | retention,        | the state fits free
                      | accelerator       | memory, else compare
                      | memory            | transfer against
                      |                   | recompute and take the
                      |                   | cheaper
      Accumulated     | accelerator       | require free memory for
      context         | memory, prefill   | weights plus context
                      | rate              | state; if unmet, degrade
                      |                   | or choose a larger-memory
                      |                   | candidate
      Tool fan-out    | endpoint cost,    | value the path to each
                      | path egress       | endpoint; bound the turn
                      |                   | by the slowest call; let
                      |                   | fan-out paths differ
      Long horizon    | retention, host   | prefer retention that
                      | memory            | covers the expected idle;
                      |                   | keep the commitment
                      |                   | stable across steps
      Wake-up after a | state residency,  | re-check residency; treat
      pause           | retention         | an evicted element as
                      |                   | absent and re-run the
                      |                   | comparison
      Capability tier | accelerator       | exclude on type,
                      | capability,       | precision, or window; do
                      | context window    | not trade against
                      |                   | distance
      Compute load    | waiting work,     | prefer the lowest p95
                      | prefill rate      | wait at equal capability,
                      |                   | and re-check at each step
                      |                   | boundary
      Bulk artifact   | artifact          | charge the load time to
                      | locality, path    | the change cost; exclude
                      | egress            | candidates whose load
                      |                   | exceeds the step budget
      Governance      | state residency,  | filter before ranking;
                      | permitted regions | keep an element that may
                      |                   | not move in its permitted
                      |                   | set
      Budget          | all three         | charge transfer once per
                      | dimensions        | change, storage per
                      |                   | element-step, compute per
                      |                   | step; refuse a move
                      |                   | beyond the budget
      Tail objective  | path delay,       | compare at p95 over the
                      | waiting work,     | window; exclude
                      | retrieval         | candidates whose tail
                      |                   | breaches a constraint
Mo, et al.               Expires 3 April 2027                  [Page 20]
Internet-Draft           Agent Selection Mapping          September 2026

   The routing step uses the comparisons below. They are written as
   relations between resource quantities to state what has to be
   comparable. They define no encoding, and the values used later are
   illustrative.

      Resource-side quantities used by the routing step:

      state_bytes         | size of the working set that must be usable
      weights_bytes       | size of the artifact that must be resident
      hbm_free            | accelerator memory free at the candidate
      bw                  | available path egress towards the candidate
      rtt                 | path delay towards the candidate
      prefill_rate        | context tokens processable per second
      tokens_to_reprocess | tokens a re-computation would redo
      wait_p95            | p95 waiting time for a comparable step
      step_latency_p95    | p95 service time for a comparable step

      Derived comparisons, illustrative and not normative:

      state_fits          | state_bytes <= hbm_free - weights_bytes
      transfer_time       | state_bytes / bw + rtt
      recompute_time      | tokens_to_reprocess / prefill_rate
      state_cost          | min(transfer_time, recompute_time)
      tail_ok             | wait_p95 + step_latency_p95 <= objective
      budget_ok           | session_cost + change_cost <= budget

6.8.  Resource-Level Routing Recipe

   1. Fit. Test the memory and capability relations at every candidate:
   drop the candidates where the working set cannot be made usable, and
   the candidates whose context window is too small for the step.

   2. Cost of not holding state. Compute the state cost of each
   survivor, and drop the candidates whose state cost exceeds the budget
   of the step.

   3. Path. Value the endpoints that this step will contact, not only
   the path to the instance; the slowest intended call bounds the turn.

   4. Tail and budget. Test the tail objective at the stated percentile
   over the stated window, and test the cost of the change against the
   session budget that the change is intended to serve.

   5. Commit through Section 8.3. The resource comparison does not
   select on its own: a candidate that wins on resources is still
   subject to the stability requirement, because moving a session costs
   state and budget.

Mo, et al.               Expires 3 April 2027                  [Page 21]
Internet-Draft           Agent Selection Mapping          September 2026

   When step 1 or step 2 leaves no candidate, the degradation classes of
   the step are applied in the order of Section 8.4 before the session
   is failed.

      Step 30 of the session of Section 8.5, with illustrative values.

      Candidate C holds the state:
        hbm_free 20 GB, weights_bytes 15.23 GB, state_bytes 7.5 GB
        state_fits is true (7.5 <= 20 - 15.23)
      Candidate B does not hold it, but has a better path:
        bw 2.5 GB/s, rtt 8 ms, prefill_rate 1,000 tokens/s
        transfer_time  = 7.5 GB / 2.5 GB/s + 8 ms  ~ 3.0 s
        recompute_time = 128,000 tokens / 1,000 tokens/s ~ 128 s
        state_cost at B ~ 3.0 s

      Transfer is far cheaper than re-computation here, so holding the
      state does not by itself pin the session to C.
      The routing step therefore falls through to path and budget: B
      wins only if its p95 path and its remaining budget are better
      by more than the stability requirement of Section 8.3 allows.

7.  Information Elements and Their Abstract Syntax

   This section states what the mapping has to be able to express. It
   defines no concrete syntax: no encoding, no field layout, and no
   registry is defined here, and the concrete representation of these
   elements belongs to the data model work [I-D.ietf-cats-data-model].
   What is defined here is the meaning of each element, the abstract
   class of its value, the scope in which it applies, and what the
   absence of the element means.

   The classes below are abstract. They exist so that the semantics of
   an element can be stated without choosing a representation.

Mo, et al.               Expires 3 April 2027                  [Page 22]
Internet-Draft           Agent Selection Mapping          September 2026

      Class      | Domain                     | Typical use
      -----------+----------------------------+-------------------------
      BOOL       | true or false              | may this traffic be
                 |                            | buffered
      ENUM       | one of a named set         | capability tier,
                 |                            | rejection reason
      RANGE      | [lo, hi] with a stated     | context length, budget
                 | unit                       | limit
      COUNT      | non-negative integer       | remaining steps of a
                 |                            | horizon
      QUANTILE   | value at rank q over       | p95 step latency
                 | window W                   |
      TIME       | instant, UTC               | collection time of a
                 |                            | value
      DURATION   | length of an interval      | state age, freshness
                 |                            | bound
      SIZE       | bytes or tokens            | element size, working
                 |                            | set size
      HANDLE     | opaque token, compared for | state handle, reuse key
                 | equality                   |
      SET        | unordered, no duplicates   | elements of a working
                 |                            | set
      MAP        | key to value               | reuse key to presence
      TEXT       | short string for humans    | reason code, operator
                 |                            | label

7.1.  Descriptor Elements

Mo, et al.               Expires 3 April 2027                  [Page 23]
Internet-Draft           Agent Selection Mapping          September 2026

      Element              | Class     | Semantics
      ---------------------+-----------+--------------------------------
      service-id           | HANDLE    | the CS-ID of the service,
                           |           | session scope
      session-id           | HANDLE    | attribution key of the session
      turn-id              | COUNT     | attribution key of the turn
      step-id              | COUNT     | attribution key of the step
      capability-tier      | ENUM      | the lowest tier the step may be
                           |           | served at
      required-precision   | SET       | precision formats that are
                           |           | acceptable
      context-length       | RANGE     | tokens that the step keeps in
                           |           | context
      working-set          | SET       | the elements needed, each with
                           |           | a reuse key
      element-class        | ENUM      | class of an element, per
                           |           | Section 6.3
      locality             | MAP       | element class to permitted
                           |           | regions
      tenant               | HANDLE    | tenant under which the session
                           |           | runs
      budget               | MAP       | dimension to limit, ceiling or
                           |           | target
      objective            | QUANTILE  | quantity, percentile, window,
                           |           | weight
      tail-tolerance       | RANGE     | probability of missing the
                           |           | objective
      degradation          | SET       | reductions that the step
                           |           | tolerates
      freshness-need       | MAP       | element to the maximum usable
                           |           | age
      horizon              | COUNT     | steps that the session is
                           |           | expected to run
      mode                 | ENUM      | request-response, streaming,
                           |           | tool call, fan-out

7.2.  Resource View Elements

      Element              | Class     | Semantics
      ---------------------+-----------+--------------------------------
      path-set             | SET       | the candidate paths towards the
                           |           | instance
      endpoint-cost        | MAP       | per-prefix cost to tools and
                           |           | data sources
      determinism          | ENUM      | whether the path offers bounded
                           |           | delay
      transfer-window      | DURATION  | time in which bulk transfer may
                           |           | start

Mo, et al.               Expires 3 April 2027                  [Page 24]
Internet-Draft           Agent Selection Mapping          September 2026

      Element              | Class     | Semantics
      ---------------------+-----------+--------------------------------
      tier-offered         | ENUM      | the highest tier the instance
                           |           | can serve
      queue-delay          | QUANTILE  | wait for a comparable step, at
                           |           | p95
      memory-headroom      | SIZE      | accelerator memory free for
                           |           | state
      step-latency         | QUANTILE  | recent comparable steps, median
                           |           | and p95

      Element              | Class     | Semantics
      ---------------------+-----------+--------------------------------
      presence             | MAP       | reuse key to present, absent,
                           |           | or unknown
      residency            | ENUM      | tier at which an element
                           |           | resides
      element-size         | SIZE      | size of the element as it would
                           |           | be moved
      retrieval-cost       | DURATION  | time to make the element usable
                           |           | there
      recompute-cost       | DURATION  | time to rebuild it without
                           |           | transfer
      state-age            | DURATION  | age of the element as held at
                           |           | the instance
      retention            | DURATION  | how long the instance will keep
                           |           | it
      governance           | MAP       | element to the constraint that
                           |           | applies

7.3.  Accounting Elements

      Element              | Class     | Semantics
      ---------------------+-----------+--------------------------------
      consumed             | MAP       | tokens, time, and cost consumed
                           |           | so far
      remaining            | MAP       | limit minus consumed, per
                           |           | dimension
      reserve              | MAP       | budget held back for the
                           |           | remaining steps
      attribution          | HANDLE    | session, turn, and step to
                           |           | charge

7.4.  Exchange Metadata Elements

   Every element that moves between components carries the metadata
   below.

Mo, et al.               Expires 3 April 2027                  [Page 25]
Internet-Draft           Agent Selection Mapping          September 2026

      Element              | Class     | Semantics
      ---------------------+-----------+--------------------------------
      collected-at         | TIME      | when the value was observed
      freshness-bound      | DURATION  | how long the value may be used
      source               | HANDLE    | the component that asserted the
                           |           | value
      confidence           | ENUM      | asserted, derived, or estimated

7.5.  Semantics Rules Common to the Elements

   The following rules apply to every use of the elements above. They
   are the part of this document that a concrete data model has to
   preserve.

   *  Absence of a value from the descriptor, the resource view, or an
      exchange means that the value is unknown. Unknown is not a value: 
      a constraint that cannot be evaluated is not satisfied.

   *  Every quantity carries the time at which it was observed and a
      freshness bound. A value whose age exceeds its bound must not be
      used to satisfy a constraint, and its use in a preference must be
      reported in the decision record.

   *  A quantity that is a percentile states the rank and the window. A
      quantity that carries no rank is a mean and must be identified as a
      mean.

   *  The scope of an element is explicit. A constraint that is scoped
      to the session applies to every step of that session, and a constraint
      that is scoped to a step does not outlive it.

   *  Constraint and quantity are separate roles. The same physical
      property may appear in both roles with different semantics, as
      when a delay is a hard ceiling in one descriptor and a preference
      in another.

   *  A state handle is opaque. Testing the presence of an element,
      requesting its transfer, and attaching affinity to it must not
      require the network to learn its content, and two equal handles
      must denote the same element.

Mo, et al.               Expires 3 April 2027                  [Page 26]
Internet-Draft           Agent Selection Mapping          September 2026

   *  Consumption is monotone. A budget decreases as a session runs, is
      attributed to a session, a turn, and a step, and a ceiling cannot be
      traded against a preference.

   *  Banding is stated, not implied. Where values are compared in
      bands, the band edges are known to the decision, and banding must
      not hide a violation of a constraint.

   *  Units and populations are explicit. A quantity states its unit,
      and an aggregated quantity states the class, the key space, and
      the number of instances that it covers.

8.  Selection Procedure

   This section states the phases through which the mapping runs, the
   points at which a decision is required, the events that cause a
   re-evaluation, and the order in which the procedure degrades when it
   cannot proceed as intended. The phases are written in operational
   terms, and they assume that the information of Section 7 is available
   to the C-PS.

8.1.  Phases

   P0. Session admission. Input: the session objective, the budget, and
   the standing constraints. Action: establish that at least one
   candidate can serve the session, and create the selection context.
   Output: the session context and an initial commitment. On failure:
   the fallback ladder of Section 8.4.

   P1. Step boundary detection. Input: the traffic of the session, and,
   where available, the step signal from the agent gateway. Action:
   determine that a new step begins, and record the boundary. Output: a
   step identifier and the accumulated working set. On failure: treat
   the turn as a single step, which forbids per-step migration and is
   therefore the conservative choice.

   P2. Descriptor derivation. Input: the step, the accumulated state,
   and the session context. Action: derive the descriptor, and request
   from the gateway any element that the network cannot derive (MSG-3).
   Output: the descriptor for the step. On failure: if a required
   element cannot be obtained, the step is not evaluable and Section 8.4
   applies.

Mo, et al.               Expires 3 April 2027                  [Page 27]
Internet-Draft           Agent Selection Mapping          September 2026

   P3. Candidate enumeration. Input: the service identifier, the
   instance set, and the endpoints that the step is expected to contact.
   Action: form the candidate set of instances and paths, including the
   endpoints discovered for this step. Output: the candidate set. On
   failure: an empty candidate set triggers Section 8.4.

   P4. Constraint filtering. Input: the descriptor and the resource
   view. Action: remove every candidate that violates a constraint,
   treating an unknown value as not satisfying the constraint. Output:
   the eligible set, with the reason for each exclusion. On failure: an
   empty eligible set triggers Section 8.4.

   P5. Per-dimension valuation. Input: the eligible candidates and the
   descriptor. Action: value the forwarding, computing, and storage
   dimensions, element by element for storage, and at the stated
   percentile where the objective is tail-scoped. Output: one valuation
   per dimension per candidate. On failure: a dimension that cannot be
   valued at all is treated as unknown and is reported in the decision
   record.

   P6. Combination and tail awareness. Input: the per-dimension
   valuations and the preferences. Action: combine by the stated rule,
   weighted or lexicographic, and apply the tail objective as a bound on
   the outcome rather than as one more term in a sum. Output: a
   suitability per candidate. On failure: if no combination rule
   applies, the descriptor is malformed and the step is refused rather
   than served on a default.

   P7. Change cost and stability. Input: the suitability values, the
   current commitment, and the selection context. Action: compare the
   gain of moving against the cost of moving, that is, transfer or
   reconstruction plus the disturbance of a step in flight, and apply
   the stability requirement in force. Output: keep, hold-and-watch, or
   move. On failure: an unquantified cost of change leads to the
   conservative choice, which is not to move the session.

   P8. Commitment. Input: the outcome of P7. Action: bind the step, or
   the range of steps, to the selected instance with an explicit
   lifetime, and hand the steering instruction to the forwarders
   (MSG-10). Output: the commitment and the steering instruction. On
   failure: a commitment that cannot be established is a failure of the
   step, and it is reported as one.

   P9. Handoff orchestration. Input: a move decision and the difference
   between the working sets of the two instances. Action: reserve at the
   target (MSG-11), transfer or reconstruct the elements that the target
   lacks (MSG-13), and commit only when the working set is usable;
   otherwise abort, remain at the source, and record why. Output: a
   handoff record, committed or aborted. On failure: the session stays
   where it is, and the migration is retried only at a later decision
   point.
Mo, et al.               Expires 3 April 2027                  [Page 28]
Internet-Draft           Agent Selection Mapping          September 2026

   P10. Execution monitoring and tail detection. Input: the progress of
   the running step against its expectation. Action: detect a step whose
   cost exceeds that expectation by the stated factor, and apply the
   degradation class of the step within what the constraints allow
   (MSG-14). Output: a degradation decision, a mid-step migration
   request, or a decision to fail the step. On failure: the step is
   failed with a report that names the constraint that could not be met,
   rather than continuing to overrun the budget.

   P11. Accounting, feedback, and close. Input: the consumption of the
   completed step. Action: attribute the consumption to session, turn,
   and step, update the remaining budget and the reserve, append the
   decision record, and at the end of the session release the state that
   is no longer needed (MSG-15, MSG-16). Output: an updated selection
   context. On failure: accounting that cannot be attributed degrades
   only the precision of later budget decisions, and it is reported.

8.2.  Decision Points

   A decision is required at the points below. Between them the mapping
   is not required to act, which is what keeps a stream of metric
   updates from becoming a stream of decisions.

   *  DP1: session admission -- establish the initial commitment (P0).

   *  DP2: start of a turn -- re-evaluate the commitment (P2 to P8).

   *  DP3: step boundary -- re-evaluate for the next step (P2 to P8).

   *  DP4: wake-up after a pause -- re-verify state presence (P4, P5).

   *  DP5: metric freshness expiry -- re-validate the values used (P5).

   *  DP6: sustained improvement -- consider a move (P7, P8).

   *  DP7: constraint at risk -- re-filter and move if needed (P4, P9).

   *  DP8: tail event detected -- degrade, migrate, or fail (P10).

   *  DP9: budget threshold crossed -- pace or refuse the step (P11).

   *  DP10: session close -- release state and account (P11).

8.3.  Re-Evaluation Triggers and Stability

   The triggers below cause a re-evaluation. Each is subject to the
   stability requirement, which exists because a session that is moved
   for a gain that does not survive the move is worse off than a session
   that is not moved.

Mo, et al.               Expires 3 April 2027                  [Page 29]
Internet-Draft           Agent Selection Mapping          September 2026

   *  T1: step boundary -- re-evaluate unless the commitment covers the
      step.

   *  T2: turn start -- re-evaluate with the accumulated state.

   *  T3: wake-up after a pause -- re-verify state presence before
      reuse.

   *  T4: freshness expiry -- re-validate the values the decision used.

   *  T5: sustained improvement -- move only if the stability rule is
      met.

   *  T6: constraint at risk -- re-filter and move; stability does not
      apply.

   *  T7: instance degradation -- re-evaluate, and move if no step is in
      flight.

   *  T8: tail event -- apply the degradation class, or fail with a
      report.

   *  T9: budget threshold -- pace the remaining steps or refuse the
      step.

   *  T10: session close -- release state and close the accounting.

   Four stability rules apply to every trigger except T6 and T8, where a
   constraint is already at risk and delay is the risk itself.

   *  A minimum interval must elapse between two moves of the same
      session, unless a constraint is violated.

   *  A move requires a minimum improvement of suitability, so that a
      value which fluctuates inside a band does not move a session.

   *  The cost of change is compared with the gain over the remaining
      steps of the look-ahead horizon, not with the gain of one step.

   *  A step in flight is not moved on a preference. Only a violated
      constraint, or a tail event that the degradation class cannot
      absorb, justifies moving work that has already started.

8.4.  Fallback Ladder

   When the procedure cannot produce a selection that satisfies every
   constraint, it degrades in the order below rather than failing the
   session or silently overrunning its budget. The order follows the
   CATS fallback decision framework
   [I-D.pang-cats-fallback-decision-framework].

   1. Re-band the descriptor. Relax a preference and keep every
   constraint. This is the only step that changes what the session asked
   for, and it changes the least significant part of it.

Mo, et al.               Expires 3 April 2027                  [Page 30]
Internet-Draft           Agent Selection Mapping          September 2026

   2. Degrade within the degradation class. Serve the step at a lower
   precision, at a lower capability tier, with a truncated working set,
   or with additional delay, where the descriptor permits that
   reduction.

   3. Move the work. If the step has not started, select the best
   remaining eligible candidate; if it has started, move only when a
   constraint is violated and the degradation class cannot absorb the
   violation.

   4. Fail the step with a report. The report names the constraint that
   could not be satisfied, the trigger that led to the attempt, and the
   candidates that were rejected, with the reason for each rejection.

   5. Fail the session. Where the budget is exhausted or no candidate
   can serve the session at all, the failure is reported to the gateway
   so that the session can be deferred, approved by a human, or ended,
   rather than being served by retries that consume what is left of the
   budget.

   The following example illustrates phases P2 to P7, with illustrative
   values, for a retrieval-augmented assistant session in which an
   earlier turn has already produced a KV cache and a retrieved document
   set.

      Step: answer a follow-up question, 2 s objective, data must remain
            in region R, cost ceiling applies to the session.

      Candidate A (near edge, small model, holds session KV cache):
        forwarding good, computing: capability constraint NOT met.
      Candidate B (regional, capable, holds document set, no KV):
        forwarding ok, computing ok, storage: document set present,
        KV absent, recompute cost of KV = 0.6 s.
      Candidate C (regional, capable, holds KV cache and documents):
        forwarding ok, computing ok, storage: all present, cost ~0.

      A is eliminated by the capability constraint.  C is preferred over
      B unless C fails a constraint or its forwarding cost exceeds the
      recompute saving of 0.6 s.

   A second illustration shows the same session later, when the state
   that has accumulated changes the answer and a tail event occurs.

      The same session at step 30, with illustrative values.

      Accumulated state: context state for about 60,000 tokens, a
      document set, and the plan state of the current task.

Mo, et al.               Expires 3 April 2027                  [Page 31]
Internet-Draft           Agent Selection Mapping          September 2026

      Candidate B (regional, capable, holds the document set, no
      context state): storage: the context state is absent, and the
      recompute cost for it is large because the step needs the
      earlier context. Candidate C (regional, capable, holds the full
      context state): storage: all elements present.

      C is kept.  The gain from moving to B does not cover the cost
      of re-establishing the context state, and the stability
      requirement therefore keeps the session at C even where B
      reports more free capacity.

      During the step, a retrieval tool that normally answers in 100
      ms is still running after 4 s.  The step is a tail event (T8).
      The degradation class of the step permits a reduced retrieval
      set, so the procedure narrows the set and the step completes.
      The session is not migrated mid-step, because the context state
      is already resident and moving it would cost more than the
      delay it would remove.

9.  Message Semantics

   The phases of Section 8 require information at defined moments. This
   section names the exchanges through which that information is
   obtained, and states what each exchange has to carry, when it is
   expected, and what it must do when it cannot be answered.

   The names below are for exposition only. This document defines no
   encoding, no transport, and no message format; what it requires is
   that the semantic content of each exchange is available at the point
   in the procedure where the exchange is used. Carriage belongs to the
   documents that define the distribution mechanisms of the CATS
   framework.

9.1.  Exchange Catalogue

   MSG-1. Descriptor submission (agent gateway to C-TC). Trigger:
   session admission, and each step that the gateway can describe.
   Content: the descriptor elements of Section 7, with their scope.
   Response: MSG-9 at the end of the procedure. Rules: the descriptor is
   local to the domain, and it carries no session content beyond what a
   decision needs.

   MSG-2. Step boundary notification (agent gateway to C-TC and C-PS).
   Trigger: the start of a step, where the gateway knows it. Content:
   the attribution keys and the accumulated working set. Response: none.
   Rules: a step that is never announced is treated as part of the turn.

Mo, et al.               Expires 3 April 2027                  [Page 32]
Internet-Draft           Agent Selection Mapping          September 2026

   MSG-3. Descriptor refinement request (C-PS to agent gateway).
   Trigger: the network cannot derive an element that a decision needs.
   Content: the element, the step, and the scope for which it is needed.
   Response: MSG-1 for that step. Rules: a refinement that is not
   answered leaves the element unknown.

   MSG-4. Capability and load advertisement (C-SMA to C-PS). Trigger: on
   change, and at a period bounded by the freshness bound of what is
   advertised. Content: capability elements and computing quantities
   (Section 7), aggregated where the metric definition allows it.
   Response: none. Rules: an advertisement is not a commitment, and a
   stale advertisement must not satisfy a constraint.

   MSG-5. State availability query (C-PS to C-SMA). Trigger: the summary
   in an advertisement does not decide the question for the step, or the
   presence information has expired. Content: the reuse keys for which
   presence is needed, and the scope. Response: MSG-6. Rules: a query
   names keys, and never asks for the content of an element.

   MSG-6. State availability response (C-SMA to C-PS). Trigger: receipt
   of MSG-5. Content: presence per reuse key, the residency tier, the
   retrieval cost, the recompute cost, the age of the element, and the
   retention. Response: none. Rules: a key that the instance cannot
   answer for is reported as unknown and not as absent.

   MSG-7. Forwarding view advertisement (C-NMA to C-PS). Trigger: on
   change, and at a period bounded by the freshness bound of the
   forwarding quantities. Content: paths, per-prefix cost to the
   endpoints that a step may contact, determinism class, and the
   transfer window. Response: none. Rules: the endpoints are discovered
   while the session runs, so the view is extended rather than decided
   once.

   MSG-8. Evaluation request (C-TC to C-PS). Trigger: a decision point
   of Section 8.2. Content: the descriptor, the attribution, and the
   trigger. Response: MSG-9. Rules: one request per decision point, not
   one per metric update.

   MSG-9. Selection result (C-PS to C-TC). Trigger: completion of the
   procedure. Content: the selected candidate, the reason codes of the
   candidates that were excluded, and the suitability that decided the
   outcome. Response: none. Rules: a result that could not be produced
   is a failure of the step and is reported as one, with the constraint
   that could not be met.

   MSG-10. Steering instruction (C-PS to forwarders). Trigger: a
   commitment. Content: the candidate, the lifetime of the commitment,
   and the endpoints that the step is expected to contact. Response:
   none. Rules: the instruction is idempotent, and a repeated
   instruction must not create a second commitment.

Mo, et al.               Expires 3 April 2027                  [Page 33]
Internet-Draft           Agent Selection Mapping          September 2026

   MSG-11. State reservation request (C-PS to the target C-SMA).
   Trigger: a decision to move a session, taken before any transfer.
   Content: the reuse keys that the target must hold, the size, and the
   time by which they are needed. Response: MSG-12. Rules: reservation
   is idempotent; reserving twice must not reserve two copies, and a
   reservation that cannot be honoured must be refused rather than
   accepted and broken.

   MSG-12. Reservation response (C-SMA to C-PS). Trigger: receipt of
   MSG-11. Content: accepted or refused, the validity window of the
   reservation, and the reason for a refusal. Response: none. Rules: an
   expired reservation is not a reservation.

   MSG-13. State transfer or reconstruction notice (source or target to
   C-PS). Trigger: completion, failure, or abort of the movement of an
   element. Content: the reuse keys that are now usable, the keys that
   are not, the divergence that was accepted, and the time. Response:
   none. Rules: the commit of a handoff follows this notice; a transfer
   that fails leaves the source usable and is reported.

   MSG-14. Invalidation and degradation notice (C-SMA, C-NMA, or
   forwarders to C-PS). Trigger: a value that a decision relied on has
   become stale or false, an instance has degraded, or a running step
   has exceeded its expected cost. Content: the element or the
   constraint affected, the observed condition, and the time. Response:
   MSG-9 after re-evaluation, where the step is still running. Rules:
   this exchange is what turns a tail event into a decision; a notice
   that is not acted on must still be recorded.

   MSG-15. Accounting feedback (C-TC or forwarders to C-PS). Trigger:
   completion of a step, and at the end of a turn. Content: consumption
   per dimension, attributed to session, turn, and step. Response: none.
   Rules: consumption is monotone and is not revised downwards.

   MSG-16. Session close and state release (agent gateway to C-TC; C-PS
   to C-SMA). Trigger: the end of the session, or its abandonment.
   Content: the attribution keys and the elements that are no longer
   needed. Response: none. Rules: release is best effort, and an element
   that is not released is subject to the retention that the instance
   stated in MSG-6.

9.2.  Exchange Sequences

   A step-boundary decision uses the exchanges in this order.

Mo, et al.               Expires 3 April 2027                  [Page 34]
Internet-Draft           Agent Selection Mapping          September 2026

      time
       |
       |  gateway: step boundary and descriptor           (MSG-1/2)
       |  C-TC:    attribute the step to session and turn   (MSG-8)
       |  C-PS:    detail state if the summary is stale      (MSG-5)
       |  C-SMA:   presence, tier, cost, age, retention      (MSG-6)
       |  C-NMA:   paths and per-prefix cost                 (MSG-7)
       |  C-PS:    filter, value, combine, apply stability  (8.1-8.3)
       |  C-PS:    steering instruction to the forwarders    (MSG-10)
       |  C-PS:    decision record and accounting        (MSG-9/15)
       v
      next decision point

   A handoff uses the exchanges in this order, and it can be abandoned
   at any point before the commit.

      time
       |
       |  C-PS:    decide to move the session to candidate B       (8.3)
       |  C-PS:    reserve at B, with a validity window      (MSG-11)
       |  C-SMA:   accept or refuse                          (MSG-12)
       |  source:  send or rebuild what the target lacks     (MSG-13)
       |  C-PS:    commit at B only when the set is usable   (MSG-10)
       |           on failure: abort, stay at the source, record why
       |  C-PS:    release what is no longer needed          (MSG-16)
       |

9.3.  Rules Common to the Exchanges

   *  Every exchange states the scope of what it carries, and a consumer
      that receives an element without a scope treats it as scoped to the
      step.

   *  Reservation and commit are idempotent. A repeated request must not
      produce a second reservation, a second commitment, or a second
      transfer.

   *  The order is reservation, transfer, commit. A commit that follows
      a failed or aborted transfer must be refused, and an abort must leave
      the source usable.

   *  Advertisements are aggregated, and detail is obtained by query. A
      deployment must be able to bound the rate of both, because the
      decision points of one session are not the only demand on the
      components.

Mo, et al.               Expires 3 April 2027                  [Page 35]
Internet-Draft           Agent Selection Mapping          September 2026

   *  A component that cannot answer reports unknown. Omission is read
      as unknown by the consumer, and never as a value.

   *  An exchange that crosses an administrative boundary carries the
      least information that the decision needs, and no session content.
      The elements of Section 7 are chosen so that this is possible.

10.  Aggregation, Normalization, and Freshness

   The mapping consumes metrics that are defined elsewhere. Three
   requirements on that consumption follow from the procedure above.

   Aggregation is needed in the storage dimension in the same way as in
   the other dimensions. An instance cannot advertise every element of
   every working set that it holds, so state availability needs to be
   summarized, for example per key space or per state class, and the
   summary needs to be sufficient for the comparison in step 5.

   Normalization is not applied across dimensions when a dimension
   carries a constraint: eligibility is decided first (see M1 and M6).
   The definition of a single normalized metric
   [I-D.ietf-cats-metric-definition] remains useful for ranking eligible
   candidates and for reporting, but the eligibility decision in step 2
   is made before normalization.

   Freshness is part of the input. A state availability value and a
   computing load value that were true several seconds ago may lead to a
   selection that is worse than a static default, so the freshness of
   the information used must be available to the procedure and must be
   able to trigger a re-evaluation [I-D.zhu-cats-metric-semantics].

11.  Requirements

   The following requirements are derived from the framework, the
   information elements, the procedure, and the exchanges above. They
   are stated using the conventions of BCP 14 [RFC2119] [RFC8174], are
   requirements on the design of a CATS system that supports agent
   services, and are not protocol requirements. They are grouped by the
   part of the mapping that they constrain, and they are numbered in one
   sequence so that a requirement can be cited without its group.

   M1. A CATS system SHOULD be able to evaluate a candidate instance
   against a set of constraints before it ranks candidates against each
   other.

Mo, et al.               Expires 3 April 2027                  [Page 36]
Internet-Draft           Agent Selection Mapping          September 2026

   M2. A CATS system SHOULD be able to express the capability that a
   step requires in terms that can be matched against the capability
   that an instance offers.

   M3. A CATS system SHOULD be able to evaluate the forwarding dimension
   of a step over the paths that the step will use, including paths to
   tools and data sources, rather than only over the path to the
   instance.

   M4. A CATS system SHOULD be able to represent, for a candidate
   instance, the presence and cost of the elements of the working set
   that the step needs.

   M5. A CATS system SHOULD be able to compare the cost of transferring
   a working set element with the cost of recomputing or re-retrieving
   it at the candidate.

   M6. A CATS system MUST be able to enforce locality, tenant, and
   jurisdiction constraints as constraints, and not as terms in a scalar
   score.

   M7. A CATS system SHOULD support more than one combination rule for
   the three dimensions, including at least a weighted rule and a
   lexicographic rule.

   M8. A CATS system SHOULD record, for each selection, the input that
   produced it, to the extent needed for audit and troubleshooting.

   M9. A CATS system SHOULD re-evaluate a selection when the metrics
   that supported it become stale, when the selected instance degrades,
   or when the step for which the selection was made completes.

   M10. A CATS system SHOULD account for the cost of changing the
   selected instance, so that a marginal improvement in suitability does
   not cause unnecessary migration of a session.

   M11. A CATS system SHOULD allow the operator to express the objective
   of a session, so that per-step decisions do not systematically
   disadvantage the completion of the session as a whole.

   M12. A CATS system SHOULD limit the state that it exposes to what the
   selection decision requires, and SHOULD pass it through aggregation
   where possible.

   M13. A CATS system SHOULD be able to attribute the traffic of a step
   to the session and turn to which it belongs, to the extent needed to
   apply a per-session objective or budget.

11.1.  Framework and Descriptor Requirements

Mo, et al.               Expires 3 April 2027                  [Page 37]
Internet-Draft           Agent Selection Mapping          September 2026

   M14. A CATS system SHOULD be able to derive a descriptor for a step,
   and to revise it between the steps of a session, rather than only for
   the session as a whole.

   M15. A CATS system SHOULD be able to express a budget with its
   dimension, such as time, tokens, cost, or energy, and to state
   whether that budget is a ceiling or a target.

   M16. A CATS system SHOULD be able to express the objective of a
   session and the horizon over which the objective is evaluated.

   M17. A CATS system SHOULD be able to express the objective of a step
   as a percentile over a stated window, together with the probability
   of exceeding that objective that the session tolerates.

   M18. A CATS system MUST keep constraints and preferences
   distinguishable throughout the procedure, including in the decision
   record and in the reports that the procedure produces.

   M19. A CATS system SHOULD be able to express a capability requirement
   as a tier or as a set membership test, and SHOULD preserve the reason
   when a candidate fails that test.

   M20. A CATS system MUST treat a value that is absent, stale, or of
   unknown provenance as unknown, and MUST NOT allow an unknown value to
   satisfy a constraint. Exchanges SHOULD indicate unknown explicitly
   rather than by omission.

   M21. A CATS system SHOULD be able to compare quantities in bands
   whose coarseness is stated in the selection context, and MUST NOT
   compare a banded value in a way that hides a violation of a
   constraint.

11.2.  State and Storage Requirements

   M22. A CATS system SHOULD be able to distinguish the classes of
   working set element defined in Section 6.3 when it evaluates the
   storage dimension.

   M23. A CATS system SHOULD be able to obtain the presence of a working
   set element per reuse key, down to a single element where the
   aggregated information does not decide the question.

   M24. A CATS system SHOULD be able to obtain, for a candidate and a
   reuse key, the residency tier, the retrieval cost, the recompute
   cost, and the age of the element.

   M25. A CATS system SHOULD re-verify the presence of the state that a
   session depends on after a pause that exceeds the freshness bound of
   the presence information.

Mo, et al.               Expires 3 April 2027                  [Page 38]
Internet-Draft           Agent Selection Mapping          September 2026

   M26. A CATS system SHOULD be able to execute a change of instance as
   a reservation, a transfer or reconstruction, and a commit, with an
   abort that leaves the source instance usable.

   M27. A CATS system SHOULD be able to state which part of a working
   set is sufficient for a step, and what divergence is acceptable when
   the whole set cannot be moved.

   M28. A CATS system MUST apply the constraint that governs each
   working set element to that element, and MUST NOT move an element
   whose constraint forbids the move.

11.3.  Discrete Decision Requirements

   M29. A CATS system SHOULD take its decisions at explicit decision
   points, SHOULD NOT require a decision for every update of every
   quantity, and SHOULD NOT split the work of one step across instances.

   M30. A CATS system SHOULD be able to obtain, or to detect, the
   boundary of a step, so that a selection can be revisited between
   steps.

   M31. A CATS system SHOULD apply a stability requirement before it
   selects a different instance for a session that is already placed.

   M32. A CATS system SHOULD compare the gain of changing the selected
   instance with the cost of the change, where that cost includes the
   transfer or reconstruction and the disturbance of a step in flight.

11.4.  Tail and Long-Horizon Requirements

   M33. A CATS system SHOULD evaluate an objective that is stated as a
   percentile at that percentile over the stated window, and SHOULD
   identify a quantity that is a mean as a mean.

   M34. A CATS system MUST NOT conclude that a constraint is satisfied
   from a mean value where the tail of the distribution can violate the
   constraint.

   M35. A CATS system SHOULD detect a running step whose cost exceeds
   its expected cost by the factor stated for that step.

   M36. A CATS system SHOULD apply the degradation class of a step in
   the order of the fallback ladder, and SHOULD fail a step with a
   report rather than allow it to overrun the session budget silently.

   M37. A CATS system SHOULD weigh state affinity against load balancing
   with the cost of re-establishing the state, considered over the
   look-ahead horizon rather than for one step.

Mo, et al.               Expires 3 April 2027                  [Page 39]
Internet-Draft           Agent Selection Mapping          September 2026

   M38. A CATS system SHOULD pace the consumption of a session budget
   across the steps of the session, and SHOULD hold back a reserve for
   the steps that remain.

   M39. A CATS system SHOULD retain the decision records of a session
   for the life of that session, to the extent needed by the later
   decisions of the same session and by the accounting of its budget.

11.5.  Exchange Requirements

   M40. A CATS system MUST be able to obtain each element of Section 7
   at the point in the procedure where that element is required.

   M41. A CATS system MUST convey, with every quantity, the time at
   which the quantity was observed and the bound within which it may be
   used.

   M42. A CATS system SHOULD support both a query mode and a
   notification mode for the availability of state, so that detail is
   obtained on demand and change is signalled without polling.

   M43. A CATS system MUST make the exchanges that reserve, commit, and
   move state idempotent and safe to retry or to reorder.

   M44. A CATS system SHOULD aggregate advertisements per class or per
   key space, and SHOULD obtain detail by query rather than by
   broadcasting it.

11.6.  Mapping Requirements

   M45. A CATS system SHOULD be able to state, for each characteristic
   of a step, which of the transfer, storage, and compute dimensions it
   affects, and whether it acts as a constraint or as a quantity in that
   dimension.

   M46. A CATS system SHOULD charge the cost of a change of instance to
   the session budget that the change is intended to serve.

   M47. A CATS system SHOULD treat the dimensions asymmetrically at
   selection time: a transfer cost is paid once per change of instance,
   while storage and compute costs are paid per step.

   M48. A CATS system SHOULD be able to express, for each resource that
   it compares, the unit in which the resource is observed and the
   component that observes it.

   M49. A CATS system SHOULD be able to derive the cost of not holding
   the state that a step needs at a candidate as the cheaper of the
   transfer time, that is, size over available bandwidth plus delay, and
   the recomputation time, that is, tokens over rate.

Mo, et al.               Expires 3 April 2027                  [Page 40]
Internet-Draft           Agent Selection Mapping          September 2026

   M50. A CATS system SHOULD be able to reject a candidate on the basis
   of a resource comparison, such as insufficient free memory, a step
   cost that exceeds the step budget, or a breached tail objective, and
   SHOULD report which comparison rejected it.

12.  Relationship to the CATS Framework

   The mapping framework uses the existing components. The following
   table states the responsibility of each component with respect to the
   three layers and to the exchanges of Section 9.

      Component  | Responsibility in the mapping
      -----------+------------------------------------------------------
      C-TC       | Classifies session traffic; attributes it to session,
                 | turn, and step; carries the descriptor and the
                 | accounting.
      C-SMA      | Provides capability, load, and state presence for its
                 | site; answers state queries; honours and reports
                 | reservations; releases state.
      C-NMA      | Provides the forwarding view, including the
                 | per-prefix cost to the endpoints of a step.
      C-PS       | Derives the descriptor, applies constraints, values
                 | and combines the dimensions, applies the stability
                 | requirement, commits, and records.
      Forwarders | Execute the steering instruction, including the
                 | endpoints of a step, and report progress that
                 | indicates a tail event.

   No new functional component is required. What the framework adds is
   the explicit distinction between constraints and quantities, the
   storage dimension of the resource view, and the definition of the
   Selection Point at which a decision is valid.

12.1.  Traceability to the Agent Service Requirements

   The requirements of [I-D.mo-cats-agent-service-characteristics] are
   addressed as follows. The last four rows map the four properties of
   Section 5, which are taken from the same source but are not single
   requirement items there.

Mo, et al.               Expires 3 April 2027                  [Page 41]
Internet-Draft           Agent Selection Mapping          September 2026

      Agent service  | Addressed by
      requirement    |
      ---------------+--------------------------------------------------
      R1, R2         | M1, M6, M7, M18, M19, M20
      R3             | M13; Selection Context (Layer 3)
      R4             | M9, M10, M31, M32
      R5, R6         | M3; Forwarding dimension (Layer 2)
      R7             | M36; tolerance element of the descriptor
      R8, R9, R10    | M2, M4, M19; Computing dimension
      R11, R12       | M4, M5, M23, M24
      R13            | M6, M28
      R14, R19, R20  | M22, M23, M24, M26, M27, M43
      R15            | M11, M15, M38
      R16, R21, R22  | Section 7 of the characteristics; M8, M12, M41,
                     | M44
      R17            | M33, M34, M35
      R18            | M42; mode element of the descriptor
      long-horizon   | M16, M25, M37, M38, M39
      (Section 5.1)  |
      stateful       | M22 to M28
      (Section 5.2)  |
      discrete       | M14, M21, M29, M30
      (Section 5.3)  |
      heavy-tailed   | M17, M33 to M36, M40
      (Section 5.4)  |
      resource       | M45, M46, M47
      mapping        |
      (Section 6.6)  |
      resource       | M48, M49, M50
      routing        |
      (Section 6.7)  |

13.  Operational Considerations

   Operators enabling this mapping should expect five consequences.

   *  The decision logic is more complex than a score comparison, which
      costs configuration and troubleshooting effort. Start with one
      constraint set and one combination rule, and add rules only where
      experience justifies them.

   *  State affinity concentrates load. If every session is steered to
      the instance that holds its state, a subset of instances absorbs
      the load, so the policy that overrides affinity must be expressible
      in the descriptor or in the selection context.

Mo, et al.               Expires 3 April 2027                  [Page 42]
Internet-Draft           Agent Selection Mapping          September 2026

   *  Correctness depends on the honesty of the resource view. An
      instance that over-reports capability or state attracts the sessions
      whose state it can then observe.

   *  The number of decisions grows with steps rather than with
      sessions, so the decision rate and the state query rate need a bound,
      and that bound belongs in the selection context.

   *  The mapping depends on summaries. A percentile or an aggregate
      computed from too few samples propagates its noise into the decision,
      which is one of the uses of the decision record.

14.  Security Considerations

   The decision depends on data supplied by the instances, which raises
   the value of misrepresenting it. Claimed state availability should be
   verified in the same way as claimed computing capability, and an
   implementation should not let a single dimension supplied by an
   untrusted party dominate the combination.

   The descriptor and the resource view expose the constraints and the
   working set of a session, so least disclosure applies: a descriptor
   should not be exposed beyond the components that need it.

   Reservation and handoff need their own limits. A reservation consumes
   the storage of an instance without moving a session to it, and a
   transfer consumes bandwidth between sites, so both need the
   authorization of the selection that they implement and a per-session
   rate limit.

   Tail-aware degradation can be gamed by an instance that under-reports
   its high percentiles or that reports an optimistic expectation.
   Percentile and expected-cost estimates should therefore be attributed
   to a source in the resource view, so that a source which is
   repeatedly wrong can be discounted.

15.  Privacy Considerations

   The descriptor and the resource view can reveal user behaviour: the
   number of steps, the sensitivity expressed through locality
   constraints, the size of the working set, and the tools in use.
   Aggregation and short retention are recommended, and the mapping does
   not require the network to learn the content of the session.

Mo, et al.               Expires 3 April 2027                  [Page 43]
Internet-Draft           Agent Selection Mapping          September 2026

   The classes and the reuse keys of a working set reveal structure even
   where the content does not, so reuse keys should be scoped to their
   tenant, and a handle should not be linkable across tenants or between
   the sessions of different users.

16.  IANA Considerations

   This document has no IANA actions.

17.  Normative References

   [I-D.ietf-cats-framework] Li, C., Du, Z., Boucadair, M., Contreras,
   L. M., et al., "A Framework for Computing-Aware Traffic Steering
   (CATS)", Work in Progress, Internet-Draft,
   draft-ietf-cats-framework-24, September 2026.

   [I-D.ietf-cats-metric-definition] Yao, K., et al., "CATS Metrics
   Definition", Work in Progress, Internet-Draft,
   draft-ietf-cats-metric-definition-11, September 2026.

   [I-D.mo-cats-agent-service-characteristics] Mo, Y., Yang, D., Zhou,
   C., "AI Agent Service Characteristics and Their Implications for
   Computing-Aware Traffic Steering", Work in Progress, Internet-Draft,
   draft-mo-cats-agent-service-characteristics-00, September 2026.

18.  Informative References

   [CATS-CHARTER] IETF, "Computing-Aware Traffic Steering (CATS) Working
   Group Charter",
   <https://datatracker.ietf.org/doc/charter-ietf-cats/>.

   [I-D.ietf-cats-data-model] Yao, H., Lin, C., et al., "Data Model for
   Computing-Aware Traffic Steering (CATS)", Work in Progress,
   draft-ietf-cats-data-model-00, September 2026.

   [I-D.ietf-cats-oam-fw] Fu, H., Xiong, Q., Du, Z., et al., "Computing-
   Aware Traffic Steering (CATS) Operations, Administration, and
   Maintenance (OAM) Framework", Work in Progress, Internet-Draft,
   draft-ietf-cats-oam-fw-01, July 2026.

   [I-D.ietf-cats-usecases-requirements] Yao, K., et al.,
   "Computing-Aware Traffic Steering (CATS) Problem Statement, Use
   Cases, and Requirements", Work in Progress,
   draft-ietf-cats-usecases-requirements-14, September 2026.

   [I-D.li-cats-kv-cache-distribution] Li, Z., et al., "KV Cache
   Distribution for Distributed LLM Inference: Use Case and
   Requirements", Work in Progress,
   draft-li-cats-kv-cache-distribution-00, July 2026.

Mo, et al.               Expires 3 April 2027                  [Page 44]
Internet-Draft           Agent Selection Mapping          September 2026

   [I-D.luan-cats-catpts] Li, Q., Luan, Z., et al., "A Timescale-Aware
   Framework for Compute-Aware Task Placement and Traffic Steering",
   Work in Progress, draft-luan-cats-catpts-01, August 2026.

   [I-D.pang-cats-fallback-decision-framework] Pang, R., Ed., Han, M.,
   Ed., Huang, T., Ed., "CATS Fallback Decision Framework", Work in
   Progress, Internet-Draft,
   draft-pang-cats-fallback-decision-framework-00, July 2026.

   [I-D.zhu-cats-metric-semantics] Zhu, M., "Operational Semantics for
   CATS Metric Consumption", Work in Progress,
   draft-zhu-cats-metric-semantics-01, August 2026.

   [RFC2119] Bradner, S., "Key words for use in RFCs to Indicate
   Requirement Levels", BCP 14, RFC 2119, DOI 10.17487/RFC2119, March
   1997, <https://www.rfc-editor.org/info/rfc2119>.

   [RFC8174] Leiba, B., "Ambiguity of Uppercase vs Lowercase in RFC 2119
   Key Words", BCP 14, RFC 8174, DOI 10.17487/RFC8174, May 2017,
   <https://www.rfc-editor.org/info/rfc8174>.
Acknowledgments

   The authors would like to thank the participants of the CATS working
   group for the discussions that shaped this document.

Authors' Addresses

   Y. Mo
   Huazhong University of Science and Technology
   Email: moyj@hust.edu.cn

   D. Yang
   Huazhong University of Science and Technology
   Email: d202581903@hust.edu.cn

   C. Zhou
   Huazhong University of Science and Technology
   Email: m202474228@hust.edu.cn

Mo, et al.               Expires 3 April 2027                  [Page 45]