CATS                                                               Y. Mo
Internet-Draft                                                   D. Yang
Intended status: Informational                                   C. Zhou
Expires: 3 April 2027      Huazhong University of Science and Technology
                                                       30 September 2026


         

A Selection Mapping Framework for AI Agent Services in

Computing-Aware Traffic Steering

draft-mo-cats-agent-selection-mapping-00 Abstract The Computing-Aware Traffic Steering (CATS) framework selects a service contact instance for a service request by combining computing and network metrics that are distributed by CATS Service Metric Agents and CATS Network Metric Agents. For AI agent services, the request that arrives at the network is not a single unit of work: it is a session that expands into multiple steps, each of which may require a different capability, a different state, and a different path. This document describes a mapping framework that turns the characteristics of an agent step into (1) a set of hard constraints and (2) a per-dimension valuation over a three-dimensional resource view composed of forwarding, computing, and storage, and that feeds the resulting suitability of each candidate service contact instance into the existing CATS selection function. The framework introduces no new functional component. Status of This Memo This Internet-Draft is submitted in full conformance with the provisions of BCP 78 and BCP 79. Internet-Drafts are working documents of the Internet Engineering Task Force (IETF). Note that other groups may also distribute working documents as Internet-Drafts. The list of current Internet- Drafts is at https://datatracker.ietf.org/drafts/current/. Internet-Drafts are draft documents valid for a maximum of six months and may be updated, replaced, or obsoleted by other documents at any time. It is inappropriate to use Internet-Drafts as reference material or to cite them other than as "work in progress." This Internet-Draft will expire on 3 April 2027. Copyright Notice Mo, et al. Expires 3 April 2027 [Page 1]

Internet-Draft Agent Selection Mapping September 2026 Copyright (c) 2026 IETF Trust and the persons identified as the document authors. All rights reserved. This document is subject to BCP 78 and the IETF Trust's Legal Provisions Relating to IETF Documents (https://trustee.ietf.org/ license-info) in effect on the date of publication of this document. Please review these documents carefully, as they describe your rights and restrictions with respect to this document. Code Components extracted from this document must include Revised BSD License text as described in Section 4.e of the Trust Legal Provisions and are provided without warranty as described in the Revised BSD License. Discussion Venues Discussion of this document takes place on the CATS Working Group mailing list (cats@ietf.org), which is archived at https://mailarchive.ietf.org/arch/browse/cats/. Table of Contents Mo, et al. Expires 3 April 2027 [Page 2]

Internet-Draft Agent Selection Mapping September 2026 1. Introduction . . . . . . . . . . . . . . . . . . . . . . . . . 4 2. Conventions and Definitions . . . . . . . . . . . . . . . . . .5 3. Terminology . . . . . . . . . . . . . . . . . . . . . . . . . .5 4. Problem Statement . . . . . . . . . . . . . . . . . . . . . . .7 5. Agent Service Properties and Their Selection Implications . . .8 5.1. Long-Horizon Sessions . . . . . . . . . . . . . . . . . . 8 5.2. Stateful Execution . . . . . . . . . . . . . . . . . . . .9 5.3. Discrete Progress . . . . . . . . . . . . . . . . . . . . 9 5.4. Heavy-Tailed Consumption . . . . . . . . . . . . . . . . .9 6. Selection Mapping Framework . . . . . . . . . . . . . . . . . .9 6.1. Framework Entities and Information Flow . . . . . . . . . 9 6.2. Layer 1: Agent Service Descriptor . . . . . . . . . . . .10 6.3. Layer 2: Resource View . . . . . . . . . . . . . . . . . 12 6.4. Layer 3: Selection Context . . . . . . . . . . . . . . . 13 6.5. Decision Artifacts . . . . . . . . . . . . . . . . . . . 14 6.6. Characteristic-to-Resource Mapping . . . . . . . . . . . 14 6.7. Characteristic-to-Resource Routing . . . . . . . . . . . 18 6.8. Resource-Level Routing Recipe . . . . . . . . . . . . . .21 7. Information Elements and Their Abstract Syntax . . . . . . . .22 7.1. Descriptor Elements . . . . . . . . . . . . . . . . . . .23 7.2. Resource View Elements . . . . . . . . . . . . . . . . . 24 7.3. Accounting Elements . . . . . . . . . . . . . . . . . . .25 7.4. Exchange Metadata Elements . . . . . . . . . . . . . . . 25 7.5. Semantics Rules Common to the Elements . . . . . . . . . 26 8. Selection Procedure . . . . . . . . . . . . . . . . . . . . . 27 8.1. Phases . . . . . . . . . . . . . . . . . . . . . . . . . 27 8.2. Decision Points . . . . . . . . . . . . . . . . . . . . .29 8.3. Re-Evaluation Triggers and Stability . . . . . . . . . . 29 8.4. Fallback Ladder . . . . . . . . . . . . . . . . . . . . .30 9. Message Semantics . . . . . . . . . . . . . . . . . . . . . . 32 9.1. Exchange Catalogue . . . . . . . . . . . . . . . . . . . 32 9.2. Exchange Sequences . . . . . . . . . . . . . . . . . . . 34 9.3. Rules Common to the Exchanges . . . . . . . . . . . . . .35 10. Aggregation, Normalization, and Freshness . . . . . . . . . .36 11. Requirements . . . . . . . . . . . . . . . . . . . . . . . . 36 11.1. Framework and Descriptor Requirements . . . . . . . . . 37 11.2. State and Storage Requirements . . . . . . . . . . . . .38 11.3. Discrete Decision Requirements . . . . . . . . . . . . .39 11.4. Tail and Long-Horizon Requirements . . . . . . . . . . .39 11.5. Exchange Requirements . . . . . . . . . . . . . . . . . 40 11.6. Mapping Requirements . . . . . . . . . . . . . . . . . .40 12. Relationship to the CATS Framework . . . . . . . . . . . . . 41 12.1. Traceability to the Agent Service Requirements . . . . .41 13. Operational Considerations . . . . . . . . . . . . . . . . . 42 14. Security Considerations . . . . . . . . . . . . . . . . . . .43 15. Privacy Considerations . . . . . . . . . . . . . . . . . . . 43 16. IANA Considerations . . . . . . . . . . . . . . . . . . . . .44 17. Normative References . . . . . . . . . . . . . . . . . . . . 44 18. Informative References . . . . . . . . . . . . . . . . . . . 44 Acknowledgments . . . . . . . . . . . . . . . . . . . . . . . . . 45 Authors' Addresses . . . . . . . . . . . . . . . . . . . . . . . .45 Mo, et al. Expires 3 April 2027 [Page 3]

Internet-Draft Agent Selection Mapping September 2026

1. Introduction

A CATS system classifies the traffic of a service request, selects a service contact instance, and steers the traffic of the request towards the selected instance [I-D.ietf-cats-framework]. The selection is made by the CATS Path Selector (C-PS) using computing metrics collected by a CATS Service Metric Agent (C-SMA) and network metrics collected by a CATS Network Metric Agent (C-NMA). The metrics that are exchanged for that purpose are specified in [I-D.ietf-cats-metric-definition] at three levels of abstraction. The use cases and requirements document already anticipates distributed AI training and inference as a use case of CATS, and notes that the resources that matter for inference include processor cores and the memory used for cache [I-D.ietf-cats-usecases-requirements]. For agent services, two properties of that model need to be extended without changing it. First, the unit that arrives at the network is a session or a step within a session, and the requirements of a step are richer than a request for a service: they include a capability constraint, a state affinity, a budget, and a locality constraint. Second, the state that determines how quickly a step can be served is not part of the current metric view, as has been observed for the specific case of a key-value (KV) cache [I-D.li-cats-kv-cache-distribution]. This document describes how the characteristics of an agent step [I-D.mo-cats-agent-service-characteristics] are mapped onto a selection decision over three dimensions -- forwarding, computing, and storage -- using the existing CATS components. The mapping framework is intended to be used as follows: * as input to a future extension of the metric definition, by identifying the information that the mapping needs; * as input to the data model work [I-D.ietf-cats-data-model], by identifying the objects that configuration and monitoring must cover; * as a description of a selection procedure that an implementation can follow today using metrics that are already defined. Mo, et al. Expires 3 April 2027 [Page 4]

Internet-Draft Agent Selection Mapping September 2026 The mapping is written for four properties of agent services. an agent session is long-horizon, stateful, discrete in the way that it makes progress, and heavy-tailed in what it consumes. Section 5 states what each property requires of a selection procedure. Sections 6, 7, 8, and 9 give the framework, the information elements and their abstract syntax, the procedure with its decision points, and the exchanges through which the information reaches the selection function. Section 11 collects the requirements. The intended standing of this document is informational groundwork in the sense of the CATS charter [CATS-CHARTER]: it states what a CATS system has to be able to express, so that the work can be taken up by the metric definition and the data model rather than by a protocol extension. The resource inventory of Section 6.7 is what a metric framework would have to expose, and the elements of Section 7 are what a data model would have to accommodate. This document defines no encoding and no protocol.

2. Conventions and Definitions

The key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT", "SHOULD", "SHOULD NOT", "RECOMMENDED", "NOT RECOMMENDED", "MAY", and "OPTIONAL" in this document are to be interpreted as described in BCP 14 [RFC2119] [RFC8174] when, and only when, they appear in all capitals, as shown here. In this document these key words are used to state requirements on the design of a CATS system that supports agent services. They do not describe protocol behavior, and this document defines no protocol, message format, or data model.

3. Terminology

This document uses the terms of [I-D.ietf-cats-framework] and [I-D.mo-cats-agent-service-characteristics]. The following additional terms are used. Agent Service Descriptor: The representation of what a step of an agent session needs, expressed as a set of constraints and preferences. Mo, et al. Expires 3 April 2027 [Page 5]

Internet-Draft Agent Selection Mapping September 2026 Resource View: The set of properties of a candidate instance and of the paths towards it, grouped into the forwarding, computing, and storage dimensions, that are used to evaluate a step. Affinity: A preference for a candidate instance on the grounds that it already holds session state, model artifacts, or other reusable material. Selection Point: The point in the execution of a session at which a selection is made, such as the start of a session, the start of a turn, the start of a step, or a re-evaluation triggered by a change in the resource view. Decision Point: A point in the life of a session at which the mapping is required to produce a selection, such as session admission, the start of a turn, a step boundary, the detection of a tail event, or a change of the resource view that invalidates the current selection. Decision Record: The artifact that records the descriptor, the resource view, the trigger, and the outcome of one application of the procedure, for audit, for accounting, and as input to later decisions of the same session. Working Set Class: A category of reusable state with a distinct cost of reconstruction, such as model-side context state, retrieved material, tool results, plan and scratch state, and longer-term memory. State Handle: An opaque reference to a working set element that allows a component to test presence, request a transfer, or attach affinity to the element without learning its content. Reuse Key: The identifier under which a working set element can be addressed for the purpose of presence testing, transfer, or reuse. Residency Tier: The memory tier at which a working set element currently resides at an instance, for example accelerator memory, host memory, local storage, or a remote store. Mo, et al. Expires 3 April 2027 [Page 6]

Internet-Draft Agent Selection Mapping September 2026 Stability Window: The minimum interval, or the minimum improvement of suitability, that must hold before a different instance may be selected for a session that is already placed. Commitment: The binding of a step, or of a range of steps, to a selected instance, together with the state and the budget that the binding assumes. Handoff: The planned move of a session, or of part of its working set, from one instance to another, including reservation, transfer, and commit or abort. Tail Event: An execution whose cost or duration lies in the upper tail of the distribution for comparable steps, for example a tool call that takes much longer than the same tool normally takes. Fallback Ladder: The ordered set of behaviors to which the procedure degrades when no candidate satisfies all constraints, when the selected instance fails, or when a tail event exhausts the budget of a step.

4. Problem Statement

The mapping exists because the current selection input and the current demand do not align. First, the mapping from demand to metrics is implicit. Today the C-PS selects on metrics that describe instances and paths, but the relationship between what a step needs and those metrics is not modeled: for a step that must reason over a 200 kB context, must not send that context outside a region, and must complete within two seconds, it is not obvious which metric at which level expresses the need, or which part of the need is a constraint rather than a preference. Second, a single scalar cannot represent the decision. A normalized score ranks candidates, but it cannot express that an instance is ineligible because it cannot serve the required capability, nor that two eligible instances differ in whether they already hold the session state. Mixing a hard constraint into a score either violates the constraint or distorts the ranking. Mo, et al. Expires 3 April 2027 [Page 7]

Internet-Draft Agent Selection Mapping September 2026 Third, the decision is discrete and its lifetime is not uniform, while the quantities that inform it are continuous and heavy-tailed. Placement changes slowly, because it may involve loading a model or migrating state, whereas traffic allocation and per-step routing change quickly [I-D.luan-cats-catpts], so a single validity rule for the whole decision is wrong. A session places a step, executes it, and then places the next one, so a mapping that reacts to every change of every quantity moves a long autonomous loop for gains that do not survive the move. At the same time the objective that matters is usually a percentile: an objective stated as a mean hides exactly the sessions that dominate cost and user-visible latency [I-D.mo-cats-agent-service-characteristics].

5. Agent Service Properties and Their Selection Implications

The four properties below are taken from the agent service characteristics [I-D.mo-cats-agent-service-characteristics]. Each constrains the mapping in a way that the current selection function does not face. Property | Required of the mapping | Where it is | | handled -------------+--------------------------------+------------------- Long-horizon | Decide for a session of many | 6.6 mapping; 6.4 | steps rather than for a single | commitment policy; | request, and keep decisions | 8.3 stability | stable while it runs | Stateful | Treat reusable state as both a | 6.3 and 6.6; 8.1 | constraint and a quantity, and | and 9 handoff | plan the move of that state | | before committing to it | Discrete | Decide at step boundaries on | 6.6 mapping; 8.2 | banded values, and change an | decision points; | instance only when a | 8.3 stability | constraint requires it or the | | gain is durable | Heavy-tailed | Evaluate objectives at high | 6.6 mapping; 8.1 | percentiles, and keep the | valuation; 8.4 | slowest call of a turn from | fallback | dictating the objective |

5.1. Long-Horizon Sessions

Mo, et al. Expires 3 April 2027 [Page 8]

Internet-Draft Agent Selection Mapping September 2026 A session runs for tens or hundreds of steps and pauses for minutes while a tool runs, so the decision outlives the snapshot that produced it without being valid for the whole session: establishing a placement that loads a model artifact costs seconds to minutes, while a per-step routing decision is revised in milliseconds [I-D.mo-cats-agent-service-characteristics]. The descriptor therefore carries the session objective, the remaining budget, the remaining horizon, and the freshness need of each element, and a wake-up after a pause re-checks whether the state that the session depends on is still there.

5.2. Stateful Execution

State is a first-class input rather than an optimization. Its classes differ by orders of magnitude in size and in reconstruction cost, so evaluation is per element and uses the accumulated working set rather than the working set of the first step. Where a candidate does not report whether an element is present, the presence is unknown, and an unknown value does not satisfy a constraint.

5.3. Discrete Progress

The steps of a session are not known when it starts, and one step can fan out to several endpoints. The mapping therefore faces discrete decisions rather than a continuous allocation: it acts at step boundaries, compares banded values, and changes an instance only when a constraint requires it or the gain survives both the cost of the change and the stability requirement.

5.4. Heavy-Tailed Consumption

Work is skewed across sessions and the objective is usually a percentile, so the mapping needs tail-aware valuation, that is, comparison at a stated percentile over a stated window, and tail protection: detecting a running step that exceeds its expected cost and degrading it, or failing it with a report, in the order of the fallback ladder [I-D.pang-cats-fallback-decision-framework].

6. Selection Mapping Framework

The mapping framework has three layers.

6.1. Framework Entities and Information Flow

The mapping is performed by the components that already exist in the CATS framework. No new functional component is introduced; what the mapping adds is the information that flows between them and the times at which they act. Mo, et al. Expires 3 April 2027 [Page 9]

Internet-Draft Agent Selection Mapping September 2026 agent gateway | descriptor: capability, working set, constraints, objective, | budget, tail target; step boundary; wake-up after a pause v C-TC -- classify and attribute to session / turn / step-------+ | | v | C-PS - derive descriptor -> apply constraints -> value dims---+ | ^ ^ ^ | | | | | | | C-SMA: capability C-SMA: state C-NMA: paths | | load, tail presence, tier cost, prefixes | v | forwarders <---- steering instruction-------------------------+ | v decision record -> budget accounting -> next decision point * Agent gateway: terminates the session and derives the descriptor where the network cannot; it supplies the descriptor, the step boundary, the wake-up after a pause, and the budget. * C-TC: classifies session traffic, attributes it to the session, the turn, and the step, and carries the attribution and the descriptor keys. * C-PS: applies constraints, values the dimensions, combines them, commits, and produces the selection and the decision record. * C-SMA: reports capability, load, and state presence for its site, which is the computing and state view of a candidate. * C-NMA: reports paths and the per-prefix cost towards endpoints, which is the forwarding view of a candidate. * Forwarders: execute the steering decision for a step, and report the progress that indicates a tail event.

6.2. Layer 1: Agent Service Descriptor

The descriptor is derived from the step to be served and, where unavailable from the network, is provided by the agent gateway that terminates the session. It contains: * the identity of the service, expressed as the CATS Service Identifier (CS-ID), and, where known, the session and step identities used for attribution; * the capability required by the step, expressed in terms that can be matched against the capability offered by an instance, such as the model capability tier, the required context length, and the required tool or data reachability; Mo, et al. Expires 3 April 2027 [Page 10]

Internet-Draft Agent Selection Mapping September 2026 * the working set that the step needs, expressed as the set of state elements that, if present, would allow the step to be served without re-computation, together with the keys under which those elements can be addressed; * the tolerance of the step, expressed as whether its traffic may be buffered, delayed, or carried at reduced precision when the network is congested, in the terms used by the forwarding dimension; * the constraints of the step, including locality and governance constraints that apply to the input data and to the working set, the tenant to which the session belongs, and any hard ceiling on cost or energy; * the preferences of the step, including the latency objective, the willingness to trade precision or delay, and the relative importance of state reuse against load balancing. Beyond the elements above, the descriptor carries the quantities that let the mapping work over a long session rather than over a single request: * the objective of the session and the horizon over which it is evaluated, where the gateway can estimate that horizon; * the remaining budget of the session, and whether the budget is a ceiling that may not be exceeded or a target that may be missed at a stated cost; * the tail objective of the step, expressed as a percentile, the window over which the percentile is taken, and the maximum probability of exceeding the objective that the session accepts; * the degradation class of the step, that is, which reductions the step tolerates: lower precision, a lower capability tier, a truncated working set, or additional delay; Mo, et al. Expires 3 April 2027 [Page 11]

Internet-Draft Agent Selection Mapping September 2026 * the freshness tolerance of the step, that is, how old a value may be for a decision on that step to remain valid. The descriptor is a local representation. Whether and how any element of it is carried in packets or signaled between components is explicitly out of scope for this document.

6.3. Layer 2: Resource View

The resource view is the three-dimensional description of a candidate instance and of the paths towards it. Each dimension is composed of constraints and of quantities. Dimension | Constraint elements | Quantity elements -------------+--------------------------------+------------------- Forwarding | path policy, determinism | latency, jitter, | class, reachability of tools | loss, bandwidth, | and data sources, permitted | path diversity, | transfer window | per-prefix cost to | | tool and data | | endpoints Computing | capability tier, model and | queueing delay, | adapter versions, precision | free accelerator | support, maximum context | capacity, memory | length | headroom, | | comparable-step | | latency at median | | and p95 Storage | permitted location per | presence per reuse | element, tenant isolation, | key, residency | sharing scope, retention | tier, retrieval | requirement | cost, recompute | | cost, freshness, | | element size The storage dimension is evaluated per working set element rather than per instance, because the elements of one working set differ in class, in size, and in the cost of reconstructing them. The classes that the mapping distinguishes are the following. * Model-side context state: the state that allows earlier context to be reused without reprocessing it, for example a key-value cache or a cached prefix. It is large, it grows with the context, and it may be evicted during a pause. Mo, et al. Expires 3 April 2027 [Page 12]

Internet-Draft Agent Selection Mapping September 2026 * Retrieved material: the items that a step has retrieved, for example documents or passages. It is small or moderate in size and cheap to re-retrieve while the retrieval endpoint is reachable. * Tool results: the outputs of tool invocations. They are small, and re-obtaining one may be impossible because the tool is not idempotent, which makes their presence a hard requirement rather than a preference. * Plan and scratch state: the intermediate reasoning and bookkeeping of the current task. It is small, but it is the most expensive class to reconstruct, because reconstruction requires the earlier steps to be re-executed. * Longer-term memory: state that outlives the session, for example a profile or a set of preferences. It is small, it is read-mostly, and it is frequently subject to a location constraint. The forward, computing, and storage dimensions are evaluated for the same candidate instance. A candidate is a pair of a service contact instance (identified by the CATS Service Contact Instance ID, CSCI-ID) and the path (or set of paths) towards it.

6.4. Layer 3: Selection Context

The selection context records when and how often the mapping is applied. It contains the Selection Point, the objective of the session as a whole, the remaining budget of the session, and the instruction that determines whether a change of instance is permitted. This layer exists because the same descriptor evaluated at the start of a session and in the middle of a turn can produce different results, and because moving a session between instances has a cost that must be accounted for. The selection context is also where the discrete and the long-horizon properties meet the per-step decision. It therefore carries: * the commitment policy, that is, which parts of the decision may change within a step, between the steps of a turn, and between turns, and which parts are held for the life of the session; * the stability requirement in force, that is, the minimum interval and the minimum improvement of suitability that must hold before a different instance may be selected; Mo, et al. Expires 3 April 2027 [Page 13]

Internet-Draft Agent Selection Mapping September 2026 * the banding policy, that is, the coarseness at which continuous quantities are compared, which bounds the rate at which decisions can change; * the look-ahead horizon, that is, the number of future steps over which the objective is considered when a choice made now would be expensive to reverse; * the budget pacing rule, that is, whether the session may spend its budget unevenly across steps, and how much of the remaining budget is reserved for the steps that are expected to be the most expensive.

6.5. Decision Artifacts

Five artifacts are produced or consumed by one application of the procedure. Their separation matters because they have different lifetimes, and mixing them is what makes a selection oscillate. The decision record is also the artifact through which an operations and management view of CATS can observe the mapping [I-D.ietf-cats-oam-fw]. * Descriptor: produced by the agent gateway or by C-TC; it lives for the step, and it is the input to one or more evaluations. * Resource view snapshot: produced by C-SMA and C-NMA; it lives until its freshness bound expires, and it is the input to an evaluation. * Decision record: produced by C-PS; it lives for the session, and it serves audit, accounting, and the later decisions of the same session. * Commitment: produced by C-PS; it lives until it is revised at a decision point, and it binds a step to an instance. * Handoff record: produced by C-PS; it lives from the reservation to the commit or the abort, and it drives the movement of state.

6.6. Characteristic-to-Resource Mapping

This subsection is the core of the mapping. It states, for each characteristic of an agent service, how the characteristic appears in the three dimensions at selection time and what the selection function has to do about it. The dimensions are not symmetric at selection time. * Transfer is paid once per change of instance, and it is paid in bulk: what matters is the bytes that have to be moved before the Mo, et al. Expires 3 April 2027 [Page 14]

Internet-Draft Agent Selection Mapping September 2026 step can run, not the bytes of the step itself. Transfer cost appears in the forwarding dimension, which is where the paths and their cost are evaluated. * Storage is paid per element and per step: what matters is which working set elements are present, where they reside, and what it costs to make them usable. * Compute is paid per step: what matters is the capability that the step needs and the capacity that is free for it when it arrives. Table 1 maps each characteristic onto the signal that each dimension observes. Table 2 maps it onto the decision, and states whether the characteristic acts as a constraint or as a preference. Table 3 gives the order of magnitude that the dimensions have to compare; the values are measured in [I-D.mo-cats-agent-service-characteristics], and they are indicative rather than normative. Mo, et al. Expires 3 April 2027 [Page 15]

Internet-Draft Agent Selection Mapping September 2026 Characteristic | Transfer | Storage | Compute -----------------+---------------+---------------+---------------- Session state | bytes moved | presence per | recompute cost | per change; | reuse key; | as the | per-prefix | tier; | alternative to | cost to the | retrieval | transfer | state source | cost | Multi-step and | per-endpoint | tool results | tier needed by tools | path cost; | retained per | this step | fan-out to | key | | several | | | prefixes | | Long-horizon and | transfer | retention | re-prefill cost pauses | window after | against | after eviction | a wake-up | eviction in a | | | pause | Context growth | bulk prefix | context state | prefill for the | transfer; | grows with | retained prefix | compression | tokens | | cuts bytes | | Capability tiers | not | artifact | tier match, as | applicable | residency per | a constraint | | tier | Communication | flow shape: | state updated | concurrency the mode | streaming, | per response | step needs | fan-out, | | | burst | | Memory | reachability | long-lived | lookup cost per persistence | of the memory | elements; | step | store | sharing scope | Locality and | permitted | permitted | permitted tier tenancy | paths and | location per | per element | regions | element | Budget | bytes charged | storage time | compute time | to the | charged | charged | session | | Turn and task | p95 path | p95 retrieval | p95 queue and objective | delay over | delay | step latency | the window | | Model artifact | bulk load | residency and | load time size | transfer, | headroom for | before the | seconds to | weights | first token | minutes | | Mo, et al. Expires 3 April 2027 [Page 16]

Internet-Draft Agent Selection Mapping September 2026 Characteristic | Decision rule at | Role | Req. | selection time | | ----------------+--------------------------+------------+--------- Session state | prefer the instance | preference | M4, M5, | holding the largest | | M23, M27 | share of the working | | | set, unless transfer or | | | recompute costs more | | | than the gain | | Multi-step and | value the path to each | both | M3, M34, tools | endpoint and bound the | | M40 | turn by its slowest call | | Long-horizon | re-verify state presence | preference | M25, and pauses | after a pause, and keep | | M31, M37 | the commitment stable | | | across steps | | Context growth | require memory headroom | constraint | M22, M24 | for the accumulated | | | context state, not for | | | the prompt | | Capability | exclude instances below | constraint | M6, M19 tiers | the required tier; do | | | not trade the tier for | | | distance | | Communication | keep a streaming | both | M29, M42 mode | response on one | | | instance; let fan-out | | | paths differ | | Memory | prefer proximity to the | preference | M23, M24 persistence | memory store | | Locality and | filter before ranking; | constraint | M6, M20, tenancy | an element that may not | | M28 | move stays where it is | | Budget | charge the cost of the | both | M15, | change to the session | | M32, M38 | budget before moving | | Turn and task | compare at the stated | preference | M17, objective | percentile over the | | M33, M34 | stated window | | Model artifact | exclude instances that | both | M2, M24, size | cannot hold weights and | | M32 | state, then account load | | | time in the change cost | | Mo, et al. Expires 3 April 2027 [Page 17]

Internet-Draft Agent Selection Mapping September 2026 Anchor quantity | Value and unit | Source -----------------+----------------------------+---------------- Context state | 55 MiB per 1k tokens at | Section 5.4 size | 7B; 313 MiB at 72B, half | | precision | Model artifacts | 0.88 to 688 GB over | Section 6.1 | fourteen models; 15.23 GB | | to 5.57 GB at 4-bit | Bulk transfer | seconds to minutes per | Section 6.1 cost | change of instance, not | | milliseconds | Tool call | heterogeneous and | Section 5.2 latency | heavy-tailed; the slowest | | call bounds the turn | Objective | a percentile over a | Section 5.3, statement | window; a mean is not a | R17 | substitute | Three rules follow from the tables and apply to every characteristic. * A characteristic that cannot be observed in a dimension leaves that dimension unknown for the decision. Unknown is not a favorable value, and a constraint that rests on it is not satisfied. * A characteristic that appears as a constraint is filtered before ranking; the same characteristic may appear as a preference in another session, and the descriptor states which it is. * A cost that is paid once per change is compared with the remaining budget and the remaining horizon, not with the gain of the single step that triggers it. Section 6.7 takes the same characteristics to the level of the concrete resources that have to be compared.

6.7. Characteristic-to-Resource Routing

Section 6.6 maps a characteristic onto a dimension. This subsection takes the same characteristics to the level of the concrete resources that a decision has to compare, because that is the level at which a CATS system can route: a dimension is not observable, whereas accelerator memory free, path egress, and state residency are. Mo, et al. Expires 3 April 2027 [Page 18]

Internet-Draft Agent Selection Mapping September 2026 The resource list below is also the list that a metric framework would have to expose for this mapping to be operational. Stating it at the level of the physical resource and its unit is what allows the requirements of Section 11 to be met by a metric definition and a data model rather than by an implementation. Table 4 is the resource inventory. T, S, and C denote the transfer, storage, and compute dimensions. Resource | Dim. | Observed as | Observed | | | by ------------------+------+----------------------------+----------- Path egress | T | available bytes/s towards | C-NMA capacity | | the candidate | Path delay and | T | ms, median and p95 | C-NMA jitter | | | Endpoint cost | T | per-prefix ms to a tool or | C-NMA | | data endpoint | Transfer window | T | class and permitted start | operator | | window | Accelerator | C/S | bytes free for weights and | C-SMA memory free | | state | Host memory free | S | bytes free | C-SMA Local persistent | S | bytes free and read MB/s | C-SMA store | | | Remote state | S | RTT and throughput | C-SMA store | | | State residency | S | tier per reuse key | C-SMA State retention | S | seconds the instance will | C-SMA | | keep it | Accelerator | C | type, count, precision set | C-SMA capability | | | Waiting work | C | queue depth and p95 wait | C-SMA Prefill and | C | available tokens/s | C-SMA decode rate | | | Context window | C | tokens the instance can | C-SMA | | hold | Artifact locality | C/S | registry RTT and load MB/s | C-SMA Table 5 states, for each characteristic, which of those resources carries it and what the routing step does about it. Mo, et al. Expires 3 April 2027 [Page 19]

Internet-Draft Agent Selection Mapping September 2026 Characteristic | Resource that | Routing action | carries it | ----------------+-------------------+-------------------------- Session state | state residency, | route to a holder when | retention, | the state fits free | accelerator | memory, else compare | memory | transfer against | | recompute and take the | | cheaper Accumulated | accelerator | require free memory for context | memory, prefill | weights plus context | rate | state; if unmet, degrade | | or choose a larger-memory | | candidate Tool fan-out | endpoint cost, | value the path to each | path egress | endpoint; bound the turn | | by the slowest call; let | | fan-out paths differ Long horizon | retention, host | prefer retention that | memory | covers the expected idle; | | keep the commitment | | stable across steps Wake-up after a | state residency, | re-check residency; treat pause | retention | an evicted element as | | absent and re-run the | | comparison Capability tier | accelerator | exclude on type, | capability, | precision, or window; do | context window | not trade against | | distance Compute load | waiting work, | prefer the lowest p95 | prefill rate | wait at equal capability, | | and re-check at each step | | boundary Bulk artifact | artifact | charge the load time to | locality, path | the change cost; exclude | egress | candidates whose load | | exceeds the step budget Governance | state residency, | filter before ranking; | permitted regions | keep an element that may | | not move in its permitted | | set Budget | all three | charge transfer once per | dimensions | change, storage per | | element-step, compute per | | step; refuse a move | | beyond the budget Tail objective | path delay, | compare at p95 over the | waiting work, | window; exclude | retrieval | candidates whose tail | | breaches a constraint Mo, et al. Expires 3 April 2027 [Page 20]

Internet-Draft Agent Selection Mapping September 2026 The routing step uses the comparisons below. They are written as relations between resource quantities to state what has to be comparable. They define no encoding, and the values used later are illustrative. Resource-side quantities used by the routing step: state_bytes | size of the working set that must be usable weights_bytes | size of the artifact that must be resident hbm_free | accelerator memory free at the candidate bw | available path egress towards the candidate rtt | path delay towards the candidate prefill_rate | context tokens processable per second tokens_to_reprocess | tokens a re-computation would redo wait_p95 | p95 waiting time for a comparable step step_latency_p95 | p95 service time for a comparable step Derived comparisons, illustrative and not normative: state_fits | state_bytes <= hbm_free - weights_bytes transfer_time | state_bytes / bw + rtt recompute_time | tokens_to_reprocess / prefill_rate state_cost | min(transfer_time, recompute_time) tail_ok | wait_p95 + step_latency_p95 <= objective budget_ok | session_cost + change_cost <= budget

6.8. Resource-Level Routing Recipe

1. Fit. Test the memory and capability relations at every candidate: drop the candidates where the working set cannot be made usable, and the candidates whose context window is too small for the step. 2. Cost of not holding state. Compute the state cost of each survivor, and drop the candidates whose state cost exceeds the budget of the step. 3. Path. Value the endpoints that this step will contact, not only the path to the instance; the slowest intended call bounds the turn. 4. Tail and budget. Test the tail objective at the stated percentile over the stated window, and test the cost of the change against the session budget that the change is intended to serve. 5. Commit through Section 8.3. The resource comparison does not select on its own: a candidate that wins on resources is still subject to the stability requirement, because moving a session costs state and budget. Mo, et al. Expires 3 April 2027 [Page 21]

Internet-Draft Agent Selection Mapping September 2026 When step 1 or step 2 leaves no candidate, the degradation classes of the step are applied in the order of Section 8.4 before the session is failed. Step 30 of the session of Section 8.5, with illustrative values. Candidate C holds the state: hbm_free 20 GB, weights_bytes 15.23 GB, state_bytes 7.5 GB state_fits is true (7.5 <= 20 - 15.23) Candidate B does not hold it, but has a better path: bw 2.5 GB/s, rtt 8 ms, prefill_rate 1,000 tokens/s transfer_time = 7.5 GB / 2.5 GB/s + 8 ms ~ 3.0 s recompute_time = 128,000 tokens / 1,000 tokens/s ~ 128 s state_cost at B ~ 3.0 s Transfer is far cheaper than re-computation here, so holding the state does not by itself pin the session to C. The routing step therefore falls through to path and budget: B wins only if its p95 path and its remaining budget are better by more than the stability requirement of Section 8.3 allows.

7. Information Elements and Their Abstract Syntax

This section states what the mapping has to be able to express. It defines no concrete syntax: no encoding, no field layout, and no registry is defined here, and the concrete representation of these elements belongs to the data model work [I-D.ietf-cats-data-model]. What is defined here is the meaning of each element, the abstract class of its value, the scope in which it applies, and what the absence of the element means. The classes below are abstract. They exist so that the semantics of an element can be stated without choosing a representation. Mo, et al. Expires 3 April 2027 [Page 22]

Internet-Draft Agent Selection Mapping September 2026 Class | Domain | Typical use -----------+----------------------------+------------------------- BOOL | true or false | may this traffic be | | buffered ENUM | one of a named set | capability tier, | | rejection reason RANGE | [lo, hi] with a stated | context length, budget | unit | limit COUNT | non-negative integer | remaining steps of a | | horizon QUANTILE | value at rank q over | p95 step latency | window W | TIME | instant, UTC | collection time of a | | value DURATION | length of an interval | state age, freshness | | bound SIZE | bytes or tokens | element size, working | | set size HANDLE | opaque token, compared for | state handle, reuse key | equality | SET | unordered, no duplicates | elements of a working | | set MAP | key to value | reuse key to presence TEXT | short string for humans | reason code, operator | | label

7.1. Descriptor Elements

Mo, et al. Expires 3 April 2027 [Page 23]

Internet-Draft Agent Selection Mapping September 2026 Element | Class | Semantics ---------------------+-----------+-------------------------------- service-id | HANDLE | the CS-ID of the service, | | session scope session-id | HANDLE | attribution key of the session turn-id | COUNT | attribution key of the turn step-id | COUNT | attribution key of the step capability-tier | ENUM | the lowest tier the step may be | | served at required-precision | SET | precision formats that are | | acceptable context-length | RANGE | tokens that the step keeps in | | context working-set | SET | the elements needed, each with | | a reuse key element-class | ENUM | class of an element, per | | Section 6.3 locality | MAP | element class to permitted | | regions tenant | HANDLE | tenant under which the session | | runs budget | MAP | dimension to limit, ceiling or | | target objective | QUANTILE | quantity, percentile, window, | | weight tail-tolerance | RANGE | probability of missing the | | objective degradation | SET | reductions that the step | | tolerates freshness-need | MAP | element to the maximum usable | | age horizon | COUNT | steps that the session is | | expected to run mode | ENUM | request-response, streaming, | | tool call, fan-out

7.2. Resource View Elements

Element | Class | Semantics ---------------------+-----------+-------------------------------- path-set | SET | the candidate paths towards the | | instance endpoint-cost | MAP | per-prefix cost to tools and | | data sources determinism | ENUM | whether the path offers bounded | | delay transfer-window | DURATION | time in which bulk transfer may | | start Mo, et al. Expires 3 April 2027 [Page 24]

Internet-Draft Agent Selection Mapping September 2026 Element | Class | Semantics ---------------------+-----------+-------------------------------- tier-offered | ENUM | the highest tier the instance | | can serve queue-delay | QUANTILE | wait for a comparable step, at | | p95 memory-headroom | SIZE | accelerator memory free for | | state step-latency | QUANTILE | recent comparable steps, median | | and p95 Element | Class | Semantics ---------------------+-----------+-------------------------------- presence | MAP | reuse key to present, absent, | | or unknown residency | ENUM | tier at which an element | | resides element-size | SIZE | size of the element as it would | | be moved retrieval-cost | DURATION | time to make the element usable | | there recompute-cost | DURATION | time to rebuild it without | | transfer state-age | DURATION | age of the element as held at | | the instance retention | DURATION | how long the instance will keep | | it governance | MAP | element to the constraint that | | applies

7.3. Accounting Elements

Element | Class | Semantics ---------------------+-----------+-------------------------------- consumed | MAP | tokens, time, and cost consumed | | so far remaining | MAP | limit minus consumed, per | | dimension reserve | MAP | budget held back for the | | remaining steps attribution | HANDLE | session, turn, and step to | | charge

7.4. Exchange Metadata Elements

Every element that moves between components carries the metadata below. Mo, et al. Expires 3 April 2027 [Page 25]

Internet-Draft Agent Selection Mapping September 2026 Element | Class | Semantics ---------------------+-----------+-------------------------------- collected-at | TIME | when the value was observed freshness-bound | DURATION | how long the value may be used source | HANDLE | the component that asserted the | | value confidence | ENUM | asserted, derived, or estimated

7.5. Semantics Rules Common to the Elements

The following rules apply to every use of the elements above. They are the part of this document that a concrete data model has to preserve. * Absence of a value from the descriptor, the resource view, or an exchange means that the value is unknown. Unknown is not a value: a constraint that cannot be evaluated is not satisfied. * Every quantity carries the time at which it was observed and a freshness bound. A value whose age exceeds its bound must not be used to satisfy a constraint, and its use in a preference must be reported in the decision record. * A quantity that is a percentile states the rank and the window. A quantity that carries no rank is a mean and must be identified as a mean. * The scope of an element is explicit. A constraint that is scoped to the session applies to every step of that session, and a constraint that is scoped to a step does not outlive it. * Constraint and quantity are separate roles. The same physical property may appear in both roles with different semantics, as when a delay is a hard ceiling in one descriptor and a preference in another. * A state handle is opaque. Testing the presence of an element, requesting its transfer, and attaching affinity to it must not require the network to learn its content, and two equal handles must denote the same element. Mo, et al. Expires 3 April 2027 [Page 26]

Internet-Draft Agent Selection Mapping September 2026 * Consumption is monotone. A budget decreases as a session runs, is attributed to a session, a turn, and a step, and a ceiling cannot be traded against a preference. * Banding is stated, not implied. Where values are compared in bands, the band edges are known to the decision, and banding must not hide a violation of a constraint. * Units and populations are explicit. A quantity states its unit, and an aggregated quantity states the class, the key space, and the number of instances that it covers.

8. Selection Procedure

This section states the phases through which the mapping runs, the points at which a decision is required, the events that cause a re-evaluation, and the order in which the procedure degrades when it cannot proceed as intended. The phases are written in operational terms, and they assume that the information of Section 7 is available to the C-PS.

8.1. Phases

P0. Session admission. Input: the session objective, the budget, and the standing constraints. Action: establish that at least one candidate can serve the session, and create the selection context. Output: the session context and an initial commitment. On failure: the fallback ladder of Section 8.4. P1. Step boundary detection. Input: the traffic of the session, and, where available, the step signal from the agent gateway. Action: determine that a new step begins, and record the boundary. Output: a step identifier and the accumulated working set. On failure: treat the turn as a single step, which forbids per-step migration and is therefore the conservative choice. P2. Descriptor derivation. Input: the step, the accumulated state, and the session context. Action: derive the descriptor, and request from the gateway any element that the network cannot derive (MSG-3). Output: the descriptor for the step. On failure: if a required element cannot be obtained, the step is not evaluable and Section 8.4 applies. Mo, et al. Expires 3 April 2027 [Page 27]

Internet-Draft Agent Selection Mapping September 2026 P3. Candidate enumeration. Input: the service identifier, the instance set, and the endpoints that the step is expected to contact. Action: form the candidate set of instances and paths, including the endpoints discovered for this step. Output: the candidate set. On failure: an empty candidate set triggers Section 8.4. P4. Constraint filtering. Input: the descriptor and the resource view. Action: remove every candidate that violates a constraint, treating an unknown value as not satisfying the constraint. Output: the eligible set, with the reason for each exclusion. On failure: an empty eligible set triggers Section 8.4. P5. Per-dimension valuation. Input: the eligible candidates and the descriptor. Action: value the forwarding, computing, and storage dimensions, element by element for storage, and at the stated percentile where the objective is tail-scoped. Output: one valuation per dimension per candidate. On failure: a dimension that cannot be valued at all is treated as unknown and is reported in the decision record. P6. Combination and tail awareness. Input: the per-dimension valuations and the preferences. Action: combine by the stated rule, weighted or lexicographic, and apply the tail objective as a bound on the outcome rather than as one more term in a sum. Output: a suitability per candidate. On failure: if no combination rule applies, the descriptor is malformed and the step is refused rather than served on a default. P7. Change cost and stability. Input: the suitability values, the current commitment, and the selection context. Action: compare the gain of moving against the cost of moving, that is, transfer or reconstruction plus the disturbance of a step in flight, and apply the stability requirement in force. Output: keep, hold-and-watch, or move. On failure: an unquantified cost of change leads to the conservative choice, which is not to move the session. P8. Commitment. Input: the outcome of P7. Action: bind the step, or the range of steps, to the selected instance with an explicit lifetime, and hand the steering instruction to the forwarders (MSG-10). Output: the commitment and the steering instruction. On failure: a commitment that cannot be established is a failure of the step, and it is reported as one. P9. Handoff orchestration. Input: a move decision and the difference between the working sets of the two instances. Action: reserve at the target (MSG-11), transfer or reconstruct the elements that the target lacks (MSG-13), and commit only when the working set is usable; otherwise abort, remain at the source, and record why. Output: a handoff record, committed or aborted. On failure: the session stays where it is, and the migration is retried only at a later decision point. Mo, et al. Expires 3 April 2027 [Page 28]

Internet-Draft Agent Selection Mapping September 2026 P10. Execution monitoring and tail detection. Input: the progress of the running step against its expectation. Action: detect a step whose cost exceeds that expectation by the stated factor, and apply the degradation class of the step within what the constraints allow (MSG-14). Output: a degradation decision, a mid-step migration request, or a decision to fail the step. On failure: the step is failed with a report that names the constraint that could not be met, rather than continuing to overrun the budget. P11. Accounting, feedback, and close. Input: the consumption of the completed step. Action: attribute the consumption to session, turn, and step, update the remaining budget and the reserve, append the decision record, and at the end of the session release the state that is no longer needed (MSG-15, MSG-16). Output: an updated selection context. On failure: accounting that cannot be attributed degrades only the precision of later budget decisions, and it is reported.

8.2. Decision Points

A decision is required at the points below. Between them the mapping is not required to act, which is what keeps a stream of metric updates from becoming a stream of decisions. * DP1: session admission -- establish the initial commitment (P0). * DP2: start of a turn -- re-evaluate the commitment (P2 to P8). * DP3: step boundary -- re-evaluate for the next step (P2 to P8). * DP4: wake-up after a pause -- re-verify state presence (P4, P5). * DP5: metric freshness expiry -- re-validate the values used (P5). * DP6: sustained improvement -- consider a move (P7, P8). * DP7: constraint at risk -- re-filter and move if needed (P4, P9). * DP8: tail event detected -- degrade, migrate, or fail (P10). * DP9: budget threshold crossed -- pace or refuse the step (P11). * DP10: session close -- release state and account (P11).

8.3. Re-Evaluation Triggers and Stability

The triggers below cause a re-evaluation. Each is subject to the stability requirement, which exists because a session that is moved for a gain that does not survive the move is worse off than a session that is not moved. Mo, et al. Expires 3 April 2027 [Page 29]

Internet-Draft Agent Selection Mapping September 2026 * T1: step boundary -- re-evaluate unless the commitment covers the step. * T2: turn start -- re-evaluate with the accumulated state. * T3: wake-up after a pause -- re-verify state presence before reuse. * T4: freshness expiry -- re-validate the values the decision used. * T5: sustained improvement -- move only if the stability rule is met. * T6: constraint at risk -- re-filter and move; stability does not apply. * T7: instance degradation -- re-evaluate, and move if no step is in flight. * T8: tail event -- apply the degradation class, or fail with a report. * T9: budget threshold -- pace the remaining steps or refuse the step. * T10: session close -- release state and close the accounting. Four stability rules apply to every trigger except T6 and T8, where a constraint is already at risk and delay is the risk itself. * A minimum interval must elapse between two moves of the same session, unless a constraint is violated. * A move requires a minimum improvement of suitability, so that a value which fluctuates inside a band does not move a session. * The cost of change is compared with the gain over the remaining steps of the look-ahead horizon, not with the gain of one step. * A step in flight is not moved on a preference. Only a violated constraint, or a tail event that the degradation class cannot absorb, justifies moving work that has already started.

8.4. Fallback Ladder

When the procedure cannot produce a selection that satisfies every constraint, it degrades in the order below rather than failing the session or silently overrunning its budget. The order follows the CATS fallback decision framework [I-D.pang-cats-fallback-decision-framework]. 1. Re-band the descriptor. Relax a preference and keep every constraint. This is the only step that changes what the session asked for, and it changes the least significant part of it. Mo, et al. Expires 3 April 2027 [Page 30]

Internet-Draft Agent Selection Mapping September 2026 2. Degrade within the degradation class. Serve the step at a lower precision, at a lower capability tier, with a truncated working set, or with additional delay, where the descriptor permits that reduction. 3. Move the work. If the step has not started, select the best remaining eligible candidate; if it has started, move only when a constraint is violated and the degradation class cannot absorb the violation. 4. Fail the step with a report. The report names the constraint that could not be satisfied, the trigger that led to the attempt, and the candidates that were rejected, with the reason for each rejection. 5. Fail the session. Where the budget is exhausted or no candidate can serve the session at all, the failure is reported to the gateway so that the session can be deferred, approved by a human, or ended, rather than being served by retries that consume what is left of the budget. The following example illustrates phases P2 to P7, with illustrative values, for a retrieval-augmented assistant session in which an earlier turn has already produced a KV cache and a retrieved document set. Step: answer a follow-up question, 2 s objective, data must remain in region R, cost ceiling applies to the session. Candidate A (near edge, small model, holds session KV cache): forwarding good, computing: capability constraint NOT met. Candidate B (regional, capable, holds document set, no KV): forwarding ok, computing ok, storage: document set present, KV absent, recompute cost of KV = 0.6 s. Candidate C (regional, capable, holds KV cache and documents): forwarding ok, computing ok, storage: all present, cost ~0. A is eliminated by the capability constraint. C is preferred over B unless C fails a constraint or its forwarding cost exceeds the recompute saving of 0.6 s. A second illustration shows the same session later, when the state that has accumulated changes the answer and a tail event occurs. The same session at step 30, with illustrative values. Accumulated state: context state for about 60,000 tokens, a document set, and the plan state of the current task. Mo, et al. Expires 3 April 2027 [Page 31]

Internet-Draft Agent Selection Mapping September 2026 Candidate B (regional, capable, holds the document set, no context state): storage: the context state is absent, and the recompute cost for it is large because the step needs the earlier context. Candidate C (regional, capable, holds the full context state): storage: all elements present. C is kept. The gain from moving to B does not cover the cost of re-establishing the context state, and the stability requirement therefore keeps the session at C even where B reports more free capacity. During the step, a retrieval tool that normally answers in 100 ms is still running after 4 s. The step is a tail event (T8). The degradation class of the step permits a reduced retrieval set, so the procedure narrows the set and the step completes. The session is not migrated mid-step, because the context state is already resident and moving it would cost more than the delay it would remove.

9. Message Semantics

The phases of Section 8 require information at defined moments. This section names the exchanges through which that information is obtained, and states what each exchange has to carry, when it is expected, and what it must do when it cannot be answered. The names below are for exposition only. This document defines no encoding, no transport, and no message format; what it requires is that the semantic content of each exchange is available at the point in the procedure where the exchange is used. Carriage belongs to the documents that define the distribution mechanisms of the CATS framework.

9.1. Exchange Catalogue

MSG-1. Descriptor submission (agent gateway to C-TC). Trigger: session admission, and each step that the gateway can describe. Content: the descriptor elements of Section 7, with their scope. Response: MSG-9 at the end of the procedure. Rules: the descriptor is local to the domain, and it carries no session content beyond what a decision needs. MSG-2. Step boundary notification (agent gateway to C-TC and C-PS). Trigger: the start of a step, where the gateway knows it. Content: the attribution keys and the accumulated working set. Response: none. Rules: a step that is never announced is treated as part of the turn. Mo, et al. Expires 3 April 2027 [Page 32]

Internet-Draft Agent Selection Mapping September 2026 MSG-3. Descriptor refinement request (C-PS to agent gateway). Trigger: the network cannot derive an element that a decision needs. Content: the element, the step, and the scope for which it is needed. Response: MSG-1 for that step. Rules: a refinement that is not answered leaves the element unknown. MSG-4. Capability and load advertisement (C-SMA to C-PS). Trigger: on change, and at a period bounded by the freshness bound of what is advertised. Content: capability elements and computing quantities (Section 7), aggregated where the metric definition allows it. Response: none. Rules: an advertisement is not a commitment, and a stale advertisement must not satisfy a constraint. MSG-5. State availability query (C-PS to C-SMA). Trigger: the summary in an advertisement does not decide the question for the step, or the presence information has expired. Content: the reuse keys for which presence is needed, and the scope. Response: MSG-6. Rules: a query names keys, and never asks for the content of an element. MSG-6. State availability response (C-SMA to C-PS). Trigger: receipt of MSG-5. Content: presence per reuse key, the residency tier, the retrieval cost, the recompute cost, the age of the element, and the retention. Response: none. Rules: a key that the instance cannot answer for is reported as unknown and not as absent. MSG-7. Forwarding view advertisement (C-NMA to C-PS). Trigger: on change, and at a period bounded by the freshness bound of the forwarding quantities. Content: paths, per-prefix cost to the endpoints that a step may contact, determinism class, and the transfer window. Response: none. Rules: the endpoints are discovered while the session runs, so the view is extended rather than decided once. MSG-8. Evaluation request (C-TC to C-PS). Trigger: a decision point of Section 8.2. Content: the descriptor, the attribution, and the trigger. Response: MSG-9. Rules: one request per decision point, not one per metric update. MSG-9. Selection result (C-PS to C-TC). Trigger: completion of the procedure. Content: the selected candidate, the reason codes of the candidates that were excluded, and the suitability that decided the outcome. Response: none. Rules: a result that could not be produced is a failure of the step and is reported as one, with the constraint that could not be met. MSG-10. Steering instruction (C-PS to forwarders). Trigger: a commitment. Content: the candidate, the lifetime of the commitment, and the endpoints that the step is expected to contact. Response: none. Rules: the instruction is idempotent, and a repeated instruction must not create a second commitment. Mo, et al. Expires 3 April 2027 [Page 33]

Internet-Draft Agent Selection Mapping September 2026 MSG-11. State reservation request (C-PS to the target C-SMA). Trigger: a decision to move a session, taken before any transfer. Content: the reuse keys that the target must hold, the size, and the time by which they are needed. Response: MSG-12. Rules: reservation is idempotent; reserving twice must not reserve two copies, and a reservation that cannot be honoured must be refused rather than accepted and broken. MSG-12. Reservation response (C-SMA to C-PS). Trigger: receipt of MSG-11. Content: accepted or refused, the validity window of the reservation, and the reason for a refusal. Response: none. Rules: an expired reservation is not a reservation. MSG-13. State transfer or reconstruction notice (source or target to C-PS). Trigger: completion, failure, or abort of the movement of an element. Content: the reuse keys that are now usable, the keys that are not, the divergence that was accepted, and the time. Response: none. Rules: the commit of a handoff follows this notice; a transfer that fails leaves the source usable and is reported. MSG-14. Invalidation and degradation notice (C-SMA, C-NMA, or forwarders to C-PS). Trigger: a value that a decision relied on has become stale or false, an instance has degraded, or a running step has exceeded its expected cost. Content: the element or the constraint affected, the observed condition, and the time. Response: MSG-9 after re-evaluation, where the step is still running. Rules: this exchange is what turns a tail event into a decision; a notice that is not acted on must still be recorded. MSG-15. Accounting feedback (C-TC or forwarders to C-PS). Trigger: completion of a step, and at the end of a turn. Content: consumption per dimension, attributed to session, turn, and step. Response: none. Rules: consumption is monotone and is not revised downwards. MSG-16. Session close and state release (agent gateway to C-TC; C-PS to C-SMA). Trigger: the end of the session, or its abandonment. Content: the attribution keys and the elements that are no longer needed. Response: none. Rules: release is best effort, and an element that is not released is subject to the retention that the instance stated in MSG-6.

9.2. Exchange Sequences

A step-boundary decision uses the exchanges in this order. Mo, et al. Expires 3 April 2027 [Page 34]

Internet-Draft Agent Selection Mapping September 2026 time | | gateway: step boundary and descriptor (MSG-1/2) | C-TC: attribute the step to session and turn (MSG-8) | C-PS: detail state if the summary is stale (MSG-5) | C-SMA: presence, tier, cost, age, retention (MSG-6) | C-NMA: paths and per-prefix cost (MSG-7) | C-PS: filter, value, combine, apply stability (8.1-8.3) | C-PS: steering instruction to the forwarders (MSG-10) | C-PS: decision record and accounting (MSG-9/15) v next decision point A handoff uses the exchanges in this order, and it can be abandoned at any point before the commit. time | | C-PS: decide to move the session to candidate B (8.3) | C-PS: reserve at B, with a validity window (MSG-11) | C-SMA: accept or refuse (MSG-12) | source: send or rebuild what the target lacks (MSG-13) | C-PS: commit at B only when the set is usable (MSG-10) | on failure: abort, stay at the source, record why | C-PS: release what is no longer needed (MSG-16) |

9.3. Rules Common to the Exchanges

* Every exchange states the scope of what it carries, and a consumer that receives an element without a scope treats it as scoped to the step. * Reservation and commit are idempotent. A repeated request must not produce a second reservation, a second commitment, or a second transfer. * The order is reservation, transfer, commit. A commit that follows a failed or aborted transfer must be refused, and an abort must leave the source usable. * Advertisements are aggregated, and detail is obtained by query. A deployment must be able to bound the rate of both, because the decision points of one session are not the only demand on the components. Mo, et al. Expires 3 April 2027 [Page 35]

Internet-Draft Agent Selection Mapping September 2026 * A component that cannot answer reports unknown. Omission is read as unknown by the consumer, and never as a value. * An exchange that crosses an administrative boundary carries the least information that the decision needs, and no session content. The elements of Section 7 are chosen so that this is possible.

10. Aggregation, Normalization, and Freshness

The mapping consumes metrics that are defined elsewhere. Three requirements on that consumption follow from the procedure above. Aggregation is needed in the storage dimension in the same way as in the other dimensions. An instance cannot advertise every element of every working set that it holds, so state availability needs to be summarized, for example per key space or per state class, and the summary needs to be sufficient for the comparison in step 5. Normalization is not applied across dimensions when a dimension carries a constraint: eligibility is decided first (see M1 and M6). The definition of a single normalized metric [I-D.ietf-cats-metric-definition] remains useful for ranking eligible candidates and for reporting, but the eligibility decision in step 2 is made before normalization. Freshness is part of the input. A state availability value and a computing load value that were true several seconds ago may lead to a selection that is worse than a static default, so the freshness of the information used must be available to the procedure and must be able to trigger a re-evaluation [I-D.zhu-cats-metric-semantics].

11. Requirements

The following requirements are derived from the framework, the information elements, the procedure, and the exchanges above. They are stated using the conventions of BCP 14 [RFC2119] [RFC8174], are requirements on the design of a CATS system that supports agent services, and are not protocol requirements. They are grouped by the part of the mapping that they constrain, and they are numbered in one sequence so that a requirement can be cited without its group. M1. A CATS system SHOULD be able to evaluate a candidate instance against a set of constraints before it ranks candidates against each other. Mo, et al. Expires 3 April 2027 [Page 36]

Internet-Draft Agent Selection Mapping September 2026 M2. A CATS system SHOULD be able to express the capability that a step requires in terms that can be matched against the capability that an instance offers. M3. A CATS system SHOULD be able to evaluate the forwarding dimension of a step over the paths that the step will use, including paths to tools and data sources, rather than only over the path to the instance. M4. A CATS system SHOULD be able to represent, for a candidate instance, the presence and cost of the elements of the working set that the step needs. M5. A CATS system SHOULD be able to compare the cost of transferring a working set element with the cost of recomputing or re-retrieving it at the candidate. M6. A CATS system MUST be able to enforce locality, tenant, and jurisdiction constraints as constraints, and not as terms in a scalar score. M7. A CATS system SHOULD support more than one combination rule for the three dimensions, including at least a weighted rule and a lexicographic rule. M8. A CATS system SHOULD record, for each selection, the input that produced it, to the extent needed for audit and troubleshooting. M9. A CATS system SHOULD re-evaluate a selection when the metrics that supported it become stale, when the selected instance degrades, or when the step for which the selection was made completes. M10. A CATS system SHOULD account for the cost of changing the selected instance, so that a marginal improvement in suitability does not cause unnecessary migration of a session. M11. A CATS system SHOULD allow the operator to express the objective of a session, so that per-step decisions do not systematically disadvantage the completion of the session as a whole. M12. A CATS system SHOULD limit the state that it exposes to what the selection decision requires, and SHOULD pass it through aggregation where possible. M13. A CATS system SHOULD be able to attribute the traffic of a step to the session and turn to which it belongs, to the extent needed to apply a per-session objective or budget.

11.1. Framework and Descriptor Requirements

Mo, et al. Expires 3 April 2027 [Page 37]

Internet-Draft Agent Selection Mapping September 2026 M14. A CATS system SHOULD be able to derive a descriptor for a step, and to revise it between the steps of a session, rather than only for the session as a whole. M15. A CATS system SHOULD be able to express a budget with its dimension, such as time, tokens, cost, or energy, and to state whether that budget is a ceiling or a target. M16. A CATS system SHOULD be able to express the objective of a session and the horizon over which the objective is evaluated. M17. A CATS system SHOULD be able to express the objective of a step as a percentile over a stated window, together with the probability of exceeding that objective that the session tolerates. M18. A CATS system MUST keep constraints and preferences distinguishable throughout the procedure, including in the decision record and in the reports that the procedure produces. M19. A CATS system SHOULD be able to express a capability requirement as a tier or as a set membership test, and SHOULD preserve the reason when a candidate fails that test. M20. A CATS system MUST treat a value that is absent, stale, or of unknown provenance as unknown, and MUST NOT allow an unknown value to satisfy a constraint. Exchanges SHOULD indicate unknown explicitly rather than by omission. M21. A CATS system SHOULD be able to compare quantities in bands whose coarseness is stated in the selection context, and MUST NOT compare a banded value in a way that hides a violation of a constraint.

11.2. State and Storage Requirements

M22. A CATS system SHOULD be able to distinguish the classes of working set element defined in Section 6.3 when it evaluates the storage dimension. M23. A CATS system SHOULD be able to obtain the presence of a working set element per reuse key, down to a single element where the aggregated information does not decide the question. M24. A CATS system SHOULD be able to obtain, for a candidate and a reuse key, the residency tier, the retrieval cost, the recompute cost, and the age of the element. M25. A CATS system SHOULD re-verify the presence of the state that a session depends on after a pause that exceeds the freshness bound of the presence information. Mo, et al. Expires 3 April 2027 [Page 38]

Internet-Draft Agent Selection Mapping September 2026 M26. A CATS system SHOULD be able to execute a change of instance as a reservation, a transfer or reconstruction, and a commit, with an abort that leaves the source instance usable. M27. A CATS system SHOULD be able to state which part of a working set is sufficient for a step, and what divergence is acceptable when the whole set cannot be moved. M28. A CATS system MUST apply the constraint that governs each working set element to that element, and MUST NOT move an element whose constraint forbids the move.

11.3. Discrete Decision Requirements

M29. A CATS system SHOULD take its decisions at explicit decision points, SHOULD NOT require a decision for every update of every quantity, and SHOULD NOT split the work of one step across instances. M30. A CATS system SHOULD be able to obtain, or to detect, the boundary of a step, so that a selection can be revisited between steps. M31. A CATS system SHOULD apply a stability requirement before it selects a different instance for a session that is already placed. M32. A CATS system SHOULD compare the gain of changing the selected instance with the cost of the change, where that cost includes the transfer or reconstruction and the disturbance of a step in flight.

11.4. Tail and Long-Horizon Requirements

M33. A CATS system SHOULD evaluate an objective that is stated as a percentile at that percentile over the stated window, and SHOULD identify a quantity that is a mean as a mean. M34. A CATS system MUST NOT conclude that a constraint is satisfied from a mean value where the tail of the distribution can violate the constraint. M35. A CATS system SHOULD detect a running step whose cost exceeds its expected cost by the factor stated for that step. M36. A CATS system SHOULD apply the degradation class of a step in the order of the fallback ladder, and SHOULD fail a step with a report rather than allow it to overrun the session budget silently. M37. A CATS system SHOULD weigh state affinity against load balancing with the cost of re-establishing the state, considered over the look-ahead horizon rather than for one step. Mo, et al. Expires 3 April 2027 [Page 39]

Internet-Draft Agent Selection Mapping September 2026 M38. A CATS system SHOULD pace the consumption of a session budget across the steps of the session, and SHOULD hold back a reserve for the steps that remain. M39. A CATS system SHOULD retain the decision records of a session for the life of that session, to the extent needed by the later decisions of the same session and by the accounting of its budget.

11.5. Exchange Requirements

M40. A CATS system MUST be able to obtain each element of Section 7 at the point in the procedure where that element is required. M41. A CATS system MUST convey, with every quantity, the time at which the quantity was observed and the bound within which it may be used. M42. A CATS system SHOULD support both a query mode and a notification mode for the availability of state, so that detail is obtained on demand and change is signalled without polling. M43. A CATS system MUST make the exchanges that reserve, commit, and move state idempotent and safe to retry or to reorder. M44. A CATS system SHOULD aggregate advertisements per class or per key space, and SHOULD obtain detail by query rather than by broadcasting it.

11.6. Mapping Requirements

M45. A CATS system SHOULD be able to state, for each characteristic of a step, which of the transfer, storage, and compute dimensions it affects, and whether it acts as a constraint or as a quantity in that dimension. M46. A CATS system SHOULD charge the cost of a change of instance to the session budget that the change is intended to serve. M47. A CATS system SHOULD treat the dimensions asymmetrically at selection time: a transfer cost is paid once per change of instance, while storage and compute costs are paid per step. M48. A CATS system SHOULD be able to express, for each resource that it compares, the unit in which the resource is observed and the component that observes it. M49. A CATS system SHOULD be able to derive the cost of not holding the state that a step needs at a candidate as the cheaper of the transfer time, that is, size over available bandwidth plus delay, and the recomputation time, that is, tokens over rate. Mo, et al. Expires 3 April 2027 [Page 40]

Internet-Draft Agent Selection Mapping September 2026 M50. A CATS system SHOULD be able to reject a candidate on the basis of a resource comparison, such as insufficient free memory, a step cost that exceeds the step budget, or a breached tail objective, and SHOULD report which comparison rejected it.

12. Relationship to the CATS Framework

The mapping framework uses the existing components. The following table states the responsibility of each component with respect to the three layers and to the exchanges of Section 9. Component | Responsibility in the mapping -----------+------------------------------------------------------ C-TC | Classifies session traffic; attributes it to session, | turn, and step; carries the descriptor and the | accounting. C-SMA | Provides capability, load, and state presence for its | site; answers state queries; honours and reports | reservations; releases state. C-NMA | Provides the forwarding view, including the | per-prefix cost to the endpoints of a step. C-PS | Derives the descriptor, applies constraints, values | and combines the dimensions, applies the stability | requirement, commits, and records. Forwarders | Execute the steering instruction, including the | endpoints of a step, and report progress that | indicates a tail event. No new functional component is required. What the framework adds is the explicit distinction between constraints and quantities, the storage dimension of the resource view, and the definition of the Selection Point at which a decision is valid.

12.1. Traceability to the Agent Service Requirements

The requirements of [I-D.mo-cats-agent-service-characteristics] are addressed as follows. The last four rows map the four properties of Section 5, which are taken from the same source but are not single requirement items there. Mo, et al. Expires 3 April 2027 [Page 41]

Internet-Draft Agent Selection Mapping September 2026 Agent service | Addressed by requirement | ---------------+-------------------------------------------------- R1, R2 | M1, M6, M7, M18, M19, M20 R3 | M13; Selection Context (Layer 3) R4 | M9, M10, M31, M32 R5, R6 | M3; Forwarding dimension (Layer 2) R7 | M36; tolerance element of the descriptor R8, R9, R10 | M2, M4, M19; Computing dimension R11, R12 | M4, M5, M23, M24 R13 | M6, M28 R14, R19, R20 | M22, M23, M24, M26, M27, M43 R15 | M11, M15, M38 R16, R21, R22 | Section 7 of the characteristics; M8, M12, M41, | M44 R17 | M33, M34, M35 R18 | M42; mode element of the descriptor long-horizon | M16, M25, M37, M38, M39 (Section 5.1) | stateful | M22 to M28 (Section 5.2) | discrete | M14, M21, M29, M30 (Section 5.3) | heavy-tailed | M17, M33 to M36, M40 (Section 5.4) | resource | M45, M46, M47 mapping | (Section 6.6) | resource | M48, M49, M50 routing | (Section 6.7) |

13. Operational Considerations

Operators enabling this mapping should expect five consequences. * The decision logic is more complex than a score comparison, which costs configuration and troubleshooting effort. Start with one constraint set and one combination rule, and add rules only where experience justifies them. * State affinity concentrates load. If every session is steered to the instance that holds its state, a subset of instances absorbs the load, so the policy that overrides affinity must be expressible in the descriptor or in the selection context. Mo, et al. Expires 3 April 2027 [Page 42]

Internet-Draft Agent Selection Mapping September 2026 * Correctness depends on the honesty of the resource view. An instance that over-reports capability or state attracts the sessions whose state it can then observe. * The number of decisions grows with steps rather than with sessions, so the decision rate and the state query rate need a bound, and that bound belongs in the selection context. * The mapping depends on summaries. A percentile or an aggregate computed from too few samples propagates its noise into the decision, which is one of the uses of the decision record.

14. Security Considerations

The decision depends on data supplied by the instances, which raises the value of misrepresenting it. Claimed state availability should be verified in the same way as claimed computing capability, and an implementation should not let a single dimension supplied by an untrusted party dominate the combination. The descriptor and the resource view expose the constraints and the working set of a session, so least disclosure applies: a descriptor should not be exposed beyond the components that need it. Reservation and handoff need their own limits. A reservation consumes the storage of an instance without moving a session to it, and a transfer consumes bandwidth between sites, so both need the authorization of the selection that they implement and a per-session rate limit. Tail-aware degradation can be gamed by an instance that under-reports its high percentiles or that reports an optimistic expectation. Percentile and expected-cost estimates should therefore be attributed to a source in the resource view, so that a source which is repeatedly wrong can be discounted.

15. Privacy Considerations

The descriptor and the resource view can reveal user behaviour: the number of steps, the sensitivity expressed through locality constraints, the size of the working set, and the tools in use. Aggregation and short retention are recommended, and the mapping does not require the network to learn the content of the session. Mo, et al. Expires 3 April 2027 [Page 43]

Internet-Draft Agent Selection Mapping September 2026 The classes and the reuse keys of a working set reveal structure even where the content does not, so reuse keys should be scoped to their tenant, and a handle should not be linkable across tenants or between the sessions of different users.

16. IANA Considerations

This document has no IANA actions.

17. Normative References

[I-D.ietf-cats-framework] Li, C., Du, Z., Boucadair, M., Contreras, L. M., et al., "A Framework for Computing-Aware Traffic Steering (CATS)", Work in Progress, Internet-Draft, draft-ietf-cats-framework-24, September 2026. [I-D.ietf-cats-metric-definition] Yao, K., et al., "CATS Metrics Definition", Work in Progress, Internet-Draft, draft-ietf-cats-metric-definition-11, September 2026. [I-D.mo-cats-agent-service-characteristics] Mo, Y., Yang, D., Zhou, C., "AI Agent Service Characteristics and Their Implications for Computing-Aware Traffic Steering", Work in Progress, Internet-Draft, draft-mo-cats-agent-service-characteristics-00, September 2026.

18. Informative References

[CATS-CHARTER] IETF, "Computing-Aware Traffic Steering (CATS) Working Group Charter", <https://datatracker.ietf.org/doc/charter-ietf-cats/>. [I-D.ietf-cats-data-model] Yao, H., Lin, C., et al., "Data Model for Computing-Aware Traffic Steering (CATS)", Work in Progress, draft-ietf-cats-data-model-00, September 2026. [I-D.ietf-cats-oam-fw] Fu, H., Xiong, Q., Du, Z., et al., "Computing- Aware Traffic Steering (CATS) Operations, Administration, and Maintenance (OAM) Framework", Work in Progress, Internet-Draft, draft-ietf-cats-oam-fw-01, July 2026. [I-D.ietf-cats-usecases-requirements] Yao, K., et al., "Computing-Aware Traffic Steering (CATS) Problem Statement, Use Cases, and Requirements", Work in Progress, draft-ietf-cats-usecases-requirements-14, September 2026. [I-D.li-cats-kv-cache-distribution] Li, Z., et al., "KV Cache Distribution for Distributed LLM Inference: Use Case and Requirements", Work in Progress, draft-li-cats-kv-cache-distribution-00, July 2026. Mo, et al. Expires 3 April 2027 [Page 44]

Internet-Draft Agent Selection Mapping September 2026 [I-D.luan-cats-catpts] Li, Q., Luan, Z., et al., "A Timescale-Aware Framework for Compute-Aware Task Placement and Traffic Steering", Work in Progress, draft-luan-cats-catpts-01, August 2026. [I-D.pang-cats-fallback-decision-framework] Pang, R., Ed., Han, M., Ed., Huang, T., Ed., "CATS Fallback Decision Framework", Work in Progress, Internet-Draft, draft-pang-cats-fallback-decision-framework-00, July 2026. [I-D.zhu-cats-metric-semantics] Zhu, M., "Operational Semantics for CATS Metric Consumption", Work in Progress, draft-zhu-cats-metric-semantics-01, August 2026. [RFC2119] Bradner, S., "Key words for use in RFCs to Indicate Requirement Levels", BCP 14, RFC 2119, DOI 10.17487/RFC2119, March 1997, <https://www.rfc-editor.org/info/rfc2119>. [RFC8174] Leiba, B., "Ambiguity of Uppercase vs Lowercase in RFC 2119 Key Words", BCP 14, RFC 8174, DOI 10.17487/RFC8174, May 2017, <https://www.rfc-editor.org/info/rfc8174>. Acknowledgments The authors would like to thank the participants of the CATS working group for the discussions that shaped this document. Authors' Addresses Y. Mo Huazhong University of Science and Technology Email: moyj@hust.edu.cn D. Yang Huazhong University of Science and Technology Email: d202581903@hust.edu.cn C. Zhou Huazhong University of Science and Technology Email: m202474228@hust.edu.cn Mo, et al. Expires 3 April 2027 [Page 45]