AgentOps Observability for Failure Detection and Attribution
draft-pei-opsawg-agentops-observability-00
This document is an Internet-Draft (I-D).
Anyone may submit an I-D to the IETF.
This I-D is not endorsed by the IETF and has no formal standing in the
IETF standards process.
| Document | Type | Active Internet-Draft (individual) | |
|---|---|---|---|
| Authors | Changhua Pei , Zexin Wang , Jingjing Li | ||
| Last updated | 2026-07-21 | ||
| RFC stream | (None) | ||
| Intended RFC status | (None) | ||
| Formats | |||
| Stream | Stream state | (No stream defined) | |
| Consensus boilerplate | Unknown | ||
| RFC Editor Note | (None) | ||
| IESG | IESG state | I-D Exists | |
| Telechat date | (None) | ||
| Responsible AD | (None) | ||
| Send notices to | (None) |
draft-pei-opsawg-agentops-observability-00
Operations and Management Area Working Group C. Pei
Internet-Draft CNIC, Chinese Academy of Sciences
Intended status: Standards Track 21 July 2026
Expires: 22 January 2027
AgentOps Observability for Failure Detection and Attribution
draft-pei-opsawg-agentops-observability-00
Abstract
Agentic systems execute tasks through long-running sequences of model
inference, planning, delegation, tool use, state updates,
verification, and recovery. Conventional metrics, logs, and traces
can identify a failed request but often cannot determine whether a
failure is emerging, which actor introduced it, or which earlier
event was its root cause. This document specifies an AgentOps
observability event model and processing requirements for
interoperable early failure detection and post-failure root-cause
attribution. It defines common event, anomaly, assertion, and
diagnosis records; distinguishes causal origins from propagated
symptoms; and provides the evidence model required by separate
benchmarks of detection lead time, responsible-actor attribution, and
root-cause localization. The model is transport-neutral and can be
carried by existing telemetry systems.
Status of This Memo
This Internet-Draft is submitted in full conformance with the
provisions of BCP 78 and BCP 79.
Internet-Drafts are working documents of the Internet Engineering
Task Force (IETF). Note that other groups may also distribute
working documents as Internet-Drafts. The list of current Internet-
Drafts is at https://datatracker.ietf.org/drafts/current/.
Internet-Drafts are draft documents valid for a maximum of six months
and may be updated, replaced, or obsoleted by other documents at any
time. It is inappropriate to use Internet-Drafts as reference
material or to cite them other than as "work in progress."
This Internet-Draft will expire on 22 January 2027.
Copyright Notice
Copyright (c) 2026 IETF Trust and the persons identified as the
document authors. All rights reserved.
Pei Expires 22 January 2027 [Page 1]
Internet-Draft AgentOps Failure Observability July 2026
This document is subject to BCP 78 and the IETF Trust's Legal
Provisions Relating to IETF Documents (https://trustee.ietf.org/
license-info) in effect on the date of publication of this document.
Please review these documents carefully, as they describe your rights
and restrictions with respect to this document. Code Components
extracted from this document must include Revised BSD License text as
described in Section 4.e of the Trust Legal Provisions and are
provided without warranty as described in the Revised BSD License.
Table of Contents
1. Introduction . . . . . . . . . . . . . . . . . . . . . . . . 3
1.1. Requirements Language . . . . . . . . . . . . . . . . . . 4
1.2. Scope . . . . . . . . . . . . . . . . . . . . . . . . . . 4
2. Terminology . . . . . . . . . . . . . . . . . . . . . . . . . 4
3. AgentOps Observability Architecture . . . . . . . . . . . . . 5
4. Event Model . . . . . . . . . . . . . . . . . . . . . . . . . 6
4.1. Event Envelope . . . . . . . . . . . . . . . . . . . . . 6
4.2. Actor and Phase . . . . . . . . . . . . . . . . . . . . . 7
4.3. Event Types . . . . . . . . . . . . . . . . . . . . . . . 7
4.4. Evidence and Content Handling . . . . . . . . . . . . . . 9
4.5. CDDL Definition . . . . . . . . . . . . . . . . . . . . . 9
5. Early Failure Detection . . . . . . . . . . . . . . . . . . . 11
5.1. Baselines and Partial-Trajectory Evaluation . . . . . . . 11
5.2. Anomaly Signal . . . . . . . . . . . . . . . . . . . . . 11
5.3. Assertions and Trust Levels . . . . . . . . . . . . . . . 12
6. Root-Cause Attribution . . . . . . . . . . . . . . . . . . . 13
6.1. Attribution Principles . . . . . . . . . . . . . . . . . 13
6.2. Diagnosis Record . . . . . . . . . . . . . . . . . . . . 14
6.3. Failure Categories . . . . . . . . . . . . . . . . . . . 15
7. Benchmarking Considerations . . . . . . . . . . . . . . . . . 15
8. Operational Processing Requirements . . . . . . . . . . . . . 16
8.1. Collection and Sampling . . . . . . . . . . . . . . . . . 16
8.2. Clock and Ordering . . . . . . . . . . . . . . . . . . . 16
8.3. Schema Evolution . . . . . . . . . . . . . . . . . . . . 16
9. Privacy Considerations . . . . . . . . . . . . . . . . . . . 16
10. Security Considerations . . . . . . . . . . . . . . . . . . . 17
11. IANA Considerations . . . . . . . . . . . . . . . . . . . . . 17
12. Implementation Status . . . . . . . . . . . . . . . . . . . . 17
13. References . . . . . . . . . . . . . . . . . . . . . . . . . 18
13.1. Normative References . . . . . . . . . . . . . . . . . . 18
13.2. Informative References . . . . . . . . . . . . . . . . . 18
Appendix A. Example Records . . . . . . . . . . . . . . . . . . 19
A.1. Early Anomaly Example . . . . . . . . . . . . . . . . . . 19
A.2. Attribution Example . . . . . . . . . . . . . . . . . . . 20
Acknowledgements . . . . . . . . . . . . . . . . . . . . . . . . 21
Author's Address . . . . . . . . . . . . . . . . . . . . . . . . 21
Pei Expires 22 January 2027 [Page 2]
Internet-Draft AgentOps Failure Observability July 2026
1. Introduction
Large language model (LLM)-based agents increasingly act as software
components that plan, call tools, coordinate with other agents,
update memory, inspect environmental feedback, and revise
intermediate results. An execution is therefore not a single
request-response pair. It is a distributed, stateful trajectory
involving heterogeneous actors and a mixture of deterministic and
probabilistic operations.
Traditional operations telemetry remains necessary, but it is not
sufficient for this setting. A late tool error or failed
verification may be only the visible symptom of an earlier flawed
plan, incorrect handoff, state loss, hallucinated assumption, or
ignored constraint. Furthermore, telemetry generated by an agent can
itself be incomplete or incorrect. Operational systems need evidence
that supports two different functions:
* online detection of a deviation early enough to permit mitigation;
and
* post-failure attribution of the responsible actor and earliest
decisive root-cause event.
These functions MUST NOT be conflated. An anomalous event is not
necessarily causal, and a causal event need not have the largest
anomaly score. Long-horizon evidence also creates a practical
context problem for LLM-based diagnostic systems: a complete
trajectory may contain hundreds of steps, while the relevant evidence
is sparse and separated from the final failure by many intervening
events.
This document defines a common event model that preserves semantic,
topological, temporal, and responsibility-transfer evidence. It also
defines interoperable records for anomaly signals, operational
assertions, and final diagnoses. Detection and attribution
algorithms are outside the scope of this document.
The design is informed by work on agent-and-step failure attribution
[WHO-WHEN], anomaly-first trajectory filtering [DUOTRACE], and long-
horizon root-cause benchmarks [LONGRCA]. It is intended to
complement, rather than replace, distributed tracing and telemetry
semantic conventions such as [W3C-TRACE-CONTEXT] and [OTEL-SEMCONV].
Pei Expires 22 January 2027 [Page 3]
Internet-Draft AgentOps Failure Observability July 2026
1.1. Requirements Language
The key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT",
"SHOULD", "SHOULD NOT", "RECOMMENDED", "NOT RECOMMENDED", "MAY", and
"OPTIONAL" in this document are to be interpreted as described in
BCP 14 [RFC2119] [RFC8174] when, and only when, they appear in all
capitals, as shown here.
1.2. Scope
This document specifies:
* a logical event envelope for agentic executions;
* minimum evidence needed for early failure detection;
* a representation for anomaly and policy-assertion signals;
* a representation for actor and root-cause attribution;
* rules for causal-chain roles and delegation-aware responsibility;
and
* evidence fields needed by AgentOps diagnosis benchmarks.
This document does not specify a telemetry transport, storage engine,
sampling algorithm, anomaly-detection model, or automated remediation
policy. It does not require disclosure of model weights, hidden
chain-of-thought, or plaintext user content.
2. Terminology
Agentic system: A software system in which one or more model-driven
actors select or execute actions over multiple steps, possibly
using tools, memory, delegation, and environmental feedback.
Actor: An entity that participates in an execution. An actor can be
an agent, orchestrator, model, tool, verifier, human, or
environment component.
Trajectory: The ordered event graph associated with one agentic task
or session. A trajectory can contain branches, joins, retries,
and nested calls.
Root-cause event: The earliest decisive event that introduced a
defect without which the observed failure would not have occurred
under the recorded execution.
Pei Expires 22 January 2027 [Page 4]
Internet-Draft AgentOps Failure Observability July 2026
Responsible actor: The actor accountable for the root-cause event.
This can be an agent, a non-agent component, or the external input
or environment.
Propagation event: An event that carries forward an already
introduced defect without independently introducing the decisive
defect.
Exposure event: An event that makes a latent defect observable, such
as a failed test or verifier rejection.
Terminal event: The event at which the system declares or
irrecoverably reaches failure.
Early failure detection: Production of an anomaly or assertion
signal before the terminal event, with sufficient evidence to
support an operational response.
Trust level agreement (TLA): An application-defined set of runtime
confidence or safety levels and assertions that establish
operational boundaries for an agentic system.
3. AgentOps Observability Architecture
An implementation typically contains the following logical roles:
1. Event producers instrument agents, orchestrators, model gateways,
tools, memory systems, verifiers, and environments.
2. A collector correlates and stores event records without changing
their causal identifiers.
3. A detector compares the partial trajectory with one or more
baselines and emits anomaly or assertion records.
4. An attribution engine retrieves evidence, reconstructs
propagation, and emits a diagnosis record.
5. An operator or automated controller decides whether to continue,
checkpoint, retry, reroute, constrain, or terminate the task.
The same component can perform more than one role. The architecture
does not require a central collector, but all records used together
MUST share a stable session identifier and SHOULD share distributed
trace identifiers.
Pei Expires 22 January 2027 [Page 5]
Internet-Draft AgentOps Failure Observability July 2026
Event producers SHOULD reuse W3C Trace Context identifiers when an
event is part of a distributed trace. An AgentOps event is more
specific than a generic span: it also records actor identity, agentic
phase, action kind, delegation, state references, and evidence needed
to distinguish causal origin from downstream manifestation.
4. Event Model
4.1. Event Envelope
Each event record MUST contain:
* schema_version, identifying the version of this event model;
* event_id, unique within the telemetry domain;
* session_id, stable across the complete task trajectory;
* timestamp, formatted according to [RFC3339];
* sequence, a monotonically increasing value within the producing
actor;
* actor, identifying the actor and its operational role;
* phase, identifying the agentic lifecycle phase;
* event_type, identifying the operation represented by the event;
and
* status, describing the event outcome known at emission time.
Each event SHOULD contain trace_id, span_id, and parent_span_id when
distributed trace context exists. It SHOULD contain parent_event_ids
when the execution topology cannot be represented by a single parent
span.
The sequence field orders events from one producer. Consumers MUST
NOT assume that it establishes a total order across producers.
Consumers SHOULD use causal links, timestamps, and producer-local
sequence numbers together.
Pei Expires 22 January 2027 [Page 6]
Internet-Draft AgentOps Failure Observability July 2026
4.2. Actor and Phase
The actor.kind value MUST be one of human, agent, orchestrator,
model, tool, verifier, environment, or other. The actor.id value
MUST be stable within the session. The actor.role value SHOULD
describe the functional role rather than a vendor-specific class
name.
The phase value MUST be one of:
* pre-execution: request interpretation, constraint extraction,
planning, actor selection, and initial state construction;
* execution: inference, delegation, tool use, memory access, state
update, and intermediate verification; or
* post-execution: final verification, result publication, diagnosis,
remediation, and audit.
4.3. Event Types
This document defines the following initial event types:
Pei Expires 22 January 2027 [Page 7]
Internet-Draft AgentOps Failure Observability July 2026
+=================================+================================+
| Event type | Meaning |
+=================================+================================+
| session.start, session.end | Task or session boundary |
+---------------------------------+--------------------------------+
| task.received | Input and constraints accepted |
+---------------------------------+--------------------------------+
| plan.created, plan.revised | Plan or decomposition changed |
+---------------------------------+--------------------------------+
| agent.delegated, agent.returned | Responsibility transferred or |
| | returned |
+---------------------------------+--------------------------------+
| model.request, model.response | Model interaction boundary |
+---------------------------------+--------------------------------+
| tool.call, tool.result | Tool interaction boundary |
+---------------------------------+--------------------------------+
| memory.read, memory.write | Persistent or session state |
| | access |
+---------------------------------+--------------------------------+
| checkpoint.created, | Recoverable state boundary |
| checkpoint.restored | |
+---------------------------------+--------------------------------+
| assertion.evaluated | TLA or policy assertion |
| | evaluated |
+---------------------------------+--------------------------------+
| verification.result | Intermediate or final |
| | verification |
+---------------------------------+--------------------------------+
| recovery.attempt | Retry, rollback, reroute, or |
| | repair action |
+---------------------------------+--------------------------------+
| human.intervention | Human decision or correction |
+---------------------------------+--------------------------------+
| anomaly.signal | Detector output |
+---------------------------------+--------------------------------+
| diagnosis.result | Attribution engine output |
+---------------------------------+--------------------------------+
Table 1: Initial AgentOps Event Types
Unknown event types MUST be preserved by collectors and MUST NOT
cause the containing trajectory to be rejected. Their semantics are
extension data.
Pei Expires 22 January 2027 [Page 8]
Internet-Draft AgentOps Failure Observability July 2026
4.4. Evidence and Content Handling
Events SHOULD record structured attributes required for diagnosis,
including action name, input and output references, tool result
status, model and tool versions, token usage, duration, delegation
target, verifier result, and checkpoint reference when applicable.
Large or sensitive content SHOULD be represented by an evidence_ref
containing a content hash, media type, access-controlled location,
and sensitivity label. A collector MUST NOT dereference protected
evidence unless authorized for the relevant purpose.
Implementations MUST NOT require hidden model chain-of-thought. An
implementation MAY record a model-generated decision summary,
reflection, or explanation when the model and deployment policy
permit it. Such text MUST be labeled as self-reported evidence and
MUST NOT be treated as an authoritative description of internal model
computation.
4.5. CDDL Definition
The following Concise Data Definition Language (CDDL) [RFC8610]
fragment defines the JSON-compatible logical record. JSON
serialization MUST follow [RFC8259].
agentops-record = event / anomaly / assertion / diagnosis
event = {
schema_version: tstr,
event_id: tstr,
session_id: tstr,
timestamp: tstr,
sequence: uint,
? trace_id: tstr,
? span_id: tstr,
? parent_span_id: tstr,
? parent_event_ids: [* tstr],
actor: actor,
phase: "pre-execution" / "execution" / "post-execution",
event_type: tstr,
status: "started" / "ok" / "error" / "cancelled" /
"unknown",
? action: action,
? delegation: delegation,
? model: component,
? tool: component,
? checkpoint_id: tstr,
? evidence: [* evidence-ref],
Pei Expires 22 January 2027 [Page 9]
Internet-Draft AgentOps Failure Observability July 2026
? attributes: { * tstr => any }
}
actor = {
id: tstr,
kind: "human" / "agent" / "orchestrator" / "model" /
"tool" / "verifier" / "environment" / "other",
? role: tstr,
? instance: tstr,
? version: tstr
}
action = {
kind: tstr,
? name: tstr,
? input_ref: evidence-ref,
? output_ref: evidence-ref,
? duration_ms: uint,
? token_input: uint,
? token_output: uint
}
delegation = {
from_actor: tstr,
to_actor: tstr,
instruction_ref: evidence-ref,
? constraints_ref: evidence-ref
}
component = {
name: tstr,
? provider: tstr,
? version: tstr,
? call_id: tstr
}
evidence-ref = {
id: tstr,
? media_type: tstr,
? hash: tstr,
? location: tstr,
? sensitivity: "public" / "internal" / "confidential" /
"restricted",
? provenance: "observed" / "self-reported" / "derived"
}
Figure 1: AgentOps Core Event Model
Pei Expires 22 January 2027 [Page 10]
Internet-Draft AgentOps Failure Observability July 2026
5. Early Failure Detection
5.1. Baselines and Partial-Trajectory Evaluation
A detector operates on a partial trajectory and compares it with an
explicit baseline. The baseline can represent successful executions,
policy-conforming executions, a known model or tool version, or an
application-defined TLA level.
A detector SHOULD consider both:
* semantic evidence, such as actor role, operation, constraint, tool
result, and state transition; and
* structural evidence, such as delegation topology, call tree,
branch, retry, checkpoint, and dependency relationships.
This dual view is important because an individually plausible action
can be anomalous in its topological or role context, while a
structurally unusual action can be semantically correct. The
specification does not mandate a particular detector.
Reconstruction, prediction, assertion, spectrum, statistical, and
hybrid detectors can all produce conforming anomaly records.
5.2. Anomaly Signal
An anomaly record MUST identify the detector, baseline, detection
time, score, and candidate event set. It MUST state whether a higher
or lower score indicates greater deviation. It SHOULD include the
threshold, feature groups, confidence, and the earliest candidate
event.
Pei Expires 22 January 2027 [Page 11]
Internet-Draft AgentOps Failure Observability July 2026
anomaly = {
schema_version: tstr,
anomaly_id: tstr,
session_id: tstr,
detected_at: tstr,
detector: component,
baseline_id: tstr,
score: number,
score_direction: "higher-is-more-anomalous" /
"lower-is-more-anomalous",
candidate_event_ids: [+ tstr],
? earliest_candidate_event_id: tstr,
? threshold: number,
? confidence: float,
? feature_groups: [* ("semantic" / "topology" / "resource" /
"state" / "assertion" / tstr)],
? explanation_ref: evidence-ref,
? recommended_action: "observe" / "checkpoint" / "retry" /
"reroute" / "constrain" / "terminate"
}
number = int / float
Figure 2: Anomaly Signal Record
The candidate_event_ids array is a ranked set. A detector SHOULD
emit a compact, evidence-preserving candidate set rather than copying
the complete trajectory into the anomaly record. The original events
MUST remain retrievable for audit and attribution.
5.3. Assertions and Trust Levels
Application assertions provide a deterministic complement to
statistical detection. An assertion record MUST identify the
applicable TLA level, assertion identifier, evaluation result, and
events used as evidence.
Pei Expires 22 January 2027 [Page 12]
Internet-Draft AgentOps Failure Observability July 2026
assertion = {
schema_version: tstr,
assertion_event_id: tstr,
session_id: tstr,
timestamp: tstr,
tla_level: tstr,
assertion_id: tstr,
result: "pass" / "approaching" / "violation" / "unknown",
evidence_event_ids: [* tstr],
? measured_value: any,
? boundary_value: any,
? remediation: tstr
}
Figure 3: Trust Level Assertion Record
The approaching result supports risk prediction before a hard
boundary is crossed. Trust levels and boundaries are application-
specific; this document standardizes their observability, not their
values.
6. Root-Cause Attribution
6.1. Attribution Principles
An attribution engine MUST distinguish at least four causal-chain
roles: root-cause, propagation, exposure, and terminal. It MUST NOT
select an exposure or terminal event solely because it is closest to
the observed failure.
Attribution SHOULD apply the following responsibility rules:
* If an executor faithfully implements a defective delegated
instruction, the delegating actor and instruction event are the
primary root-cause candidates. The executor event is propagation
evidence.
* If the executor departs from a valid instruction or introduces a
new defect, responsibility remains with the executor.
* A verifier that exposes a defect is not the root cause. A
verifier that misses a detectable defect can be a contributing
factor.
* If an external requirement is self-contradictory or an environment
failure is unrecoverable, a non-agent source can be the
responsible actor.
Pei Expires 22 January 2027 [Page 13]
Internet-Draft AgentOps Failure Observability July 2026
* Among causally sufficient candidates, the engine SHOULD select the
earliest decisive event for which a feasible recovery existed
before the failure propagated.
These rules allow responsibility to move across planner-executor-
verifier handoffs without automatically blaming either the first or
last actor.
6.2. Diagnosis Record
A diagnosis record MUST identify the responsible actor, root-cause
event, terminal event, supporting evidence, and method. It SHOULD
include the failure chain, category, confidence, and validation
status.
diagnosis = {
schema_version: tstr,
diagnosis_id: tstr,
session_id: tstr,
created_at: tstr,
method: component,
responsible_actor_id: tstr,
root_cause_event_id: tstr,
terminal_event_id: tstr,
evidence_event_ids: [+ tstr],
failure_chain: [+ chain-link],
? category: tstr,
? rationale_ref: evidence-ref,
? confidence: float,
? validation: "unvalidated" / "machine-validated" /
"human-adjudicated",
? contributing_actor_ids: [* tstr]
}
chain-link = {
event_id: tstr,
role: "root-cause" / "propagation" / "exposure" / "terminal",
actor_id: tstr
}
Figure 4: Root-Cause Diagnosis Record
The evidence_event_ids values MUST resolve to events in the original
trajectory. Derived summaries can be included through rationale_ref,
but they MUST NOT replace the source evidence.
Pei Expires 22 January 2027 [Page 14]
Internet-Draft AgentOps Failure Observability July 2026
6.3. Failure Categories
Implementations MAY use the following initial analysis taxonomy. The
category is auxiliary; responsible-actor and root-cause-event labels
remain the primary interoperable targets.
+========+================+======================================+
| Prefix | Class | Example leaf categories |
+========+================+======================================+
| A | Input | misunderstood intent, ignored |
| | understanding | constraint, misread observation |
+--------+----------------+--------------------------------------+
| B | Reasoning | fabrication, flawed reasoning |
+--------+----------------+--------------------------------------+
| C | Planning and | flawed plan, wrong handoff, cross- |
| | collaboration | actor state loss |
+--------+----------------+--------------------------------------+
| D | Execution and | wrong tool or arguments, unrecovered |
| | tools | tool failure, repetition |
+--------+----------------+--------------------------------------+
| E | Verification | verification gap, premature |
| | and completion | completion |
+--------+----------------+--------------------------------------+
| X | Non-agent | contradictory input, unrecoverable |
| | | environment |
+--------+----------------+--------------------------------------+
Table 2: Initial Failure Category Classes
Future specifications can define stable category identifiers.
Collectors MUST preserve unknown categories.
7. Benchmarking Considerations
This document standardizes observability evidence, not a task suite
or scoring methodology. AgentOps early-detection and root-cause-
diagnosis benchmark requirements are specified separately in
[AGENTOPS-BENCH].
In particular, an AgentOps diagnosis benchmark is distinct from a
benchmark of an agent that performs network configuration or other
network operations. The latter evaluates the agent's task capability
and network outcome; the former evaluates whether an operational
method can detect and diagnose a failure in an agentic system. A
network-management-agent trajectory can be encoded using this event
model and used as a diagnosis workload, but its network-task score is
not an AgentOps diagnosis score.
Pei Expires 22 January 2027 [Page 15]
Internet-Draft AgentOps Failure Observability July 2026
8. Operational Processing Requirements
8.1. Collection and Sampling
Sampling decisions can destroy the evidence required for diagnosis.
Implementations SHOULD retain task boundaries, delegation edges,
assertion violations, anomaly candidates, verifier outcomes, recovery
attempts, and checkpoint references whenever any event in the session
is retained.
When full payloads cannot be retained, producers SHOULD emit stable
hashes and access-controlled evidence references. Head sampling
alone is generally insufficient for failed trajectories; collectors
SHOULD support a deferred or tail-based retention decision triggered
by anomalies, assertion violations, or terminal failures.
8.2. Clock and Ordering
All timestamps MUST follow [RFC3339] and SHOULD be expressed in UTC.
Implementations SHOULD record clock-quality information when cross-
host ordering affects diagnosis. Causal parent links take precedence
over wall clock order when the two conflict.
8.3. Schema Evolution
Major incompatible revisions MUST change schema_version. Consumers
MUST ignore unknown optional fields and preserve them when forwarding
records. Required fields defined by a future major version MUST NOT
be assumed by a consumer that only supports this version.
9. Privacy Considerations
Agent trajectories can contain user prompts, personal data,
credentials, proprietary source code, retrieved documents, model
outputs, and behavioral profiles. Centralizing this material creates
privacy risks described more generally in [RFC6973].
Producers MUST support field-level redaction or evidence references
in place of plaintext content. Access to evidence SHOULD be purpose-
limited, logged, and time-bounded. Benchmark publishers MUST remove
credentials and direct identifiers and SHOULD evaluate whether
trajectory combinations can re-identify users or organizations.
Actor identifiers SHOULD be pseudonymous outside the operational
trust domain. Content hashes can still reveal low-entropy values
through guessing and therefore MUST NOT be treated as anonymization
by themselves.
Pei Expires 22 January 2027 [Page 16]
Internet-Draft AgentOps Failure Observability July 2026
Raw prompts, outputs, memory values, and self-reported reasoning
summaries SHOULD NOT cross an administrative boundary unless needed
for a declared diagnostic purpose and protected by an applicable
policy.
10. Security Considerations
AgentOps telemetry is security-sensitive. It can expose
instructions, system topology, tool capabilities, failure patterns,
and recovery controls. The general guidance in [RFC3552] applies.
Telemetry producers and collectors MUST provide confidentiality,
integrity, and origin authentication appropriate to the deployment.
Diagnosis systems MUST treat agent-generated explanations and
attributes as untrusted input. An agent or tool can attempt to hide
a failure, frame another actor, inject instructions into a diagnostic
model, or create excessive telemetry.
Collectors SHOULD separate observed facts from self-reported and
derived claims using the provenance field. Attribution engines
SHOULD validate cited event identifiers and SHOULD prefer directly
observed tool, verifier, and state-transition evidence over
unsupported natural-language claims.
Implementations MUST apply resource limits to record size, evidence
retrieval, candidate-set expansion, and recursive delegation graphs.
Automated remediation SHOULD require policy authorization
proportional to impact. In particular, an anomaly score alone MUST
NOT authorize destructive actions.
Replay and reordering can corrupt causal reconstruction. Deployments
SHOULD use tamper-evident storage, stable identifiers, producer-local
sequence numbers, and duplicate detection. Cross-domain trace
identifiers MUST NOT be accepted as proof of trust or identity.
11. IANA Considerations
This document has no IANA actions.
12. Implementation Status
This section is to be removed before publication as an RFC.
Pei Expires 22 January 2027 [Page 17]
Internet-Draft AgentOps Failure Observability July 2026
The event model is based on telemetry fields used in ongoing AgentOps
experiments for anomaly-first failure attribution and long-horizon
root-cause localization. No interoperability implementation is
claimed for version -00. Future revisions are expected to report
mappings to existing OpenTelemetry representations and validation of
the CDDL examples.
13. References
13.1. Normative References
[RFC2119] Bradner, S., "Key words for use in RFCs to Indicate
Requirement Levels", BCP 14, RFC 2119,
DOI 10.17487/RFC2119, March 1997,
<https://www.rfc-editor.org/rfc/rfc2119>.
[RFC3339] Klyne, G. and C. Newman, "Date and Time on the Internet:
Timestamps", RFC 3339, DOI 10.17487/RFC3339, July 2002,
<https://www.rfc-editor.org/rfc/rfc3339>.
[RFC8174] Leiba, B., "Ambiguity of Uppercase vs Lowercase in RFC
2119 Key Words", BCP 14, RFC 8174, DOI 10.17487/RFC8174,
May 2017, <https://www.rfc-editor.org/rfc/rfc8174>.
[RFC8259] Bray, T., Ed., "The JavaScript Object Notation (JSON) Data
Interchange Format", STD 90, RFC 8259,
DOI 10.17487/RFC8259, December 2017,
<https://www.rfc-editor.org/rfc/rfc8259>.
[RFC8610] Birkholz, H., Vigano, C., and C. Bormann, "Concise Data
Definition Language (CDDL): A Notational Convention to
Express Concise Binary Object Representation (CBOR) and
JSON Data Structures", RFC 8610, DOI 10.17487/RFC8610,
June 2019, <https://www.rfc-editor.org/rfc/rfc8610>.
13.2. Informative References
[RFC3552] Rescorla, E. and B. Korver, "Guidelines for Writing RFC
Text on Security Considerations", BCP 72, RFC 3552,
DOI 10.17487/RFC3552, July 2003,
<https://www.rfc-editor.org/rfc/rfc3552>.
[RFC6973] Cooper, A., Tschofenig, H., Aboba, B., Peterson, J.,
Morris, J., Hansen, M., and R. Smith, "Privacy
Considerations for Internet Protocols", RFC 6973,
DOI 10.17487/RFC6973, July 2013,
<https://www.rfc-editor.org/rfc/rfc6973>.
Pei Expires 22 January 2027 [Page 18]
Internet-Draft AgentOps Failure Observability July 2026
[W3C-TRACE-CONTEXT]
World Wide Web Consortium, "Trace Context", 23 November
2021, <https://www.w3.org/TR/trace-context/>.
[OTEL-SEMCONV]
OpenTelemetry Authors, "OpenTelemetry Semantic
Conventions", 2026,
<https://opentelemetry.io/docs/specs/semconv/>.
[WHO-WHEN] Zhang, S., Yin, M., Zhang, J., Liu, J., Han, Z., Zhang,
J., Li, B., Wang, C., Wang, H., Chen, Y., and Q. Wu,
"Which Agent Causes Task Failures and When? On Automated
Failure Attribution of LLM Multi-Agent Systems", 2025,
<https://arxiv.org/abs/2505.00212>.
[DUOTRACE] Authors withheld during double-blind review, "Detect
Before You Attribute - Cascade Failure Attribution for
Multi-Agent Systems", 2026.
[LONGRCA] Authors withheld during double-blind review, "LongRCA
Bench - Diagnosing Responsible Actors and Root Causes in
Long-Horizon Agent Failures", 2026.
[AGENTOPS-BENCH]
Pei, C., "Benchmarking Failure Detection and Root-Cause
Diagnosis for Agentic Systems", 2026.
Appendix A. Example Records
A.1. Early Anomaly Example
The following example reports a deviation at an early delegated
planning event. The payload is illustrative and not derived from
benchmark data.
Pei Expires 22 January 2027 [Page 19]
Internet-Draft AgentOps Failure Observability July 2026
{
"schema_version": "1.0",
"anomaly_id": "anom-204",
"session_id": "sess-7f2a",
"detected_at": "2026-07-15T09:31:18Z",
"detector": {
"name": "dual-view-baseline-detector",
"version": "0.3"
},
"baseline_id": "software-repair-success-v5",
"score": 4.72,
"score_direction": "higher-is-more-anomalous",
"threshold": 3.50,
"candidate_event_ids": ["evt-037", "evt-041", "evt-052"],
"earliest_candidate_event_id": "evt-037",
"confidence": 0.84,
"feature_groups": ["semantic", "topology", "state"],
"recommended_action": "checkpoint"
}
Figure 5: Illustrative Early Anomaly Record
A.2. Attribution Example
Pei Expires 22 January 2027 [Page 20]
Internet-Draft AgentOps Failure Observability July 2026
{
"schema_version": "1.0",
"diagnosis_id": "diag-991",
"session_id": "sess-7f2a",
"created_at": "2026-07-15T09:46:03Z",
"method": {
"name": "evidence-grounded-attribution",
"version": "0.2"
},
"responsible_actor_id": "planner-1",
"root_cause_event_id": "evt-037",
"terminal_event_id": "evt-214",
"evidence_event_ids": ["evt-037", "evt-038", "evt-143", "evt-214"],
"failure_chain": [
{"event_id": "evt-037", "role": "root-cause",
"actor_id": "planner-1"},
{"event_id": "evt-038", "role": "propagation",
"actor_id": "executor-2"},
{"event_id": "evt-143", "role": "exposure",
"actor_id": "verifier-1"},
{"event_id": "evt-214", "role": "terminal",
"actor_id": "orchestrator-1"}
],
"category": "C1",
"confidence": 0.78,
"validation": "human-adjudicated"
}
Figure 6: Illustrative Root-Cause Attribution Record
Acknowledgements
The author thanks the researchers and practitioners whose AgentOps,
multi-agent failure attribution, and long-horizon diagnosis work
informed this document.
Author's Address
Changhua Pei
CNIC, Chinese Academy of Sciences
Beijing
China
Pei Expires 22 January 2027 [Page 21]