<?xml version="1.0" encoding="UTF-8"?>
<reference anchor="I-D.contreras-bmwg-ai-agent-benchmarking" target="https://datatracker.ietf.org/doc/html/draft-contreras-bmwg-ai-agent-benchmarking-00">
   <front>
      <title>Benchmarking Methodology for AI Agents in Network Operations</title>
      <author initials="L. M." surname="Contreras" fullname="Luis M. Contreras">
         <organization>Telefonica</organization>
      </author>
      <date month="July" day="6" year="2026" />
      <abstract>
	 <t>   This document defines a benchmarking methodology for evaluating
   Artificial Intelligence (AI) agents performing network operations
   tasks such as configuration, troubleshooting, and optimization.  This
   document focuses on task-oriented performance metrics, including task
   completion success, execution efficiency, and robustness across
   multi-step workflows, proposing benchmarking practices towards agent-
   based, closed-loop network operation scenarios.

   The proposed methodology aims to provide a reproducible and vendor-
   independent framework to compare AI agent effectiveness in controlled
   network environments.

	 </t>
      </abstract>
   </front>
   <seriesInfo name="Internet-Draft" value="draft-contreras-bmwg-ai-agent-benchmarking-00" />
   
</reference>
