Skip to main content

Mouth-to-Ear Response Latency for Conversational Voice Systems: Metric Definition and Active Measurement Method
draft-nygate-ippm-mrl-00

Document Type Active Internet-Draft (individual)
Author Daniel Nygate
Last updated 2026-09-07
RFC stream (None)
Intended RFC status (None)
Formats
Stream Stream state (No stream defined)
Consensus boilerplate Unknown
RFC Editor Note (None)
IESG IESG state I-D Exists
Telechat date (None)
Responsible AD (None)
Send notices to (None)
draft-nygate-ippm-mrl-00
IP Performance Measurement                                     D. Nygate
Internet-Draft                                          7 September 2026
Intended status: Informational                                          
Expires: 11 March 2027

 Mouth-to-Ear Response Latency for Conversational Voice Systems: Metric
                Definition and Active Measurement Method
                        draft-nygate-ippm-mrl-00

Abstract

   This document defines mouth-to-ear response latency (MRL), a
   performance metric for conversational voice systems, together with an
   active method for measuring it at the RTP reference point of the
   calling endpoint.  MRL is the interval between the transmission of
   the final speech sample of a caller's utterance and the arrival of
   the first sample of the system's response audio.  Two variants are
   defined, one taken at packet arrival and one taken behind a de-jitter
   buffer of stated target depth.  The method is specified so that both
   timestamps are drawn from a single clock on a single host, so that
   the metric requires no synchronisation between the measuring endpoint
   and the system under test.  Requirements for stimulus material,
   capture content, quality control, calibration and reporting are
   given.

About This Document

   This note is to be removed before publishing as an RFC.

   Status information for this document may be found at
   https://datatracker.ietf.org/doc/draft-nygate-ippm-mrl/.

   Discussion of this document takes place on the IP Performance
   Measurement Working Group mailing list (mailto:ippm@ietf.org), which
   is archived at https://mailarchive.ietf.org/arch/browse/ippm/.
   Subscribe at https://www.ietf.org/mailman/listinfo/ippm/.

   Source for this draft and an issue tracker can be found at
   https://github.com/dnygate/draft-mrl.

Status of This Memo

   This Internet-Draft is submitted in full conformance with the
   provisions of BCP 78 and BCP 79.

Nygate                    Expires 11 March 2027                 [Page 1]
Internet-Draft        Mouth-to-Ear Response Latency       September 2026

   Internet-Drafts are working documents of the Internet Engineering
   Task Force (IETF).  Note that other groups may also distribute
   working documents as Internet-Drafts.  The list of current Internet-
   Drafts is at https://datatracker.ietf.org/drafts/current/.

   Internet-Drafts are draft documents valid for a maximum of six months
   and may be updated, replaced, or obsoleted by other documents at any
   time.  It is inappropriate to use Internet-Drafts as reference
   material or to cite them other than as "work in progress."

   This Internet-Draft will expire on 11 March 2027.

Copyright Notice

   Copyright (c) 2026 IETF Trust and the persons identified as the
   document authors.  All rights reserved.

   This document is subject to BCP 78 and the IETF Trust's Legal
   Provisions Relating to IETF Documents (https://trustee.ietf.org/
   license-info) in effect on the date of publication of this document.
   Please review these documents carefully, as they describe your rights
   and restrictions with respect to this document.  Code Components
   extracted from this document must include Revised BSD License text as
   described in Section 4.e of the Trust Legal Provisions and are
   provided without warranty as described in the Revised BSD License.

Table of Contents

   1.  Introduction  . . . . . . . . . . . . . . . . . . . . . . . .   3
     1.1.  Motivation and decomposition  . . . . . . . . . . . . . .   4
     1.2.  Existing practice . . . . . . . . . . . . . . . . . . . .   5
     1.3.  Scope . . . . . . . . . . . . . . . . . . . . . . . . . .   6
     1.4.  Relationship to existing work . . . . . . . . . . . . . .   6
   2.  Conventions and Definitions . . . . . . . . . . . . . . . . .   7
   3.  Metric Definition . . . . . . . . . . . . . . . . . . . . . .   8
     3.1.  Metric name . . . . . . . . . . . . . . . . . . . . . . .   8
     3.2.  Metric description  . . . . . . . . . . . . . . . . . . .   8
     3.3.  Interval start: t0  . . . . . . . . . . . . . . . . . . .   8
     3.4.  Interval end: t1  . . . . . . . . . . . . . . . . . . . .   9
       3.4.1.  Ingress MRL . . . . . . . . . . . . . . . . . . . . .  10
       3.4.2.  Playout MRL . . . . . . . . . . . . . . . . . . . . .  10
       3.4.3.  Both variants required  . . . . . . . . . . . . . . .  11
     3.5.  Response onset  . . . . . . . . . . . . . . . . . . . . .  11
     3.6.  Unprompted audio  . . . . . . . . . . . . . . . . . . . .  11
     3.7.  Filler audio and response continuity  . . . . . . . . . .  12
   4.  Method of Measurement . . . . . . . . . . . . . . . . . . . .  13
     4.1.  Reference point . . . . . . . . . . . . . . . . . . . . .  13
     4.2.  Observing elsewhere . . . . . . . . . . . . . . . . . . .  13

Nygate                    Expires 11 March 2027                 [Page 2]
Internet-Draft        Mouth-to-Ear Response Latency       September 2026

       4.2.1.  Characterising the additional terms . . . . . . . . .  16
     4.3.  Stimulus requirements . . . . . . . . . . . . . . . . . .  16
     4.4.  Transmission pacing . . . . . . . . . . . . . . . . . . .  16
     4.5.  Capture contents  . . . . . . . . . . . . . . . . . . . .  17
     4.6.  Clock requirements  . . . . . . . . . . . . . . . . . . .  17
     4.7.  Sources of error and calibration  . . . . . . . . . . . .  17
   5.  Units of Measurement  . . . . . . . . . . . . . . . . . . . .  19
   6.  Measurement Points and Measurement Domain . . . . . . . . . .  19
   7.  Measurement Timing  . . . . . . . . . . . . . . . . . . . . .  19
     7.1.  Turns within a session  . . . . . . . . . . . . . . . . .  20
   8.  Quality Control and Result Validity . . . . . . . . . . . . .  20
     8.1.  Blocking conditions . . . . . . . . . . . . . . . . . . .  20
     8.2.  Advisory conditions . . . . . . . . . . . . . . . . . . .  21
     8.3.  Reporting of discards . . . . . . . . . . . . . . . . . .  21
   9.  Reporting . . . . . . . . . . . . . . . . . . . . . . . . . .  21
     9.1.  Required fields . . . . . . . . . . . . . . . . . . . . .  21
     9.2.  Reporting format  . . . . . . . . . . . . . . . . . . . .  22
       9.2.1.  instrument  . . . . . . . . . . . . . . . . . . . . .  22
       9.2.2.  measurement . . . . . . . . . . . . . . . . . . . . .  23
       9.2.3.  subject . . . . . . . . . . . . . . . . . . . . . . .  23
       9.2.4.  results . . . . . . . . . . . . . . . . . . . . . . .  23
       9.2.5.  measurements and captures . . . . . . . . . . . . . .  24
       9.2.6.  Example . . . . . . . . . . . . . . . . . . . . . . .  25
   10. Use and Applications  . . . . . . . . . . . . . . . . . . . .  26
   11. Applicability and Limitations . . . . . . . . . . . . . . . .  27
   12. Security Considerations . . . . . . . . . . . . . . . . . . .  28
   13. IANA Considerations . . . . . . . . . . . . . . . . . . . . .  28
   14. Future Work . . . . . . . . . . . . . . . . . . . . . . . . .  29
   15. References  . . . . . . . . . . . . . . . . . . . . . . . . .  29
     15.1.  Normative References . . . . . . . . . . . . . . . . . .  29
     15.2.  Informative References . . . . . . . . . . . . . . . . .  30
   Appendix A.  Draft Performance Metrics Registry Entry . . . . . .  31
   Appendix B.  Reference Implementation . . . . . . . . . . . . . .  32
   Appendix C.  Open Questions for Review  . . . . . . . . . . . . .  32
   Author's Address  . . . . . . . . . . . . . . . . . . . . . . . .  33

1.  Introduction

   Conversational voice systems built from speech recognition, language
   modelling and speech synthesis components are now widely deployed
   over SIP and RTP.  Response latency is a primary determinant of
   whether such a system is usable, and it is measured today by several
   parties who do not agree on what is being measured.  There is no
   common definition of the quantity, no agreed reference point at which
   to observe it, and no convention for reporting its uncertainty.

Nygate                    Expires 11 March 2027                 [Page 3]
Internet-Draft        Mouth-to-Ear Response Latency       September 2026

   Figures produced under different assumptions are routinely compared
   as though they were commensurable, which is the practical problem
   this document exists to address.  Section 1.1 sets out the terms that
   separate one figure from another, and Section 1.2 describes the axes
   along which current measurement practice divides.

   This document takes no position on the magnitude of any term, on the
   accuracy of any published figure, or on the merits of any existing
   measurement effort.  It defines a quantity and a method of observing
   it, so that figures produced by different parties are comparable.

1.1.  Motivation and decomposition

   For a caller connected to a conversational voice system over SIP and
   RTP, the interval between the end of the caller's utterance and the
   arrival of response audio comprises, in order of occurrence:

      +=============================================================+
      | Term                                                        |
      +=============================================================+
      | de-jitter buffer depth on the inbound leg                   |
      +-------------------------------------------------------------+
      | voice activity detection and endpointing decision           |
      +-------------------------------------------------------------+
      | speech recognition finalisation after the endpoint decision |
      +-------------------------------------------------------------+
      | orchestration, retrieval and any tool invocation            |
      +-------------------------------------------------------------+
      | language model time to first token                          |
      +-------------------------------------------------------------+
      | speech synthesis time to first audio                        |
      +-------------------------------------------------------------+
      | encoding, packetisation and media relay                     |
      +-------------------------------------------------------------+
      | de-jitter buffer depth on the outbound leg                  |
      +-------------------------------------------------------------+

                                  Table 1

   A figure covering one or two of these terms does not predict the
   interval a caller experiences, and the terms are not separable by
   observation at the caller.  MRL is defined here as the aggregate,
   observed at a single reference point, precisely because the aggregate
   is what an external observer can measure without instrumenting the
   system under test.

Nygate                    Expires 11 March 2027                 [Page 4]
Internet-Draft        Mouth-to-Ear Response Latency       September 2026

1.2.  Existing practice

   Two kinds of figure are published today.

   Component-level figures are reported by the operators of individual
   components, most commonly the interval from a request reaching a
   speech synthesis endpoint to the first audio byte returned, and less
   commonly the interval to a language model's first token.  These are
   observed at an API boundary internal to the system and cover a subset
   of the terms above.

   Caller-side figures are produced by placing a call and observing the
   response.  At least one such effort publishes both its results and
   its tooling openly [TTFAB], reporting the interval between the
   caller's speech end and the onset of the response as observed in a
   recording of the call rather than in any timestamp reported by the
   system.

   Caller-side efforts differ from one another along axes that render
   their outputs incomparable, and those axes are the reason this
   document exists:

   *  the reference point at which the response is observed, which may
      be an RTP endpoint, a recording made by a carrier or a
      conferencing bridge, or an endpoint's audio device, each placing a
      different and frequently uncharacterised quantity of transport and
      buffering inside the measured interval.  Section 4.2 enumerates
      these so that a party observing anywhere on that list can produce
      a conforming report which declares the difference, and
      Section 4.2.1 gives a way to measure it rather than disclose it;

   *  the determination of the caller's speech end, which may be derived
      from a voice activity detector applied to the transmitted audio,
      or from prior annotation of known stimulus material;

   *  the treatment of buffering, since an interval taken at packet
      arrival and an interval taken after a de-jitter buffer differ by
      the depth of that buffer;

   *  the definition of response onset, which for audio that ramps in
      gradually is a choice of threshold rather than an observation;

   *  whether the measuring instrument has itself been calibrated
      against a known delay, and therefore whether a reported figure is
      a point estimate or an upper bound.

Nygate                    Expires 11 March 2027                 [Page 5]
Internet-Draft        Mouth-to-Ear Response Latency       September 2026

   An effort that states its choices along these axes produces figures a
   reader can interpret, whereas two efforts that have chosen
   differently produce figures which cannot be placed side by side even
   when each is internally correct.  This document specifies one set of
   choices and requires that they be reported alongside any figure
   derived under them.

1.3.  Scope

   This document specifies:

   *  the definition of the MRL metric and its two required variants;

   *  the method of measurement, including the reference point, the
      determination of the interval endpoints, and the required contents
      of a capture;

   *  the conditions under which a measurement MUST be treated as
      invalid;

   *  the fields that MUST accompany a reported figure.

   This document does not specify endpointing strategy, barge-in
   handling, speech recognition or synthesis behaviour, or any property
   of the system under test.  It does not specify a signalling protocol;
   the method assumes an established RTP session and is independent of
   how that session was established.

1.4.  Relationship to existing work

   The metric defined here is an application-layer performance metric in
   the sense of [RFC6390], and this document follows the template in
   Section 5.4 of that document.  [RFC6076] defines end-to-end
   performance metrics for telephony sessions at the SIP layer and is
   the closest existing IETF work in subject matter; the metric defined
   here concerns the media plane rather than signalling, and is
   complementary.  No existing IETF document defines the quantity
   described in Section 1.1 or specifies a reference point at which to
   observe it.

   The framework of [RFC2330] and the delay metric of [RFC7679] inform
   the treatment of error and uncertainty.  The method described here is
   an active method in the taxonomy of [RFC7799], since it generates the
   stimulus whose response it measures.

Nygate                    Expires 11 March 2027                 [Page 6]
Internet-Draft        Mouth-to-Ear Response Latency       September 2026

   [RFC3611] and its extensions define a reporting mechanism by which
   endpoints convey media quality metrics in band.  Conveying MRL in
   that manner is out of scope for this document and is noted in
   Section 14 as possible follow-on work.

      [[EDITOR'S NOTE, on venue.  The IPPM charter bounds the group's
      work to "metrics and methodologies which are applicable over
      transport-layer protocols over IP", while also covering
      "applications running over transport layer protocols".  Whether a
      metric whose dominant terms are speech recognition, inference and
      synthesis falls inside that boundary is a fair question and should
      be settled before effort is spent on -01.

      The precedent in favour is draft-ietf-ippm-responsiveness, an
      adopted IPPM work item targeting Proposed Standard, which measures
      at the application layer over HTTP/2 and HTTP/3 and justifies
      itself explicitly on user experience.  The argument against is
      that responsiveness remains a property of the network under load,
      whereas most of MRL is not attributable to the network at all.

      If IPPM declines, the alternatives are to take the question to
      DISPATCH, which exists for work with no obvious home, or to pursue
      publication through the Independent Submission stream.  The
      document is written to stand in any of the three.

      Separately, decide whether to pursue registration in the
      Performance Metrics Registry ([RFC8911], initially populated by
      [RFC8912]); Appendix A holds a draft entry.  The registry has to
      date been populated with IP-layer path metrics, so eligibility
      should be confirmed before the appendix is presented as a
      proposal.]]

2.  Conventions and Definitions

   The key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT",
   "SHOULD", "SHOULD NOT", "RECOMMENDED", "NOT RECOMMENDED", "MAY", and
   "OPTIONAL" in this document are to be interpreted as described in
   BCP 14 [RFC2119] [RFC8174] when, and only when, they appear in all
   capitals, as shown here.

   The following terms are used throughout.

   Calling endpoint:  The endpoint that generates the stimulus utterance
      and observes the response.  All measurement is performed here.

   System under test (SUT):  The conversational voice system that
      receives the stimulus and generates a response.  The method treats
      the SUT as opaque.

Nygate                    Expires 11 March 2027                 [Page 7]
Internet-Draft        Mouth-to-Ear Response Latency       September 2026

   Reference point:  The point at which timestamps are taken.  See
      Section 4.1.

   Stimulus:  A prerecorded utterance transmitted by the calling
      endpoint to elicit a response.

   Speech end:  The final sample of speech in the stimulus, determined
      offline.  See Section 3.3.

   Response onset:  The first sample of the SUT's response audio,
      determined offline.  See Section 3.5.

3.  Metric Definition

3.1.  Metric name

   Mouth-to-Ear Response Latency (MRL), reported in two variants named
   Ingress MRL and Playout MRL.

3.2.  Metric description

   MRL is the interval between the instant at which the calling endpoint
   transmits the final speech sample of a stimulus utterance and the
   instant at which the first sample of the SUT's response reaches the
   calling endpoint.

   MRL is defined as:

   MRL = t1 - t0

   where t0 and t1 are as defined in Section 3.3 and Section 3.4.

   MRL is signed.  A negative value indicates that response audio
   reached the calling endpoint before the stimulus had finished being
   transmitted, which occurs with aggressive endpointing and with
   backchannel responses.  Implementations MUST report negative values
   as measured and MUST NOT clamp them to zero.

3.3.  Interval start: t0

   t0 is the instant at which the final speech sample of the stimulus is
   transmitted by the calling endpoint at the reference point.

   t0 MUST be determined by:

   1.  annotating the stimulus offline to sample precision to locate the
       speech end;

Nygate                    Expires 11 March 2027                 [Page 8]
Internet-Draft        Mouth-to-Ear Response Latency       September 2026

   2.  mapping that sample onto the RTP packet that carried it, using
       RTP timestamps as defined in [RFC3550];

   3.  interpolating within that packet at the sample rate to obtain the
       sample's offset from the packet's transmission instant.

   t0 is therefore the transmission instant of the carrying packet plus
   the target sample's offset within that packet.  The sign of the
   offset term is significant; see Section 4.7.

   t0 MUST NOT be derived from voice activity detection performed at run
   time.  A run-time detector has a decision lag of its own, that lag is
   a term within the quantity being measured, and using it to define t0
   would conceal the term.

   Annotation of the stimulus MUST NOT extend the speech end by a fixed
   constant.  A fixed extension displaces t0 later by that constant and,
   since MRL is t1 minus t0, reduces every reported figure by the same
   constant.  See Section 4.7.

   An implementation MUST NOT be required to use any particular
   annotation algorithm, and MUST be able to demonstrate that its
   annotation agrees with a published reference.  The reference
   implementation uses decay-following hysteresis with a short sliding
   RMS refinement, and publishes a frozen stimulus signal whose speech
   end is exact by construction, together with the boundary its own
   annotator locates [HARNESS].  An implementation claiming conformance
   SHOULD reproduce that boundary and MUST state the deviation where it
   does not.

   Specifying the conformance test rather than the algorithm is
   deliberate.  The quantity that matters to a reader is whether two
   parties locate the same boundary in the same audio, and an algorithm
   mandated in prose can be implemented differently by two careful
   people while a boundary in a published waveform cannot be argued
   about.

      [[EDITOR'S NOTE: the reference signal is frozen bytes rather than
      a generator seed, because a signal rebuilt by its generator
      changes when the generator does.  Whether the signal itself should
      be carried in an IANA registry, an appendix, or by reference to
      the archived implementation is a question for the working group;
      reference is assumed here.]]

3.4.  Interval end: t1

   t1 is the instant at which the first sample of the SUT's response
   reaches the calling endpoint, taken at the reference point.

Nygate                    Expires 11 March 2027                 [Page 9]
Internet-Draft        Mouth-to-Ear Response Latency       September 2026

   t1 MUST be determined by reassembling the received stream in RTP
   timestamp order, locating the response onset as specified in
   Section 3.5, and mapping the onset sample back through the packet
   that carried it.

   Two variants of t1 are defined, and both MUST be reported.

3.4.1.  Ingress MRL

   Ingress MRL takes t1 at the arrival instant of the onset sample, that
   is, at the reference point with no buffering applied.

   Ingress MRL isolates the contribution of the SUT and of the network
   path from the buffering policy of the calling endpoint.  Its variance
   under a jittered path tracks the jitter of that path, because the
   metric reports the path as it finds it, and an implementation showing
   less variance than the path exhibits is smoothing a quantity it was
   asked to observe.

3.4.2.  Playout MRL

   Playout MRL takes t1 at the instant at which the onset sample would
   be released from a de-jitter buffer of stated target depth.

   The target depth MUST be reported with the figure.  The de-jitter
   model used MUST anchor on the minimum transit delay observed over a
   stated initial window, which is the behaviour adaptive buffers
   converge toward, rather than on the arrival instant of the first
   packet received.

   Playout MRL corresponds to the interval a caller waits and is the
   variant to compare against conversational turn-taking norms.

   Playout MRL is distinct from the one-way transmission delay addressed
   by [G114], and the two are not interchangeable.  One-way transmission
   delay concerns the time taken to carry audio across a path and
   applies to a conversation between two people, whereas MRL concerns
   the time a system takes to begin responding and includes transmission
   delay as one term among several.  A system may satisfy the
   transmission delay guidance in [G114] on both legs and still exhibit
   an MRL an order of magnitude larger.

Nygate                    Expires 11 March 2027                [Page 10]
Internet-Draft        Mouth-to-Ear Response Latency       September 2026

3.4.3.  Both variants required

   An implementation MUST report both variants.  Ingress MRL on its own
   understates the interval a caller experiences, while a bare Playout
   MRL leaves the SUT confounded with whatever buffering policy the
   calling endpoint happened to apply, so neither figure can be
   interpreted without the other.

3.5.  Response onset

   The instant at which audio begins has no unique definition, because
   synthesised speech commonly ramps in over tens of milliseconds rather
   than beginning at full level.  Onset is therefore computed under
   three named variants, and the dispersion across them is a component
   of the reported uncertainty.

    +===========+===================+================+===============+
    | Variant   | Above noise floor | Absolute floor | Sustained for |
    +===========+===================+================+===============+
    | sensitive | 6 dB              | -55 dBov       | 10 ms         |
    +-----------+-------------------+----------------+---------------+
    | headline  | 10 dB             | -50 dBov       | 20 ms         |
    +-----------+-------------------+----------------+---------------+
    | strict    | 12 dB             | -45 dBov       | 30 ms         |
    +-----------+-------------------+----------------+---------------+

                                 Table 2

   Levels are expressed in dBov, referenced to full-scale RMS.

   The headline variant is the reported figure.  The spread of MRL
   across all three variants is the onset-definition uncertainty of the
   measurement and MUST be published alongside any headline figure.  On
   an abrupt onset the spread is small; on a gradual ramp it grows with
   the ramp duration and becomes the dominant uncertainty term.

   Sub-frame refinement of the onset instant MUST use a short sliding
   RMS.  Instantaneous sample magnitude crosses any fixed threshold on
   isolated noise peaks at a rate high enough to displace the boundary
   materially; see Section 4.7.

3.6.  Unprompted audio

   A conversational voice system commonly speaks before the caller does,
   opening with a greeting that arrives within a few hundred
   milliseconds of the session being established.  That audio responds
   to nothing, because the caller has not yet spoken.

Nygate                    Expires 11 March 2027                [Page 11]
Internet-Draft        Mouth-to-Ear Response Latency       September 2026

   Response-onset detection MUST NOT begin before the end of any such
   unprompted audio.  The point at which it ended MUST be recorded in
   the capture, and the noise-floor estimate required by Section 3.5
   MUST be taken from audio following that point.

   Detection across the whole received stream locates the greeting
   instead of the response.  Since the greeting precedes t0, the
   resulting MRL is large and negative, and because this document
   licenses negative values as genuine behaviour there is no bound
   against which such an error announces itself.  In the reference
   implementation, a capture carrying an 800 ms greeting with a true MRL
   of 900 ms yielded -2085 ms and satisfied every condition in
   Section 8.

   A greeting falling inside the noise-floor window is speech rather
   than channel noise, so it raises the estimate and displaces every
   onset threshold derived from it.  On the same capture the floor moved
   by 1.35 dB.

   A calling endpoint SHOULD NOT begin transmitting its stimulus while
   unprompted audio is still in progress.  A system that implements
   barge-in detection will stop speaking when it hears the caller, which
   truncates the greeting and alters the interaction under measurement.

   The interval from session establishment to the onset of unprompted
   audio is a distinct and useful quantity, since a caller who hears
   nothing for several seconds after the line opens is poorly served
   whatever the system's MRL turns out to be.  It shares no terms with
   MRL and is not defined here; see Section 14.

3.7.  Filler audio and response continuity

   t1 as defined in Section 3.4 is the onset of the system's first
   response audio, whatever that audio happens to contain.  A system
   that emits an earcon, a breath or a filled pause while its response
   is still being generated therefore records a low MRL while conveying
   nothing during that interval, and would rank above a system that
   stayed quiet and then answered.

   This document does not resolve that by identifying which audio
   carries meaning.  A metric incorporating a judgement about meaning
   cannot be re-derived from a published capture by an independent
   reviewer, and that reproducibility is the property which makes the
   rest of this specification worth having.  The discriminator is
   structural instead: filler is followed by silence before the
   substantive response begins, and continuous speech is not.

Nygate                    Expires 11 March 2027                [Page 12]
Internet-Draft        Mouth-to-Ear Response Latency       September 2026

   An implementation MUST, within a window of 2000 ms following t1,
   measure the longest interval whose level lies below the onset
   threshold of the headline variant.  Intervals carried by no packet
   MUST be excluded from that measurement, because a lost or late-
   discarded frame leaves a gap indistinguishable from a deliberate
   pause and would otherwise allow a degraded path to manufacture
   filler.

   Where that interval exceeds 150 ms:

   *  the response MUST be reported as discontiguous;

   *  a second onset MUST be reported, at the start of the final
      contiguous segment, together with the MRL derived from it.

   MRL itself is unchanged by this section, so no figure measured under
   an earlier revision becomes invalid.  A reader receives both onsets
   and can see whether they differ and by how much.

   The 2000 ms window and the 150 ms threshold are fixed by this
   document rather than left to the implementation, for the reason given
   in Section 3.5: figures derived under different parameters cannot be
   compared even when each is internally correct.  Both MUST be reported
   alongside any figure derived under them.

4.  Method of Measurement

4.1.  Reference point

   All timestamps MUST be taken at the RTP egress and ingress reference
   point of the calling endpoint, that is, immediately before a packet
   is passed to the operating system for transmission and immediately
   after a packet is received from it.

   The reference point is chosen so that the measurement includes every
   term a caller experiences downstream of the calling endpoint's own
   send path, and excludes the calling endpoint's own playout hardware,
   which is a property of the observer rather than of the SUT.

4.2.  Observing elsewhere

   Not every party measuring this quantity controls an RTP endpoint.  A
   figure derived from a carrier's call recording is a real measurement
   of something a caller experienced, and declaring it as such is more
   useful than being unable to describe it at all.

Nygate                    Expires 11 March 2027                [Page 13]
Internet-Draft        Mouth-to-Ear Response Latency       September 2026

   An implementation MAY therefore observe at a point other than the one
   in Section 4.1, provided it declares which, and provided it discloses
   what that choice places inside the measured interval.  Such a report
   conforms to this document.  What it does not do is produce a figure
   comparable with one taken at the RTP reference point, because the two
   intervals contain different terms.

Nygate                    Expires 11 March 2027                [Page 14]
Internet-Draft        Mouth-to-Ear Response Latency       September 2026

   +=======================+===================+=======================+
   | reference_point       | observed at       | additional terms      |
   |                       |                   | inside the interval   |
   +=======================+===================+=======================+
   | rtp-endpoint          | immediately       | none; this is         |
   |                       | before the        | Section 4.1           |
   |                       | packet is passed  |                       |
   |                       | to the operating  |                       |
   |                       | system, and       |                       |
   |                       | immediately       |                       |
   |                       | after it is       |                       |
   |                       | received from it  |                       |
   +-----------------------+-------------------+-----------------------+
   | host-packet-capture   | a packet capture  | the host's own        |
   |                       | facility on the   | network stack, and    |
   |                       | calling host      | asymmetrically: on    |
   |                       |                   | transmission the      |
   |                       |                   | capture is taken      |
   |                       |                   | after the send path,  |
   |                       |                   | on reception before   |
   |                       |                   | the application reads |
   +-----------------------+-------------------+-----------------------+
   | bridge-recording      | a session border  | the leg between       |
   |                       | controller or     | caller and bridge in  |
   |                       | conferencing      | both directions, and  |
   |                       | bridge in the     | the bridge's own      |
   |                       | path              | media handling        |
   +-----------------------+-------------------+-----------------------+
   | carrier-recording     | a recording made  | the leg to the        |
   |                       | by a carrier or   | carrier in both       |
   |                       | communications    | directions, any       |
   |                       | platform          | transcoding it        |
   |                       |                   | performs, its         |
   |                       |                   | buffering, and its    |
   |                       |                   | recording pipeline    |
   +-----------------------+-------------------+-----------------------+
   | endpoint-audio-device | the caller's      | the endpoint's de-    |
   |                       | audio device, or  | jitter buffer, its    |
   |                       | a loopback        | audio stack and the   |
   |                       | capture of it     | device's own latency  |
   +-----------------------+-------------------+-----------------------+

                                  Table 3

Nygate                    Expires 11 March 2027                [Page 15]
Internet-Draft        Mouth-to-Ear Response Latency       September 2026

   Where a value other than rtp-endpoint is declared, the additional
   terms MUST be disclosed and the figure MUST be reported as an upper
   bound rather than a point estimate, unless those terms have been
   characterised as below.  Figures taken at different reference points
   MUST NOT be pooled into one distribution.

4.2.1.  Characterising the additional terms

   An upper bound is a weaker result than it needs to be, and the
   additional terms are measurable rather than merely acknowledgeable.

   Where the path to the system under test traverses infrastructure the
   measuring party does not control, an implementation SHOULD place a
   second call over the same route to a responder replying at a delay it
   has programmed, and subtract that measurement from the first.  The
   infrastructure's contribution is common to both and cancels, leaving
   the system's response.  Two routes over the same carrier are not
   identical, so this bounds the additional terms rather than
   eliminating them exactly, and the bound is what should be reported.

   This is the same differential construction the calibration in
   Section 4.7 uses, applied to a term outside the instrument instead of
   inside it.  It turns a disclosed caveat into a measured quantity, and
   it is available to any party that can dial a number it controls.

4.3.  Stimulus requirements

   Stimulus material MUST be prerecorded and MUST be transmitted at the
   nominal frame rate of the codec in use.  The stimulus MUST be hashed
   and the hash MUST be recorded with the capture, so that a changed
   stimulus invalidates a comparison loudly rather than silently.

   The method is independent of the codec, and the codec in use MUST be
   reported with any figure derived under it.  Where a payload format
   from [RFC3551] is used, the sample rate and frame period follow that
   profile, and the companding of [G711] applies to the PCMU and PCMA
   formats.

4.4.  Transmission pacing

   The calling endpoint MUST measure the deviation of its own
   transmission instants from the nominal frame grid and MUST record the
   worst deviation observed during the call.  An endpoint that cannot
   pace its own transmission has an unreliable t0 and therefore an
   unreliable MRL.  See Section 8.

Nygate                    Expires 11 March 2027                [Page 16]
Internet-Draft        Mouth-to-Ear Response Latency       September 2026

4.5.  Capture contents

   A capture MUST contain raw payloads and raw timestamps only.  A
   capture MUST NOT contain any derived quantity, including any latency
   figure.

   This requirement exists so that a revised onset definition, or a
   definition proposed by a reviewer, can be applied to existing data
   without repeating a collection.  The uncertainty analysis in
   Section 3.5 is not possible otherwise.

4.6.  Clock requirements

   Both t0 and t1 MUST be taken from a single monotonic clock on the
   calling endpoint.

   Because the interval is the difference of two timestamps drawn from
   one clock on one host, the metric requires no synchronisation between
   the calling endpoint and the SUT, and is insensitive to offset
   between them.  This is a deliberate property of the definition and
   distinguishes the method from one-way delay measurement, where clock
   synchronisation between the two hosts dominates the error budget as
   described in Section 3.7.1 of [RFC7679].

   A wall-clock timestamp MAY be recorded in parallel for the sole
   purpose of correlating captures with traces obtained from the SUT.
   Any offset or skew in that clock affects such correlation only and
   MUST NOT affect the reported MRL.

4.7.  Sources of error and calibration

   An implementation MUST state its calibrated accuracy, and MUST state
   the conditions under which that calibration was obtained.

   Calibration is performed by replacing the SUT with a reference
   responder that replies at a programmed delay, so that ground truth is
   known exactly.  Under each channel condition the offset from ground
   truth that the physics of the channel requires is predictable in
   advance: for Ingress MRL it is the base transit delay plus the mean
   jitter excess, and for Playout MRL it is the base transit delay plus
   the buffer target depth.  The calibration criterion is therefore the
   residual after subtracting that predicted offset, rather than the raw
   difference from ground truth.

Nygate                    Expires 11 March 2027                [Page 17]
Internet-Draft        Mouth-to-Ear Response Latency       September 2026

   Calibration MUST NOT be performed exclusively at programmed delays
   that are integer multiples of the frame period.  A responder that
   evaluates its emission deadline once per received frame quantises its
   own output to the frame period, and that error is invisible at frame-
   commensurate delays because the deadline then falls on a frame
   boundary.

   The following error mechanisms are known to produce plausible-looking
   but incorrect figures, and an implementation is advised to test for
   each:

   Fixed annotation extension:  Extending the stimulus speech end by a
      constant biases every figure low by that constant.

   Instantaneous-magnitude thresholding, at either boundary:  Sub-frame
      refinement on sample magnitude rather than a short sliding RMS
      lets isolated noise peaks cross the threshold.  At the stimulus
      end boundary this displaces t0 late; at the response onset it
      displaces t1 early, biasing MRL low by up to one analysis window.
      The second was masked in the reference implementation for as long
      as its calibration responder placed responses on the analysis
      grid, and surfaced only once a media-clock responder placed them
      where they began: it accounted for the whole of a -0.40 ms
      headline bias and a 0.99 ms spread between onset variants on a
      hard onset.  A fix applied at one boundary has to be checked
      against its mirror.

   Frame-quantised reference responder:  Adds a uniform error between
      zero and one frame period, invisible at frame-commensurate
      calibration delays.

   Frame-offset sign error:  Using the wrong sign for the within-frame
      offset term of t0 produces a residual of twice the offset with the
      opposite sign.  A large bias accompanied by a tight spread is the
      signature of a definitional or arithmetic error rather than of
      host timing noise, and implementations are advised to report that
      discrimination automatically.

   Symmetric jitter modelling:  A simulated channel that models network
      jitter as zero-mean Gaussian permits a minimum-tracking de-jitter
      anchor to sit earlier than the minimum transit delay allows, which
      appears as a negative bias in Playout MRL.  Network jitter is one-
      sided and has a hard floor at the minimum transit delay.  Sender
      pacing deviation is a different mechanism and is symmetric,
      because a timer-driven sender can fire either early or late.

   Calibration source that is not an honest RTP sender:  A reference

Nygate                    Expires 11 March 2027                [Page 18]
Internet-Draft        Mouth-to-Ear Response Latency       September 2026

      responder whose RTP timestamps count frames rather than follow a
      media clock, or whose idle-stream pacing depends on whether it is
      receiving anything, produces playout figures with errors that
      ingress cannot see, since ingress is derived from arrival and
      playout through the timestamp.  Observed as a 44.7 ms transit slip
      that flagged late discard on jitter-free calls, and 9.54 ms of
      playout spread that no buffer target reduced.  Both surfaced only
      from kept captures over a real path.

      [[EDITOR'S NOTE: the reference implementation's calibrated figures
      are published in [HARNESS] and are deliberately not reproduced
      here.  A metric specification that carries one implementation's
      results becomes stale and invites the reader to treat those
      figures as a conformance target.  Confirm this is the right call
      in review.]]

5.  Units of Measurement

   MRL is reported in milliseconds, signed, to a resolution of not
   coarser than 0.1 ms.

   Uncertainty terms accompanying the figure are reported in the same
   units.

6.  Measurement Points and Measurement Domain

   The single measurement point is the calling endpoint's RTP reference
   point as defined in Section 4.1.  No observation of the SUT's
   internal state is required, and none is assumed to be available.

   The measurement domain is the path between the calling endpoint and
   the SUT together with the SUT itself.  The method does not separate
   the two, and reported figures are therefore properties of the pairing
   rather than of the SUT alone.  Where separation is required, the path
   contribution has to be characterised independently and reported
   alongside.

7.  Measurement Timing

   Each measurement corresponds to one stimulus and one response.  A
   reported distribution MUST state the number of measurements attempted
   and the number discarded, as required by Section 8.

Nygate                    Expires 11 March 2027                [Page 19]
Internet-Draft        Mouth-to-Ear Response Latency       September 2026

7.1.  Turns within a session

   Several exchanges within one session are permitted and are the
   realistic case, since a caller does not hang up after one question.
   They are not interchangeable, though, and a report that pools them
   without saying so is not comparable with one that does not.

   A system's first response in a session carries terms that later
   responses do not: session establishment, whatever a model does on a
   cold start, caches not yet populated, and connections not yet open.
   Later responses may in turn benefit from conversational state the
   system has accumulated.  Both effects are real properties of the
   system rather than measurement artefacts, and both are invisible in
   an aggregate that does not record which turn each measurement came
   from.

   Therefore:

   *  each measurement MUST record its turn index within the session,
      counting from one;

   *  a reported distribution MUST state which turn indices it covers;

   *  first-turn measurements MUST NOT be pooled with later ones unless
      the report states that they are pooled and gives the counts of
      each.

   This specifies what must be reported rather than how many turns to
   measure.  A study of cold-start behaviour measuring only first turns,
   and a study of steady-state behaviour discarding them, are both valid
   and are answering different questions; what makes them comparable to
   each other is that both said which they did.

   A calling endpoint conducting a multi-turn session MUST NOT transmit
   its next stimulus while the system is still speaking, for the reason
   given in Section 3.6: a system implementing barge-in detection will
   stop, which alters the interaction under measurement.  Locating the
   end of the system's turn is the same problem as locating the end of a
   greeting and admits the same solution.

8.  Quality Control and Result Validity

   Two classes of condition are defined.

8.1.  Blocking conditions

   A measurement exhibiting any of the following MUST be discarded and
   MUST NOT be reported as a figure:

Nygate                    Expires 11 March 2027                [Page 20]
Internet-Draft        Mouth-to-Ear Response Latency       September 2026

   *  no speech end could be located in the stimulus;

   *  no packets were received from the SUT;

   *  no response onset was found;

   *  the response onset fell within the first received frame, so that
      the true onset may precede the observation window;

   *  the worst transmission pacing deviation exceeded a stated
      threshold.

   The pacing threshold in the reference implementation is 5 ms.  A
   calling endpoint that cannot pace its own transmission within that
   bound has an unreliable t0.

   Discarding a measurement under these conditions is correct behaviour.
   An implementation MUST NOT relax a blocking condition in order to
   retain a measurement.

8.2.  Advisory conditions

   The following conditions do not invalidate a measurement but MUST be
   reported with it:

   *  packet loss above a stated threshold;

   *  late discard above a stated threshold.

   A call over a lossy path remains a valid measurement of a lossy path.
   Where the frame carrying the response onset is itself lost, onset
   detection is deferred by whole frames and the resulting figure is an
   upper bound rather than a point estimate, which is why the flag has
   to travel with the number.

8.3.  Reporting of discards

   Every reported distribution MUST be accompanied by the count of
   measurements discarded.  A run that discards a large fraction of its
   measurements is not comparable with one that discards none,
   irrespective of the percentiles of the survivors.

9.  Reporting

9.1.  Required fields

   A reported MRL figure MUST be accompanied by:

Nygate                    Expires 11 March 2027                [Page 21]
Internet-Draft        Mouth-to-Ear Response Latency       September 2026

   *  Ingress MRL and Playout MRL, both signed;

   *  the playout target depth;

   *  the onset variant used for the headline figure, and the spread
      across all three variants;

   *  the codec and frame period;

   *  the number of measurements attempted and the number discarded;

   *  any advisory flags raised;

   *  the calibrated accuracy of the measuring implementation and the
      conditions of that calibration;

   *  an identifier for the SUT configuration sufficient to establish
      that two figures refer to the same configuration, without
      necessarily disclosing that configuration.

9.2.  Reporting format

   A reported result MUST be expressible as a JSON object [RFC8259]
   carrying the members defined below.  The format exists so that a
   reader can establish, without contacting the party who produced a
   figure, whether two figures were derived under the same choices.

   The schema member MUST be present and MUST be the string mrl-report/1
   for reports conforming to this document.

9.2.1.  instrument

   Describes the measuring implementation and its calibration.  All
   members are REQUIRED.

   name, version:  Identify the implementation that produced the report.

   calibration:  An object recording the outcome of the procedure in
      Section 4.7.  It MUST carry conditions, a human-readable statement
      of the channel and host conditions under which calibration was
      performed, and for each of ingress and playout a bias_ms and a
      p95_abs_error_ms.  It SHOULD carry reference, a URI or DOI at
      which the calibration evidence can be inspected.  An
      implementation that has not been calibrated MUST set calibration
      to null rather than omitting it, and every figure in such a report
      is an upper bound rather than a point estimate.

Nygate                    Expires 11 March 2027                [Page 22]
Internet-Draft        Mouth-to-Ear Response Latency       September 2026

9.2.2.  measurement

   Describes the choices along the axes in Section 1.2.  All members are
   REQUIRED.

   reference_point:  Where t1 was observed, taking one of the values
      enumerated in Section 4.2.  Any value other than rtp-endpoint MUST
      be accompanied by reference_point_notes stating what that choice
      places inside the measured interval, and by
      reference_point_bound_ms where the additional terms have been
      characterised as in Section 4.2.1, or null where they have not and
      the figures are therefore upper bounds.

   codec, frame_period_ms, sample_rate_hz:  The media parameters in
      force.

   playout_target_ms:  The de-jitter buffer target depth used to derive
      Playout MRL.

   onset_variants:  An array of objects, each carrying name, margin_db,
      absolute_dbov and sustain_ms.  The parameters MUST be stated
      rather than referenced by name alone, so that a report remains
      interpretable if the defaults in Section 3.5 are ever revised.

   headline_variant:  The name of the variant whose figures are quoted
      as the headline.

   continuity_window_ms, continuity_gap_threshold_ms:  The parameters of
      Section 3.7, stated rather than assumed so that a report remains
      interpretable if the defaults are ever revised.

9.2.3.  subject

   Identifies what was measured. stimulus_id and stimulus_sha256 are
   REQUIRED. sut_identifier is REQUIRED and sut_config_sha256 is
   RECOMMENDED, the latter allowing two reports to be shown to concern
   the same configuration without that configuration being disclosed.

9.2.4.  results

   Aggregate figures.  All members are REQUIRED.

   n_attempted, n_reported, n_discarded:  Counts of measurements.
      n_attempted MUST equal n_reported plus n_discarded.

   turn_indices:  The turn indices this distribution covers, and the

Nygate                    Expires 11 March 2027                [Page 23]
Internet-Draft        Mouth-to-Ear Response Latency       September 2026

      count of measurements from each.  REQUIRED by Section 7.1, which
      forbids pooling a first turn with later ones without saying so.  A
      single-turn study reports one entry, which is the honest way to
      say that the figures describe cold starts.

   discard_reasons:  An object mapping each blocking condition in
      Section 8 to the number of measurements it discarded.  The counts
      MUST sum to n_discarded.

   ingress, playout:  Objects carrying at least mean_ms, p50_ms, p95_ms
      and max_ms, computed over the reported measurements under the
      headline variant.  Values are signed.

   onset_definition_uncertainty_ms:  The spread of MRL across all
      variants in onset_variants, carrying at least p50 and max.  This
      member MUST be present, since a headline figure quoted without it
      is incomplete under Section 3.5.

   continuity:  An object carrying gap_p50_ms and gap_max_ms over the
      reported measurements, n_discontiguous, and a contiguous object of
      the same shape as ingress giving the MRL to the start of
      uninterrupted speech.  Where n_discontiguous is zero the
      contiguous figures equal the ingress figures, which is the
      expected case and is reported rather than omitted so that its
      absence never has to be inferred.

   advisory_flags:  An object mapping each advisory condition in
      Section 8 to the number of reported measurements carrying it.

9.2.5.  measurements and captures

   measurements SHOULD carry one object per individual measurement, each
   with the measurement's identifier, its session identifier and turn
   index as required by Section 7.1, its t0 and t1 in nanoseconds on the
   instrument's monotonic clock, its ingress and playout MRL under every
   variant, the continuity gap and contiguous onset from Section 3.7,
   the end of any unprompted audio as required by Section 3.6, and any
   flags raised. captures SHOULD carry one object per capture with the
   measurement identifier, a sha256 of the capture file, and a URI at
   which it can be obtained.

   Both members are optional because a party may be unable to publish
   raw material.  Omitting them removes the reader's ability to re-
   derive the figures under a different onset definition, which
   Section 3.5 identifies as the dominant uncertainty term, so a report
   omitting them is weaker evidence than one that includes them.

Nygate                    Expires 11 March 2027                [Page 24]
Internet-Draft        Mouth-to-Ear Response Latency       September 2026

9.2.6.  Example

   The following is a report with the per-measurement and capture arrays
   elided.

   {
     "schema": "mrl-report/1",
     "instrument": {
       "name": "voice-ai-latency-harness",
       "version": "0.1.1",
       "calibration": {
         "conditions": "clean channel, 20 ms grid, PCMU",
         "ingress": { "bias_ms": -0.40, "p95_abs_error_ms": 2.38 },
         "playout": { "bias_ms": -0.40, "p95_abs_error_ms": 2.38 },
         "reference": "https://doi.org/10.5281/zenodo.22124823"
       }
     },
     "measurement": {
       "reference_point": "rtp-endpoint",
       "reference_point_notes": null,
       "reference_point_bound_ms": null,
       "codec": "PCMU",
       "frame_period_ms": 20.0,
       "sample_rate_hz": 8000,
       "playout_target_ms": 40.0,
       "onset_variants": [
         { "name": "sensitive", "margin_db": 6.0,
           "absolute_dbov": -55.0, "sustain_ms": 10.0 },
         { "name": "headline", "margin_db": 10.0,
           "absolute_dbov": -50.0, "sustain_ms": 20.0 },
         { "name": "strict", "margin_db": 12.0,
           "absolute_dbov": -45.0, "sustain_ms": 30.0 }
       ],
       "headline_variant": "headline",
       "continuity_window_ms": 2000.0,
       "continuity_gap_threshold_ms": 150.0
     },
     "subject": {
       "sut_identifier": "system-A",
       "sut_config_sha256": "9f2b...c41e",
       "stimulus_id": "eval-set-1/utt-017",
       "stimulus_sha256": "3ad1...77b0"
     },
     "results": {
       "n_attempted": 20,
       "n_reported": 18,
       "n_discarded": 2,
       "turn_indices": { "1": 8, "2": 5, "3": 5 },

Nygate                    Expires 11 March 2027                [Page 25]
Internet-Draft        Mouth-to-Ear Response Latency       September 2026

       "discard_reasons": {
         "onset_not_found": 1, "tx_pacing_deviation": 1
       },
       "ingress": {
         "mean_ms": 812.4, "p50_ms": 796.0,
         "p95_ms": 1043.2, "max_ms": 1101.7
       },
       "playout": {
         "mean_ms": 852.4, "p50_ms": 836.0,
         "p95_ms": 1083.2, "max_ms": 1141.7
       },
       "onset_definition_uncertainty_ms": { "p50": 4.7, "max": 9.8 },
       "continuity": {
         "gap_p50_ms": 48.0,
         "gap_max_ms": 512.0,
         "n_discontiguous": 4,
         "contiguous": {
           "mean_ms": 941.7, "p50_ms": 802.0,
           "p95_ms": 1418.6, "max_ms": 1461.0
         }
       },
       "advisory_flags": {
         "high_loss": 1,
         "high_late_discard": 0,
         "discontiguous_response": 4
       }
     }
   }

      [[EDITOR'S NOTE: this structure is deliberately close to what the
      reference implementation already emits, which keeps at least one
      producer honest, and it is the part of the document most likely to
      change on review.  Two questions in particular.  Whether the
      format should be registered as a media type.  And whether
      reference_point should be an enumeration with values for carrier-
      side and bridge-side recording, so that efforts observing
      elsewhere can produce conforming reports that declare the
      difference, rather than being unable to conform at all.  The
      second would widen adoption considerably and is probably worth
      doing.]]

10.  Use and Applications

   The metric supports:

   *  comparison of conversational voice systems from the position of a
      caller;

Nygate                    Expires 11 March 2027                [Page 26]
Internet-Draft        Mouth-to-Ear Response Latency       September 2026

   *  characterisation of a single system across configurations or load
      levels;

   *  regression testing of a deployed system over time.

   MRL measures an interval and carries no information about what the
   response contained.  Response quality, recognition accuracy and the
   usefulness of the reply are separate quantities needing separate
   instruments, and a system that answers quickly and wrongly will score
   well here.  Cases in which the metric does not apply, or applies only
   with care, are set out in Section 11.

11.  Applicability and Limitations

   The metric as defined applies to systems that respond to a completed
   caller utterance with a discrete response.  The following cases are
   outside its current scope or require care:

   Filler and non-lexical audio:  Addressed by the continuity
      measurement in Section 3.7, which separates first audio from first
      uninterrupted audio without interpreting content.  A caller
      comparing systems should read both figures, since MRL alone
      rewards a system for making a noise.

   Unprompted audio:  Handled by Section 3.6, which requires detection
      to begin after any greeting.  Without that constraint the
      measurement locates the greeting, and the figure it produces
      describes a different event.

   Barge-in:  Where the caller interrupts the system, the interval
      defined here is not the quantity of interest and the method does
      not address it.

   Incremental and streaming responses:  Where a system emits partial
      audio that it subsequently revises, first-audio onset does not
      characterise the response.

   Path contribution:  As noted in the measurement domain, the figure is
      a property of the pairing of path and SUT.  Where the path
      includes infrastructure the measuring party does not control,
      Section 4.2.1 bounds its contribution rather than leaving it as a
      caveat.

   Conditioning across turns:  See the open item under Measurement
      Timing.

Nygate                    Expires 11 March 2027                [Page 27]
Internet-Draft        Mouth-to-Ear Response Latency       September 2026

12.  Security Considerations

   The method described here generates traffic directed at a system
   under test and elicits responses from it.  Measurement of a system
   operated by another party without that party's authorisation may
   constitute unauthorised use, may breach terms of service, and at
   sufficient volume is indistinguishable from a denial of service
   attempt.  Parties performing measurement:

   *  SHOULD obtain authorisation from the operator of the system under
      test;

   *  SHOULD limit the measurement rate to a level agreed with that
      operator;

   *  SHOULD identify their measurement traffic where a mechanism to do
      so exists.

   Captures produced by this method contain audio.  Where stimulus
   material contains recorded human speech, that material is personal
   data in many jurisdictions and may carry biometric significance, and
   captures also hold the SUT's response audio, which can reflect the
   content of the stimulus.  Implementers:

   *  SHOULD prefer synthetic or consented stimulus material;

   *  SHOULD apply access control to stored captures;

   *  SHOULD set a retention limit rather than keeping captures
      indefinitely.

   Publication of captures as measurement evidence, which the method
   otherwise encourages, has to be weighed against these considerations.

   Reported figures may be commercially sensitive to the operator of the
   system under test.  The configuration identifier described in
   Section 9 is specified as an identifier rather than as the
   configuration itself so that reproducibility and confidentiality can
   coexist.

13.  IANA Considerations

   This document has no IANA actions.

   Registration of this metric in the Performance Metrics Registry
   established by [RFC8911] would be a reasonable later step, and
   Appendix A shows what such an entry would contain.  It is
   deliberately not requested here.  The registry has to date been

Nygate                    Expires 11 March 2027                [Page 28]
Internet-Draft        Mouth-to-Ear Response Latency       September 2026

   populated with IP-layer path metrics [RFC8912], so eligibility for a
   metric of this kind is a question for the working group rather than
   an entitlement to assert in an individual submission, and a request
   made before the venue question of Section 1.2 is settled would be
   moot if this work moves elsewhere.

14.  Future Work

   Conveying MRL between endpoints in band, by means of an RTCP Extended
   Report block in the manner of [RFC3611], would allow a system to
   report the metric about itself rather than requiring an external
   caller to measure it.  That is a separate document and a different
   working group.

   Time to greeting, the interval from session establishment to the
   onset of the unprompted audio described in Section 3.6, is a second
   quantity the method already observes without defining.  It is at
   least as visible to a caller as MRL, and it appears to be published
   by nobody.  Defining it here would widen this document beyond a
   single metric, so it is noted rather than specified, and it would sit
   naturally in whatever document this one becomes part of.

15.  References

15.1.  Normative References

   [RFC2119]  Bradner, S., "Key words for use in RFCs to Indicate
              Requirement Levels", BCP 14, RFC 2119,
              DOI 10.17487/RFC2119, March 1997,
              <https://www.rfc-editor.org/rfc/rfc2119>.

   [RFC3550]  Schulzrinne, H., Casner, S., Frederick, R., and V.
              Jacobson, "RTP: A Transport Protocol for Real-Time
              Applications", STD 64, RFC 3550, DOI 10.17487/RFC3550,
              July 2003, <https://www.rfc-editor.org/rfc/rfc3550>.

   [RFC3551]  Schulzrinne, H. and S. Casner, "RTP Profile for Audio and
              Video Conferences with Minimal Control", STD 65, RFC 3551,
              DOI 10.17487/RFC3551, July 2003,
              <https://www.rfc-editor.org/rfc/rfc3551>.

   [RFC8174]  Leiba, B., "Ambiguity of Uppercase vs Lowercase in RFC
              2119 Key Words", BCP 14, RFC 8174, DOI 10.17487/RFC8174,
              May 2017, <https://www.rfc-editor.org/rfc/rfc8174>.

Nygate                    Expires 11 March 2027                [Page 29]
Internet-Draft        Mouth-to-Ear Response Latency       September 2026

   [RFC8259]  Bray, T., Ed., "The JavaScript Object Notation (JSON) Data
              Interchange Format", STD 90, RFC 8259,
              DOI 10.17487/RFC8259, December 2017,
              <https://www.rfc-editor.org/rfc/rfc8259>.

15.2.  Informative References

   [G114]     ITU-T, "One-way transmission time", ITU-T Recommendation
              G.114, 2003.

   [G711]     ITU-T, "Pulse code modulation (PCM) of voice frequencies",
              ITU-T Recommendation G.711, 1988.

   [HARNESS]  Nygate, D., "voice-ai-latency-harness: an instrument for
              measuring mouth-to-ear response latency",
              DOI 10.5281/zenodo.22124823, 2026,
              <https://doi.org/10.5281/zenodo.22124823>.

   [RFC2330]  Paxson, V., Almes, G., Mahdavi, J., and M. Mathis,
              "Framework for IP Performance Metrics", RFC 2330,
              DOI 10.17487/RFC2330, May 1998,
              <https://www.rfc-editor.org/rfc/rfc2330>.

   [RFC3611]  Friedman, T., Ed., Caceres, R., Ed., and A. Clark, Ed.,
              "RTP Control Protocol Extended Reports (RTCP XR)",
              RFC 3611, DOI 10.17487/RFC3611, November 2003,
              <https://www.rfc-editor.org/rfc/rfc3611>.

   [RFC6076]  Malas, D. and A. Morton, "Basic Telephony SIP End-to-End
              Performance Metrics", RFC 6076, DOI 10.17487/RFC6076,
              January 2011, <https://www.rfc-editor.org/rfc/rfc6076>.

   [RFC6390]  Clark, A. and B. Claise, "Guidelines for Considering New
              Performance Metric Development", BCP 170, RFC 6390,
              DOI 10.17487/RFC6390, October 2011,
              <https://www.rfc-editor.org/rfc/rfc6390>.

   [RFC7679]  Almes, G., Kalidindi, S., Zekauskas, M., and A. Morton,
              Ed., "A One-Way Delay Metric for IP Performance Metrics
              (IPPM)", STD 81, RFC 7679, DOI 10.17487/RFC7679, January
              2016, <https://www.rfc-editor.org/rfc/rfc7679>.

   [RFC7799]  Morton, A., "Active and Passive Metrics and Methods (with
              Hybrid Types In-Between)", RFC 7799, DOI 10.17487/RFC7799,
              May 2016, <https://www.rfc-editor.org/rfc/rfc7799>.

Nygate                    Expires 11 March 2027                [Page 30]
Internet-Draft        Mouth-to-Ear Response Latency       September 2026

   [RFC8911]  Bagnulo, M., Claise, B., Eardley, P., Morton, A., and A.
              Akhter, "Registry for Performance Metrics", RFC 8911,
              DOI 10.17487/RFC8911, November 2021,
              <https://www.rfc-editor.org/rfc/rfc8911>.

   [RFC8912]  Morton, A., Bagnulo, M., Eardley, P., and K. D'Souza,
              "Initial Performance Metrics Registry Entries", RFC 8912,
              DOI 10.17487/RFC8912, November 2021,
              <https://www.rfc-editor.org/rfc/rfc8912>.

   [TTFAB]    OpenBenchmarks Labs, "Voice agent latency benchmark: time
              to first audio byte measured from real phone calls", 2026,
              <https://openbenchmarks.com/voice-agent-latency>.

Appendix A.  Draft Performance Metrics Registry Entry

      [[EDITOR'S NOTE: filled in against the column structure defined in
      Section 7 of [RFC8911], using the blank template in Section 11 of
      that document, as far as the current text supports.  Incomplete
      deliberately; the gaps are the same gaps flagged in the body.
      Note that Section 7.1.2 imposes a structured naming convention on
      registered metrics, which the Name entry below has yet to
      satisfy.]]

   Identifier:  TBD by IANA

   Name:  TBD, following the naming convention of Section 7.1.2 of
      [RFC8911]

   URI:  TBD

   Description:  The interval between transmission of the final speech
      sample of a caller's utterance and arrival of the first sample of
      a conversational voice system's response, observed at the calling
      endpoint's RTP reference point.

   Change Controller:  IETF

   Version:  1

   Reference Definition:  This document, Section 3.3 through
      Section 3.5.

   Fixed Parameters:  Onset variant thresholds as tabulated in
      Section 3.5; de-jitter anchor rule as specified in Section 3.4.2.

   Reference Method:  This document, Method of Measurement.

Nygate                    Expires 11 March 2027                [Page 31]
Internet-Draft        Mouth-to-Ear Response Latency       September 2026

   Packet Stream Generation:  Prerecorded stimulus transmitted at the
      nominal frame rate of the codec in use.

   Traffic Filter:  The RTP stream of the session under measurement.

   Sampling Distribution:  TBD; see the open item under Measurement
      Timing.

   Runtime Parameters:  Codec, frame period, playout target depth,
      stimulus identifier and hash, SUT configuration identifier.

   Roles:  Calling endpoint; system under test.

   Output Type:  Signed scalar, reported in two variants, with an
      uncertainty term.

   Metric Units:  Milliseconds.

   Calibration:  As specified in Section 4.7.  An implementation reports
      its own calibrated accuracy and the conditions of calibration.

Appendix B.  Reference Implementation

   An open-source implementation of this method exists and is archived
   at [HARNESS].  It implements the metric definition, the onset
   variants, the quality control conditions and the calibration
   procedure described here, and it publishes per-host calibration
   results.

   This appendix is informative.  Conformance is to this document, and
   any disagreement between this document and that implementation is a
   defect in the implementation.

Appendix C.  Open Questions for Review

   This appendix is retained deliberately rather than removed before
   submission.  An individual submission exists to attract review, and
   the author's own list of doubts is a more efficient way to obtain
   useful review than letting each reader rediscover them.

   1.  Filler and non-lexical audio is now specified in Section 3.7 by a
       structural discriminator rather than by content recognition, and
       implemented.  What remains open is whether 150 ms and 2000 ms are
       the right constants.  They were chosen to sit well above the
       pauses inside ordinary speech, measured at around 50 ms on the
       reference implementation's material, and no corpus of real filler
       behaviour informed them yet.  Evidence welcome.

Nygate                    Expires 11 March 2027                [Page 32]
Internet-Draft        Mouth-to-Ear Response Latency       September 2026

   2.  The stimulus annotation rule is now a conformance test against a
       published reference signal rather than a mandated algorithm, and
       the signal exists.  What remains open is where the signal should
       live: carried in this document, registered, or referenced in the
       archived implementation as it is here.

   3.  The interchange record in Section 9.2, specifically whether it
       warrants a registered media type.  The reference_point
       enumeration of Section 4.2 answers the other half of what this
       question used to ask; what remains open there is whether the list
       is the right list, since it was drawn from observed practice
       rather than from a survey.

   4.  Whether Informational is the right category, or whether the
       conformance language argues for Standards Track.

   5.  Venue.  Whether IPPM's charter admits a metric of this kind, and
       if not whether DISPATCH or the Independent Submission stream is
       the better route.  See the editor's note in the Introduction.

   6.  Whether registry registration is appropriate for an application-
       layer metric of this kind.

   7.  Whether Section 7.1 draws the line in the right place.  It
       requires the turn index to be recorded and forbids silent
       pooling, and it deliberately does not prescribe how many turns to
       measure.  An argument that a fixed number belongs in the
       specification, for comparability, is one the author would want to
       hear.

Author's Address

   Daniel Nygate
   Email: dnygate@outlook.com

Nygate                    Expires 11 March 2027                [Page 33]