Mouth-to-Ear Response Latency for Conversational Voice Systems: Metric Definition and Active Measurement Method
draft-nygate-ippm-mrl-00
This document is an Internet-Draft (I-D).
Anyone may submit an I-D to the IETF.
This I-D is not endorsed by the IETF and has no formal standing in the
IETF standards process.
| Document | Type | Active Internet-Draft (individual) | |
|---|---|---|---|
| Author | Daniel Nygate | ||
| Last updated | 2026-09-07 | ||
| RFC stream | (None) | ||
| Intended RFC status | (None) | ||
| Formats | |||
| Stream | Stream state | (No stream defined) | |
| Consensus boilerplate | Unknown | ||
| RFC Editor Note | (None) | ||
| IESG | IESG state | I-D Exists | |
| Telechat date | (None) | ||
| Responsible AD | (None) | ||
| Send notices to | (None) |
draft-nygate-ippm-mrl-00
IP Performance Measurement D. Nygate
Internet-Draft 7 September 2026
Intended status: Informational
Expires: 11 March 2027
Mouth-to-Ear Response Latency for Conversational Voice Systems: Metric
Definition and Active Measurement Method
draft-nygate-ippm-mrl-00
Abstract
This document defines mouth-to-ear response latency (MRL), a
performance metric for conversational voice systems, together with an
active method for measuring it at the RTP reference point of the
calling endpoint. MRL is the interval between the transmission of
the final speech sample of a caller's utterance and the arrival of
the first sample of the system's response audio. Two variants are
defined, one taken at packet arrival and one taken behind a de-jitter
buffer of stated target depth. The method is specified so that both
timestamps are drawn from a single clock on a single host, so that
the metric requires no synchronisation between the measuring endpoint
and the system under test. Requirements for stimulus material,
capture content, quality control, calibration and reporting are
given.
About This Document
This note is to be removed before publishing as an RFC.
Status information for this document may be found at
https://datatracker.ietf.org/doc/draft-nygate-ippm-mrl/.
Discussion of this document takes place on the IP Performance
Measurement Working Group mailing list (mailto:ippm@ietf.org), which
is archived at https://mailarchive.ietf.org/arch/browse/ippm/.
Subscribe at https://www.ietf.org/mailman/listinfo/ippm/.
Source for this draft and an issue tracker can be found at
https://github.com/dnygate/draft-mrl.
Status of This Memo
This Internet-Draft is submitted in full conformance with the
provisions of BCP 78 and BCP 79.
Nygate Expires 11 March 2027 [Page 1]
Internet-Draft Mouth-to-Ear Response Latency September 2026
Internet-Drafts are working documents of the Internet Engineering
Task Force (IETF). Note that other groups may also distribute
working documents as Internet-Drafts. The list of current Internet-
Drafts is at https://datatracker.ietf.org/drafts/current/.
Internet-Drafts are draft documents valid for a maximum of six months
and may be updated, replaced, or obsoleted by other documents at any
time. It is inappropriate to use Internet-Drafts as reference
material or to cite them other than as "work in progress."
This Internet-Draft will expire on 11 March 2027.
Copyright Notice
Copyright (c) 2026 IETF Trust and the persons identified as the
document authors. All rights reserved.
This document is subject to BCP 78 and the IETF Trust's Legal
Provisions Relating to IETF Documents (https://trustee.ietf.org/
license-info) in effect on the date of publication of this document.
Please review these documents carefully, as they describe your rights
and restrictions with respect to this document. Code Components
extracted from this document must include Revised BSD License text as
described in Section 4.e of the Trust Legal Provisions and are
provided without warranty as described in the Revised BSD License.
Table of Contents
1. Introduction . . . . . . . . . . . . . . . . . . . . . . . . 3
1.1. Motivation and decomposition . . . . . . . . . . . . . . 4
1.2. Existing practice . . . . . . . . . . . . . . . . . . . . 5
1.3. Scope . . . . . . . . . . . . . . . . . . . . . . . . . . 6
1.4. Relationship to existing work . . . . . . . . . . . . . . 6
2. Conventions and Definitions . . . . . . . . . . . . . . . . . 7
3. Metric Definition . . . . . . . . . . . . . . . . . . . . . . 8
3.1. Metric name . . . . . . . . . . . . . . . . . . . . . . . 8
3.2. Metric description . . . . . . . . . . . . . . . . . . . 8
3.3. Interval start: t0 . . . . . . . . . . . . . . . . . . . 8
3.4. Interval end: t1 . . . . . . . . . . . . . . . . . . . . 9
3.4.1. Ingress MRL . . . . . . . . . . . . . . . . . . . . . 10
3.4.2. Playout MRL . . . . . . . . . . . . . . . . . . . . . 10
3.4.3. Both variants required . . . . . . . . . . . . . . . 11
3.5. Response onset . . . . . . . . . . . . . . . . . . . . . 11
3.6. Unprompted audio . . . . . . . . . . . . . . . . . . . . 11
3.7. Filler audio and response continuity . . . . . . . . . . 12
4. Method of Measurement . . . . . . . . . . . . . . . . . . . . 13
4.1. Reference point . . . . . . . . . . . . . . . . . . . . . 13
4.2. Observing elsewhere . . . . . . . . . . . . . . . . . . . 13
Nygate Expires 11 March 2027 [Page 2]
Internet-Draft Mouth-to-Ear Response Latency September 2026
4.2.1. Characterising the additional terms . . . . . . . . . 16
4.3. Stimulus requirements . . . . . . . . . . . . . . . . . . 16
4.4. Transmission pacing . . . . . . . . . . . . . . . . . . . 16
4.5. Capture contents . . . . . . . . . . . . . . . . . . . . 17
4.6. Clock requirements . . . . . . . . . . . . . . . . . . . 17
4.7. Sources of error and calibration . . . . . . . . . . . . 17
5. Units of Measurement . . . . . . . . . . . . . . . . . . . . 19
6. Measurement Points and Measurement Domain . . . . . . . . . . 19
7. Measurement Timing . . . . . . . . . . . . . . . . . . . . . 19
7.1. Turns within a session . . . . . . . . . . . . . . . . . 20
8. Quality Control and Result Validity . . . . . . . . . . . . . 20
8.1. Blocking conditions . . . . . . . . . . . . . . . . . . . 20
8.2. Advisory conditions . . . . . . . . . . . . . . . . . . . 21
8.3. Reporting of discards . . . . . . . . . . . . . . . . . . 21
9. Reporting . . . . . . . . . . . . . . . . . . . . . . . . . . 21
9.1. Required fields . . . . . . . . . . . . . . . . . . . . . 21
9.2. Reporting format . . . . . . . . . . . . . . . . . . . . 22
9.2.1. instrument . . . . . . . . . . . . . . . . . . . . . 22
9.2.2. measurement . . . . . . . . . . . . . . . . . . . . . 23
9.2.3. subject . . . . . . . . . . . . . . . . . . . . . . . 23
9.2.4. results . . . . . . . . . . . . . . . . . . . . . . . 23
9.2.5. measurements and captures . . . . . . . . . . . . . . 24
9.2.6. Example . . . . . . . . . . . . . . . . . . . . . . . 25
10. Use and Applications . . . . . . . . . . . . . . . . . . . . 26
11. Applicability and Limitations . . . . . . . . . . . . . . . . 27
12. Security Considerations . . . . . . . . . . . . . . . . . . . 28
13. IANA Considerations . . . . . . . . . . . . . . . . . . . . . 28
14. Future Work . . . . . . . . . . . . . . . . . . . . . . . . . 29
15. References . . . . . . . . . . . . . . . . . . . . . . . . . 29
15.1. Normative References . . . . . . . . . . . . . . . . . . 29
15.2. Informative References . . . . . . . . . . . . . . . . . 30
Appendix A. Draft Performance Metrics Registry Entry . . . . . . 31
Appendix B. Reference Implementation . . . . . . . . . . . . . . 32
Appendix C. Open Questions for Review . . . . . . . . . . . . . 32
Author's Address . . . . . . . . . . . . . . . . . . . . . . . . 33
1. Introduction
Conversational voice systems built from speech recognition, language
modelling and speech synthesis components are now widely deployed
over SIP and RTP. Response latency is a primary determinant of
whether such a system is usable, and it is measured today by several
parties who do not agree on what is being measured. There is no
common definition of the quantity, no agreed reference point at which
to observe it, and no convention for reporting its uncertainty.
Nygate Expires 11 March 2027 [Page 3]
Internet-Draft Mouth-to-Ear Response Latency September 2026
Figures produced under different assumptions are routinely compared
as though they were commensurable, which is the practical problem
this document exists to address. Section 1.1 sets out the terms that
separate one figure from another, and Section 1.2 describes the axes
along which current measurement practice divides.
This document takes no position on the magnitude of any term, on the
accuracy of any published figure, or on the merits of any existing
measurement effort. It defines a quantity and a method of observing
it, so that figures produced by different parties are comparable.
1.1. Motivation and decomposition
For a caller connected to a conversational voice system over SIP and
RTP, the interval between the end of the caller's utterance and the
arrival of response audio comprises, in order of occurrence:
+=============================================================+
| Term |
+=============================================================+
| de-jitter buffer depth on the inbound leg |
+-------------------------------------------------------------+
| voice activity detection and endpointing decision |
+-------------------------------------------------------------+
| speech recognition finalisation after the endpoint decision |
+-------------------------------------------------------------+
| orchestration, retrieval and any tool invocation |
+-------------------------------------------------------------+
| language model time to first token |
+-------------------------------------------------------------+
| speech synthesis time to first audio |
+-------------------------------------------------------------+
| encoding, packetisation and media relay |
+-------------------------------------------------------------+
| de-jitter buffer depth on the outbound leg |
+-------------------------------------------------------------+
Table 1
A figure covering one or two of these terms does not predict the
interval a caller experiences, and the terms are not separable by
observation at the caller. MRL is defined here as the aggregate,
observed at a single reference point, precisely because the aggregate
is what an external observer can measure without instrumenting the
system under test.
Nygate Expires 11 March 2027 [Page 4]
Internet-Draft Mouth-to-Ear Response Latency September 2026
1.2. Existing practice
Two kinds of figure are published today.
Component-level figures are reported by the operators of individual
components, most commonly the interval from a request reaching a
speech synthesis endpoint to the first audio byte returned, and less
commonly the interval to a language model's first token. These are
observed at an API boundary internal to the system and cover a subset
of the terms above.
Caller-side figures are produced by placing a call and observing the
response. At least one such effort publishes both its results and
its tooling openly [TTFAB], reporting the interval between the
caller's speech end and the onset of the response as observed in a
recording of the call rather than in any timestamp reported by the
system.
Caller-side efforts differ from one another along axes that render
their outputs incomparable, and those axes are the reason this
document exists:
* the reference point at which the response is observed, which may
be an RTP endpoint, a recording made by a carrier or a
conferencing bridge, or an endpoint's audio device, each placing a
different and frequently uncharacterised quantity of transport and
buffering inside the measured interval. Section 4.2 enumerates
these so that a party observing anywhere on that list can produce
a conforming report which declares the difference, and
Section 4.2.1 gives a way to measure it rather than disclose it;
* the determination of the caller's speech end, which may be derived
from a voice activity detector applied to the transmitted audio,
or from prior annotation of known stimulus material;
* the treatment of buffering, since an interval taken at packet
arrival and an interval taken after a de-jitter buffer differ by
the depth of that buffer;
* the definition of response onset, which for audio that ramps in
gradually is a choice of threshold rather than an observation;
* whether the measuring instrument has itself been calibrated
against a known delay, and therefore whether a reported figure is
a point estimate or an upper bound.
Nygate Expires 11 March 2027 [Page 5]
Internet-Draft Mouth-to-Ear Response Latency September 2026
An effort that states its choices along these axes produces figures a
reader can interpret, whereas two efforts that have chosen
differently produce figures which cannot be placed side by side even
when each is internally correct. This document specifies one set of
choices and requires that they be reported alongside any figure
derived under them.
1.3. Scope
This document specifies:
* the definition of the MRL metric and its two required variants;
* the method of measurement, including the reference point, the
determination of the interval endpoints, and the required contents
of a capture;
* the conditions under which a measurement MUST be treated as
invalid;
* the fields that MUST accompany a reported figure.
This document does not specify endpointing strategy, barge-in
handling, speech recognition or synthesis behaviour, or any property
of the system under test. It does not specify a signalling protocol;
the method assumes an established RTP session and is independent of
how that session was established.
1.4. Relationship to existing work
The metric defined here is an application-layer performance metric in
the sense of [RFC6390], and this document follows the template in
Section 5.4 of that document. [RFC6076] defines end-to-end
performance metrics for telephony sessions at the SIP layer and is
the closest existing IETF work in subject matter; the metric defined
here concerns the media plane rather than signalling, and is
complementary. No existing IETF document defines the quantity
described in Section 1.1 or specifies a reference point at which to
observe it.
The framework of [RFC2330] and the delay metric of [RFC7679] inform
the treatment of error and uncertainty. The method described here is
an active method in the taxonomy of [RFC7799], since it generates the
stimulus whose response it measures.
Nygate Expires 11 March 2027 [Page 6]
Internet-Draft Mouth-to-Ear Response Latency September 2026
[RFC3611] and its extensions define a reporting mechanism by which
endpoints convey media quality metrics in band. Conveying MRL in
that manner is out of scope for this document and is noted in
Section 14 as possible follow-on work.
[[EDITOR'S NOTE, on venue. The IPPM charter bounds the group's
work to "metrics and methodologies which are applicable over
transport-layer protocols over IP", while also covering
"applications running over transport layer protocols". Whether a
metric whose dominant terms are speech recognition, inference and
synthesis falls inside that boundary is a fair question and should
be settled before effort is spent on -01.
The precedent in favour is draft-ietf-ippm-responsiveness, an
adopted IPPM work item targeting Proposed Standard, which measures
at the application layer over HTTP/2 and HTTP/3 and justifies
itself explicitly on user experience. The argument against is
that responsiveness remains a property of the network under load,
whereas most of MRL is not attributable to the network at all.
If IPPM declines, the alternatives are to take the question to
DISPATCH, which exists for work with no obvious home, or to pursue
publication through the Independent Submission stream. The
document is written to stand in any of the three.
Separately, decide whether to pursue registration in the
Performance Metrics Registry ([RFC8911], initially populated by
[RFC8912]); Appendix A holds a draft entry. The registry has to
date been populated with IP-layer path metrics, so eligibility
should be confirmed before the appendix is presented as a
proposal.]]
2. Conventions and Definitions
The key words "MUST", "MUST NOT", "REQUIRED", "SHALL", "SHALL NOT",
"SHOULD", "SHOULD NOT", "RECOMMENDED", "NOT RECOMMENDED", "MAY", and
"OPTIONAL" in this document are to be interpreted as described in
BCP 14 [RFC2119] [RFC8174] when, and only when, they appear in all
capitals, as shown here.
The following terms are used throughout.
Calling endpoint: The endpoint that generates the stimulus utterance
and observes the response. All measurement is performed here.
System under test (SUT): The conversational voice system that
receives the stimulus and generates a response. The method treats
the SUT as opaque.
Nygate Expires 11 March 2027 [Page 7]
Internet-Draft Mouth-to-Ear Response Latency September 2026
Reference point: The point at which timestamps are taken. See
Section 4.1.
Stimulus: A prerecorded utterance transmitted by the calling
endpoint to elicit a response.
Speech end: The final sample of speech in the stimulus, determined
offline. See Section 3.3.
Response onset: The first sample of the SUT's response audio,
determined offline. See Section 3.5.
3. Metric Definition
3.1. Metric name
Mouth-to-Ear Response Latency (MRL), reported in two variants named
Ingress MRL and Playout MRL.
3.2. Metric description
MRL is the interval between the instant at which the calling endpoint
transmits the final speech sample of a stimulus utterance and the
instant at which the first sample of the SUT's response reaches the
calling endpoint.
MRL is defined as:
MRL = t1 - t0
where t0 and t1 are as defined in Section 3.3 and Section 3.4.
MRL is signed. A negative value indicates that response audio
reached the calling endpoint before the stimulus had finished being
transmitted, which occurs with aggressive endpointing and with
backchannel responses. Implementations MUST report negative values
as measured and MUST NOT clamp them to zero.
3.3. Interval start: t0
t0 is the instant at which the final speech sample of the stimulus is
transmitted by the calling endpoint at the reference point.
t0 MUST be determined by:
1. annotating the stimulus offline to sample precision to locate the
speech end;
Nygate Expires 11 March 2027 [Page 8]
Internet-Draft Mouth-to-Ear Response Latency September 2026
2. mapping that sample onto the RTP packet that carried it, using
RTP timestamps as defined in [RFC3550];
3. interpolating within that packet at the sample rate to obtain the
sample's offset from the packet's transmission instant.
t0 is therefore the transmission instant of the carrying packet plus
the target sample's offset within that packet. The sign of the
offset term is significant; see Section 4.7.
t0 MUST NOT be derived from voice activity detection performed at run
time. A run-time detector has a decision lag of its own, that lag is
a term within the quantity being measured, and using it to define t0
would conceal the term.
Annotation of the stimulus MUST NOT extend the speech end by a fixed
constant. A fixed extension displaces t0 later by that constant and,
since MRL is t1 minus t0, reduces every reported figure by the same
constant. See Section 4.7.
An implementation MUST NOT be required to use any particular
annotation algorithm, and MUST be able to demonstrate that its
annotation agrees with a published reference. The reference
implementation uses decay-following hysteresis with a short sliding
RMS refinement, and publishes a frozen stimulus signal whose speech
end is exact by construction, together with the boundary its own
annotator locates [HARNESS]. An implementation claiming conformance
SHOULD reproduce that boundary and MUST state the deviation where it
does not.
Specifying the conformance test rather than the algorithm is
deliberate. The quantity that matters to a reader is whether two
parties locate the same boundary in the same audio, and an algorithm
mandated in prose can be implemented differently by two careful
people while a boundary in a published waveform cannot be argued
about.
[[EDITOR'S NOTE: the reference signal is frozen bytes rather than
a generator seed, because a signal rebuilt by its generator
changes when the generator does. Whether the signal itself should
be carried in an IANA registry, an appendix, or by reference to
the archived implementation is a question for the working group;
reference is assumed here.]]
3.4. Interval end: t1
t1 is the instant at which the first sample of the SUT's response
reaches the calling endpoint, taken at the reference point.
Nygate Expires 11 March 2027 [Page 9]
Internet-Draft Mouth-to-Ear Response Latency September 2026
t1 MUST be determined by reassembling the received stream in RTP
timestamp order, locating the response onset as specified in
Section 3.5, and mapping the onset sample back through the packet
that carried it.
Two variants of t1 are defined, and both MUST be reported.
3.4.1. Ingress MRL
Ingress MRL takes t1 at the arrival instant of the onset sample, that
is, at the reference point with no buffering applied.
Ingress MRL isolates the contribution of the SUT and of the network
path from the buffering policy of the calling endpoint. Its variance
under a jittered path tracks the jitter of that path, because the
metric reports the path as it finds it, and an implementation showing
less variance than the path exhibits is smoothing a quantity it was
asked to observe.
3.4.2. Playout MRL
Playout MRL takes t1 at the instant at which the onset sample would
be released from a de-jitter buffer of stated target depth.
The target depth MUST be reported with the figure. The de-jitter
model used MUST anchor on the minimum transit delay observed over a
stated initial window, which is the behaviour adaptive buffers
converge toward, rather than on the arrival instant of the first
packet received.
Playout MRL corresponds to the interval a caller waits and is the
variant to compare against conversational turn-taking norms.
Playout MRL is distinct from the one-way transmission delay addressed
by [G114], and the two are not interchangeable. One-way transmission
delay concerns the time taken to carry audio across a path and
applies to a conversation between two people, whereas MRL concerns
the time a system takes to begin responding and includes transmission
delay as one term among several. A system may satisfy the
transmission delay guidance in [G114] on both legs and still exhibit
an MRL an order of magnitude larger.
Nygate Expires 11 March 2027 [Page 10]
Internet-Draft Mouth-to-Ear Response Latency September 2026
3.4.3. Both variants required
An implementation MUST report both variants. Ingress MRL on its own
understates the interval a caller experiences, while a bare Playout
MRL leaves the SUT confounded with whatever buffering policy the
calling endpoint happened to apply, so neither figure can be
interpreted without the other.
3.5. Response onset
The instant at which audio begins has no unique definition, because
synthesised speech commonly ramps in over tens of milliseconds rather
than beginning at full level. Onset is therefore computed under
three named variants, and the dispersion across them is a component
of the reported uncertainty.
+===========+===================+================+===============+
| Variant | Above noise floor | Absolute floor | Sustained for |
+===========+===================+================+===============+
| sensitive | 6 dB | -55 dBov | 10 ms |
+-----------+-------------------+----------------+---------------+
| headline | 10 dB | -50 dBov | 20 ms |
+-----------+-------------------+----------------+---------------+
| strict | 12 dB | -45 dBov | 30 ms |
+-----------+-------------------+----------------+---------------+
Table 2
Levels are expressed in dBov, referenced to full-scale RMS.
The headline variant is the reported figure. The spread of MRL
across all three variants is the onset-definition uncertainty of the
measurement and MUST be published alongside any headline figure. On
an abrupt onset the spread is small; on a gradual ramp it grows with
the ramp duration and becomes the dominant uncertainty term.
Sub-frame refinement of the onset instant MUST use a short sliding
RMS. Instantaneous sample magnitude crosses any fixed threshold on
isolated noise peaks at a rate high enough to displace the boundary
materially; see Section 4.7.
3.6. Unprompted audio
A conversational voice system commonly speaks before the caller does,
opening with a greeting that arrives within a few hundred
milliseconds of the session being established. That audio responds
to nothing, because the caller has not yet spoken.
Nygate Expires 11 March 2027 [Page 11]
Internet-Draft Mouth-to-Ear Response Latency September 2026
Response-onset detection MUST NOT begin before the end of any such
unprompted audio. The point at which it ended MUST be recorded in
the capture, and the noise-floor estimate required by Section 3.5
MUST be taken from audio following that point.
Detection across the whole received stream locates the greeting
instead of the response. Since the greeting precedes t0, the
resulting MRL is large and negative, and because this document
licenses negative values as genuine behaviour there is no bound
against which such an error announces itself. In the reference
implementation, a capture carrying an 800 ms greeting with a true MRL
of 900 ms yielded -2085 ms and satisfied every condition in
Section 8.
A greeting falling inside the noise-floor window is speech rather
than channel noise, so it raises the estimate and displaces every
onset threshold derived from it. On the same capture the floor moved
by 1.35 dB.
A calling endpoint SHOULD NOT begin transmitting its stimulus while
unprompted audio is still in progress. A system that implements
barge-in detection will stop speaking when it hears the caller, which
truncates the greeting and alters the interaction under measurement.
The interval from session establishment to the onset of unprompted
audio is a distinct and useful quantity, since a caller who hears
nothing for several seconds after the line opens is poorly served
whatever the system's MRL turns out to be. It shares no terms with
MRL and is not defined here; see Section 14.
3.7. Filler audio and response continuity
t1 as defined in Section 3.4 is the onset of the system's first
response audio, whatever that audio happens to contain. A system
that emits an earcon, a breath or a filled pause while its response
is still being generated therefore records a low MRL while conveying
nothing during that interval, and would rank above a system that
stayed quiet and then answered.
This document does not resolve that by identifying which audio
carries meaning. A metric incorporating a judgement about meaning
cannot be re-derived from a published capture by an independent
reviewer, and that reproducibility is the property which makes the
rest of this specification worth having. The discriminator is
structural instead: filler is followed by silence before the
substantive response begins, and continuous speech is not.
Nygate Expires 11 March 2027 [Page 12]
Internet-Draft Mouth-to-Ear Response Latency September 2026
An implementation MUST, within a window of 2000 ms following t1,
measure the longest interval whose level lies below the onset
threshold of the headline variant. Intervals carried by no packet
MUST be excluded from that measurement, because a lost or late-
discarded frame leaves a gap indistinguishable from a deliberate
pause and would otherwise allow a degraded path to manufacture
filler.
Where that interval exceeds 150 ms:
* the response MUST be reported as discontiguous;
* a second onset MUST be reported, at the start of the final
contiguous segment, together with the MRL derived from it.
MRL itself is unchanged by this section, so no figure measured under
an earlier revision becomes invalid. A reader receives both onsets
and can see whether they differ and by how much.
The 2000 ms window and the 150 ms threshold are fixed by this
document rather than left to the implementation, for the reason given
in Section 3.5: figures derived under different parameters cannot be
compared even when each is internally correct. Both MUST be reported
alongside any figure derived under them.
4. Method of Measurement
4.1. Reference point
All timestamps MUST be taken at the RTP egress and ingress reference
point of the calling endpoint, that is, immediately before a packet
is passed to the operating system for transmission and immediately
after a packet is received from it.
The reference point is chosen so that the measurement includes every
term a caller experiences downstream of the calling endpoint's own
send path, and excludes the calling endpoint's own playout hardware,
which is a property of the observer rather than of the SUT.
4.2. Observing elsewhere
Not every party measuring this quantity controls an RTP endpoint. A
figure derived from a carrier's call recording is a real measurement
of something a caller experienced, and declaring it as such is more
useful than being unable to describe it at all.
Nygate Expires 11 March 2027 [Page 13]
Internet-Draft Mouth-to-Ear Response Latency September 2026
An implementation MAY therefore observe at a point other than the one
in Section 4.1, provided it declares which, and provided it discloses
what that choice places inside the measured interval. Such a report
conforms to this document. What it does not do is produce a figure
comparable with one taken at the RTP reference point, because the two
intervals contain different terms.
Nygate Expires 11 March 2027 [Page 14]
Internet-Draft Mouth-to-Ear Response Latency September 2026
+=======================+===================+=======================+
| reference_point | observed at | additional terms |
| | | inside the interval |
+=======================+===================+=======================+
| rtp-endpoint | immediately | none; this is |
| | before the | Section 4.1 |
| | packet is passed | |
| | to the operating | |
| | system, and | |
| | immediately | |
| | after it is | |
| | received from it | |
+-----------------------+-------------------+-----------------------+
| host-packet-capture | a packet capture | the host's own |
| | facility on the | network stack, and |
| | calling host | asymmetrically: on |
| | | transmission the |
| | | capture is taken |
| | | after the send path, |
| | | on reception before |
| | | the application reads |
+-----------------------+-------------------+-----------------------+
| bridge-recording | a session border | the leg between |
| | controller or | caller and bridge in |
| | conferencing | both directions, and |
| | bridge in the | the bridge's own |
| | path | media handling |
+-----------------------+-------------------+-----------------------+
| carrier-recording | a recording made | the leg to the |
| | by a carrier or | carrier in both |
| | communications | directions, any |
| | platform | transcoding it |
| | | performs, its |
| | | buffering, and its |
| | | recording pipeline |
+-----------------------+-------------------+-----------------------+
| endpoint-audio-device | the caller's | the endpoint's de- |
| | audio device, or | jitter buffer, its |
| | a loopback | audio stack and the |
| | capture of it | device's own latency |
+-----------------------+-------------------+-----------------------+
Table 3
Nygate Expires 11 March 2027 [Page 15]
Internet-Draft Mouth-to-Ear Response Latency September 2026
Where a value other than rtp-endpoint is declared, the additional
terms MUST be disclosed and the figure MUST be reported as an upper
bound rather than a point estimate, unless those terms have been
characterised as below. Figures taken at different reference points
MUST NOT be pooled into one distribution.
4.2.1. Characterising the additional terms
An upper bound is a weaker result than it needs to be, and the
additional terms are measurable rather than merely acknowledgeable.
Where the path to the system under test traverses infrastructure the
measuring party does not control, an implementation SHOULD place a
second call over the same route to a responder replying at a delay it
has programmed, and subtract that measurement from the first. The
infrastructure's contribution is common to both and cancels, leaving
the system's response. Two routes over the same carrier are not
identical, so this bounds the additional terms rather than
eliminating them exactly, and the bound is what should be reported.
This is the same differential construction the calibration in
Section 4.7 uses, applied to a term outside the instrument instead of
inside it. It turns a disclosed caveat into a measured quantity, and
it is available to any party that can dial a number it controls.
4.3. Stimulus requirements
Stimulus material MUST be prerecorded and MUST be transmitted at the
nominal frame rate of the codec in use. The stimulus MUST be hashed
and the hash MUST be recorded with the capture, so that a changed
stimulus invalidates a comparison loudly rather than silently.
The method is independent of the codec, and the codec in use MUST be
reported with any figure derived under it. Where a payload format
from [RFC3551] is used, the sample rate and frame period follow that
profile, and the companding of [G711] applies to the PCMU and PCMA
formats.
4.4. Transmission pacing
The calling endpoint MUST measure the deviation of its own
transmission instants from the nominal frame grid and MUST record the
worst deviation observed during the call. An endpoint that cannot
pace its own transmission has an unreliable t0 and therefore an
unreliable MRL. See Section 8.
Nygate Expires 11 March 2027 [Page 16]
Internet-Draft Mouth-to-Ear Response Latency September 2026
4.5. Capture contents
A capture MUST contain raw payloads and raw timestamps only. A
capture MUST NOT contain any derived quantity, including any latency
figure.
This requirement exists so that a revised onset definition, or a
definition proposed by a reviewer, can be applied to existing data
without repeating a collection. The uncertainty analysis in
Section 3.5 is not possible otherwise.
4.6. Clock requirements
Both t0 and t1 MUST be taken from a single monotonic clock on the
calling endpoint.
Because the interval is the difference of two timestamps drawn from
one clock on one host, the metric requires no synchronisation between
the calling endpoint and the SUT, and is insensitive to offset
between them. This is a deliberate property of the definition and
distinguishes the method from one-way delay measurement, where clock
synchronisation between the two hosts dominates the error budget as
described in Section 3.7.1 of [RFC7679].
A wall-clock timestamp MAY be recorded in parallel for the sole
purpose of correlating captures with traces obtained from the SUT.
Any offset or skew in that clock affects such correlation only and
MUST NOT affect the reported MRL.
4.7. Sources of error and calibration
An implementation MUST state its calibrated accuracy, and MUST state
the conditions under which that calibration was obtained.
Calibration is performed by replacing the SUT with a reference
responder that replies at a programmed delay, so that ground truth is
known exactly. Under each channel condition the offset from ground
truth that the physics of the channel requires is predictable in
advance: for Ingress MRL it is the base transit delay plus the mean
jitter excess, and for Playout MRL it is the base transit delay plus
the buffer target depth. The calibration criterion is therefore the
residual after subtracting that predicted offset, rather than the raw
difference from ground truth.
Nygate Expires 11 March 2027 [Page 17]
Internet-Draft Mouth-to-Ear Response Latency September 2026
Calibration MUST NOT be performed exclusively at programmed delays
that are integer multiples of the frame period. A responder that
evaluates its emission deadline once per received frame quantises its
own output to the frame period, and that error is invisible at frame-
commensurate delays because the deadline then falls on a frame
boundary.
The following error mechanisms are known to produce plausible-looking
but incorrect figures, and an implementation is advised to test for
each:
Fixed annotation extension: Extending the stimulus speech end by a
constant biases every figure low by that constant.
Instantaneous-magnitude thresholding, at either boundary: Sub-frame
refinement on sample magnitude rather than a short sliding RMS
lets isolated noise peaks cross the threshold. At the stimulus
end boundary this displaces t0 late; at the response onset it
displaces t1 early, biasing MRL low by up to one analysis window.
The second was masked in the reference implementation for as long
as its calibration responder placed responses on the analysis
grid, and surfaced only once a media-clock responder placed them
where they began: it accounted for the whole of a -0.40 ms
headline bias and a 0.99 ms spread between onset variants on a
hard onset. A fix applied at one boundary has to be checked
against its mirror.
Frame-quantised reference responder: Adds a uniform error between
zero and one frame period, invisible at frame-commensurate
calibration delays.
Frame-offset sign error: Using the wrong sign for the within-frame
offset term of t0 produces a residual of twice the offset with the
opposite sign. A large bias accompanied by a tight spread is the
signature of a definitional or arithmetic error rather than of
host timing noise, and implementations are advised to report that
discrimination automatically.
Symmetric jitter modelling: A simulated channel that models network
jitter as zero-mean Gaussian permits a minimum-tracking de-jitter
anchor to sit earlier than the minimum transit delay allows, which
appears as a negative bias in Playout MRL. Network jitter is one-
sided and has a hard floor at the minimum transit delay. Sender
pacing deviation is a different mechanism and is symmetric,
because a timer-driven sender can fire either early or late.
Calibration source that is not an honest RTP sender: A reference
Nygate Expires 11 March 2027 [Page 18]
Internet-Draft Mouth-to-Ear Response Latency September 2026
responder whose RTP timestamps count frames rather than follow a
media clock, or whose idle-stream pacing depends on whether it is
receiving anything, produces playout figures with errors that
ingress cannot see, since ingress is derived from arrival and
playout through the timestamp. Observed as a 44.7 ms transit slip
that flagged late discard on jitter-free calls, and 9.54 ms of
playout spread that no buffer target reduced. Both surfaced only
from kept captures over a real path.
[[EDITOR'S NOTE: the reference implementation's calibrated figures
are published in [HARNESS] and are deliberately not reproduced
here. A metric specification that carries one implementation's
results becomes stale and invites the reader to treat those
figures as a conformance target. Confirm this is the right call
in review.]]
5. Units of Measurement
MRL is reported in milliseconds, signed, to a resolution of not
coarser than 0.1 ms.
Uncertainty terms accompanying the figure are reported in the same
units.
6. Measurement Points and Measurement Domain
The single measurement point is the calling endpoint's RTP reference
point as defined in Section 4.1. No observation of the SUT's
internal state is required, and none is assumed to be available.
The measurement domain is the path between the calling endpoint and
the SUT together with the SUT itself. The method does not separate
the two, and reported figures are therefore properties of the pairing
rather than of the SUT alone. Where separation is required, the path
contribution has to be characterised independently and reported
alongside.
7. Measurement Timing
Each measurement corresponds to one stimulus and one response. A
reported distribution MUST state the number of measurements attempted
and the number discarded, as required by Section 8.
Nygate Expires 11 March 2027 [Page 19]
Internet-Draft Mouth-to-Ear Response Latency September 2026
7.1. Turns within a session
Several exchanges within one session are permitted and are the
realistic case, since a caller does not hang up after one question.
They are not interchangeable, though, and a report that pools them
without saying so is not comparable with one that does not.
A system's first response in a session carries terms that later
responses do not: session establishment, whatever a model does on a
cold start, caches not yet populated, and connections not yet open.
Later responses may in turn benefit from conversational state the
system has accumulated. Both effects are real properties of the
system rather than measurement artefacts, and both are invisible in
an aggregate that does not record which turn each measurement came
from.
Therefore:
* each measurement MUST record its turn index within the session,
counting from one;
* a reported distribution MUST state which turn indices it covers;
* first-turn measurements MUST NOT be pooled with later ones unless
the report states that they are pooled and gives the counts of
each.
This specifies what must be reported rather than how many turns to
measure. A study of cold-start behaviour measuring only first turns,
and a study of steady-state behaviour discarding them, are both valid
and are answering different questions; what makes them comparable to
each other is that both said which they did.
A calling endpoint conducting a multi-turn session MUST NOT transmit
its next stimulus while the system is still speaking, for the reason
given in Section 3.6: a system implementing barge-in detection will
stop, which alters the interaction under measurement. Locating the
end of the system's turn is the same problem as locating the end of a
greeting and admits the same solution.
8. Quality Control and Result Validity
Two classes of condition are defined.
8.1. Blocking conditions
A measurement exhibiting any of the following MUST be discarded and
MUST NOT be reported as a figure:
Nygate Expires 11 March 2027 [Page 20]
Internet-Draft Mouth-to-Ear Response Latency September 2026
* no speech end could be located in the stimulus;
* no packets were received from the SUT;
* no response onset was found;
* the response onset fell within the first received frame, so that
the true onset may precede the observation window;
* the worst transmission pacing deviation exceeded a stated
threshold.
The pacing threshold in the reference implementation is 5 ms. A
calling endpoint that cannot pace its own transmission within that
bound has an unreliable t0.
Discarding a measurement under these conditions is correct behaviour.
An implementation MUST NOT relax a blocking condition in order to
retain a measurement.
8.2. Advisory conditions
The following conditions do not invalidate a measurement but MUST be
reported with it:
* packet loss above a stated threshold;
* late discard above a stated threshold.
A call over a lossy path remains a valid measurement of a lossy path.
Where the frame carrying the response onset is itself lost, onset
detection is deferred by whole frames and the resulting figure is an
upper bound rather than a point estimate, which is why the flag has
to travel with the number.
8.3. Reporting of discards
Every reported distribution MUST be accompanied by the count of
measurements discarded. A run that discards a large fraction of its
measurements is not comparable with one that discards none,
irrespective of the percentiles of the survivors.
9. Reporting
9.1. Required fields
A reported MRL figure MUST be accompanied by:
Nygate Expires 11 March 2027 [Page 21]
Internet-Draft Mouth-to-Ear Response Latency September 2026
* Ingress MRL and Playout MRL, both signed;
* the playout target depth;
* the onset variant used for the headline figure, and the spread
across all three variants;
* the codec and frame period;
* the number of measurements attempted and the number discarded;
* any advisory flags raised;
* the calibrated accuracy of the measuring implementation and the
conditions of that calibration;
* an identifier for the SUT configuration sufficient to establish
that two figures refer to the same configuration, without
necessarily disclosing that configuration.
9.2. Reporting format
A reported result MUST be expressible as a JSON object [RFC8259]
carrying the members defined below. The format exists so that a
reader can establish, without contacting the party who produced a
figure, whether two figures were derived under the same choices.
The schema member MUST be present and MUST be the string mrl-report/1
for reports conforming to this document.
9.2.1. instrument
Describes the measuring implementation and its calibration. All
members are REQUIRED.
name, version: Identify the implementation that produced the report.
calibration: An object recording the outcome of the procedure in
Section 4.7. It MUST carry conditions, a human-readable statement
of the channel and host conditions under which calibration was
performed, and for each of ingress and playout a bias_ms and a
p95_abs_error_ms. It SHOULD carry reference, a URI or DOI at
which the calibration evidence can be inspected. An
implementation that has not been calibrated MUST set calibration
to null rather than omitting it, and every figure in such a report
is an upper bound rather than a point estimate.
Nygate Expires 11 March 2027 [Page 22]
Internet-Draft Mouth-to-Ear Response Latency September 2026
9.2.2. measurement
Describes the choices along the axes in Section 1.2. All members are
REQUIRED.
reference_point: Where t1 was observed, taking one of the values
enumerated in Section 4.2. Any value other than rtp-endpoint MUST
be accompanied by reference_point_notes stating what that choice
places inside the measured interval, and by
reference_point_bound_ms where the additional terms have been
characterised as in Section 4.2.1, or null where they have not and
the figures are therefore upper bounds.
codec, frame_period_ms, sample_rate_hz: The media parameters in
force.
playout_target_ms: The de-jitter buffer target depth used to derive
Playout MRL.
onset_variants: An array of objects, each carrying name, margin_db,
absolute_dbov and sustain_ms. The parameters MUST be stated
rather than referenced by name alone, so that a report remains
interpretable if the defaults in Section 3.5 are ever revised.
headline_variant: The name of the variant whose figures are quoted
as the headline.
continuity_window_ms, continuity_gap_threshold_ms: The parameters of
Section 3.7, stated rather than assumed so that a report remains
interpretable if the defaults are ever revised.
9.2.3. subject
Identifies what was measured. stimulus_id and stimulus_sha256 are
REQUIRED. sut_identifier is REQUIRED and sut_config_sha256 is
RECOMMENDED, the latter allowing two reports to be shown to concern
the same configuration without that configuration being disclosed.
9.2.4. results
Aggregate figures. All members are REQUIRED.
n_attempted, n_reported, n_discarded: Counts of measurements.
n_attempted MUST equal n_reported plus n_discarded.
turn_indices: The turn indices this distribution covers, and the
Nygate Expires 11 March 2027 [Page 23]
Internet-Draft Mouth-to-Ear Response Latency September 2026
count of measurements from each. REQUIRED by Section 7.1, which
forbids pooling a first turn with later ones without saying so. A
single-turn study reports one entry, which is the honest way to
say that the figures describe cold starts.
discard_reasons: An object mapping each blocking condition in
Section 8 to the number of measurements it discarded. The counts
MUST sum to n_discarded.
ingress, playout: Objects carrying at least mean_ms, p50_ms, p95_ms
and max_ms, computed over the reported measurements under the
headline variant. Values are signed.
onset_definition_uncertainty_ms: The spread of MRL across all
variants in onset_variants, carrying at least p50 and max. This
member MUST be present, since a headline figure quoted without it
is incomplete under Section 3.5.
continuity: An object carrying gap_p50_ms and gap_max_ms over the
reported measurements, n_discontiguous, and a contiguous object of
the same shape as ingress giving the MRL to the start of
uninterrupted speech. Where n_discontiguous is zero the
contiguous figures equal the ingress figures, which is the
expected case and is reported rather than omitted so that its
absence never has to be inferred.
advisory_flags: An object mapping each advisory condition in
Section 8 to the number of reported measurements carrying it.
9.2.5. measurements and captures
measurements SHOULD carry one object per individual measurement, each
with the measurement's identifier, its session identifier and turn
index as required by Section 7.1, its t0 and t1 in nanoseconds on the
instrument's monotonic clock, its ingress and playout MRL under every
variant, the continuity gap and contiguous onset from Section 3.7,
the end of any unprompted audio as required by Section 3.6, and any
flags raised. captures SHOULD carry one object per capture with the
measurement identifier, a sha256 of the capture file, and a URI at
which it can be obtained.
Both members are optional because a party may be unable to publish
raw material. Omitting them removes the reader's ability to re-
derive the figures under a different onset definition, which
Section 3.5 identifies as the dominant uncertainty term, so a report
omitting them is weaker evidence than one that includes them.
Nygate Expires 11 March 2027 [Page 24]
Internet-Draft Mouth-to-Ear Response Latency September 2026
9.2.6. Example
The following is a report with the per-measurement and capture arrays
elided.
{
"schema": "mrl-report/1",
"instrument": {
"name": "voice-ai-latency-harness",
"version": "0.1.1",
"calibration": {
"conditions": "clean channel, 20 ms grid, PCMU",
"ingress": { "bias_ms": -0.40, "p95_abs_error_ms": 2.38 },
"playout": { "bias_ms": -0.40, "p95_abs_error_ms": 2.38 },
"reference": "https://doi.org/10.5281/zenodo.22124823"
}
},
"measurement": {
"reference_point": "rtp-endpoint",
"reference_point_notes": null,
"reference_point_bound_ms": null,
"codec": "PCMU",
"frame_period_ms": 20.0,
"sample_rate_hz": 8000,
"playout_target_ms": 40.0,
"onset_variants": [
{ "name": "sensitive", "margin_db": 6.0,
"absolute_dbov": -55.0, "sustain_ms": 10.0 },
{ "name": "headline", "margin_db": 10.0,
"absolute_dbov": -50.0, "sustain_ms": 20.0 },
{ "name": "strict", "margin_db": 12.0,
"absolute_dbov": -45.0, "sustain_ms": 30.0 }
],
"headline_variant": "headline",
"continuity_window_ms": 2000.0,
"continuity_gap_threshold_ms": 150.0
},
"subject": {
"sut_identifier": "system-A",
"sut_config_sha256": "9f2b...c41e",
"stimulus_id": "eval-set-1/utt-017",
"stimulus_sha256": "3ad1...77b0"
},
"results": {
"n_attempted": 20,
"n_reported": 18,
"n_discarded": 2,
"turn_indices": { "1": 8, "2": 5, "3": 5 },
Nygate Expires 11 March 2027 [Page 25]
Internet-Draft Mouth-to-Ear Response Latency September 2026
"discard_reasons": {
"onset_not_found": 1, "tx_pacing_deviation": 1
},
"ingress": {
"mean_ms": 812.4, "p50_ms": 796.0,
"p95_ms": 1043.2, "max_ms": 1101.7
},
"playout": {
"mean_ms": 852.4, "p50_ms": 836.0,
"p95_ms": 1083.2, "max_ms": 1141.7
},
"onset_definition_uncertainty_ms": { "p50": 4.7, "max": 9.8 },
"continuity": {
"gap_p50_ms": 48.0,
"gap_max_ms": 512.0,
"n_discontiguous": 4,
"contiguous": {
"mean_ms": 941.7, "p50_ms": 802.0,
"p95_ms": 1418.6, "max_ms": 1461.0
}
},
"advisory_flags": {
"high_loss": 1,
"high_late_discard": 0,
"discontiguous_response": 4
}
}
}
[[EDITOR'S NOTE: this structure is deliberately close to what the
reference implementation already emits, which keeps at least one
producer honest, and it is the part of the document most likely to
change on review. Two questions in particular. Whether the
format should be registered as a media type. And whether
reference_point should be an enumeration with values for carrier-
side and bridge-side recording, so that efforts observing
elsewhere can produce conforming reports that declare the
difference, rather than being unable to conform at all. The
second would widen adoption considerably and is probably worth
doing.]]
10. Use and Applications
The metric supports:
* comparison of conversational voice systems from the position of a
caller;
Nygate Expires 11 March 2027 [Page 26]
Internet-Draft Mouth-to-Ear Response Latency September 2026
* characterisation of a single system across configurations or load
levels;
* regression testing of a deployed system over time.
MRL measures an interval and carries no information about what the
response contained. Response quality, recognition accuracy and the
usefulness of the reply are separate quantities needing separate
instruments, and a system that answers quickly and wrongly will score
well here. Cases in which the metric does not apply, or applies only
with care, are set out in Section 11.
11. Applicability and Limitations
The metric as defined applies to systems that respond to a completed
caller utterance with a discrete response. The following cases are
outside its current scope or require care:
Filler and non-lexical audio: Addressed by the continuity
measurement in Section 3.7, which separates first audio from first
uninterrupted audio without interpreting content. A caller
comparing systems should read both figures, since MRL alone
rewards a system for making a noise.
Unprompted audio: Handled by Section 3.6, which requires detection
to begin after any greeting. Without that constraint the
measurement locates the greeting, and the figure it produces
describes a different event.
Barge-in: Where the caller interrupts the system, the interval
defined here is not the quantity of interest and the method does
not address it.
Incremental and streaming responses: Where a system emits partial
audio that it subsequently revises, first-audio onset does not
characterise the response.
Path contribution: As noted in the measurement domain, the figure is
a property of the pairing of path and SUT. Where the path
includes infrastructure the measuring party does not control,
Section 4.2.1 bounds its contribution rather than leaving it as a
caveat.
Conditioning across turns: See the open item under Measurement
Timing.
Nygate Expires 11 March 2027 [Page 27]
Internet-Draft Mouth-to-Ear Response Latency September 2026
12. Security Considerations
The method described here generates traffic directed at a system
under test and elicits responses from it. Measurement of a system
operated by another party without that party's authorisation may
constitute unauthorised use, may breach terms of service, and at
sufficient volume is indistinguishable from a denial of service
attempt. Parties performing measurement:
* SHOULD obtain authorisation from the operator of the system under
test;
* SHOULD limit the measurement rate to a level agreed with that
operator;
* SHOULD identify their measurement traffic where a mechanism to do
so exists.
Captures produced by this method contain audio. Where stimulus
material contains recorded human speech, that material is personal
data in many jurisdictions and may carry biometric significance, and
captures also hold the SUT's response audio, which can reflect the
content of the stimulus. Implementers:
* SHOULD prefer synthetic or consented stimulus material;
* SHOULD apply access control to stored captures;
* SHOULD set a retention limit rather than keeping captures
indefinitely.
Publication of captures as measurement evidence, which the method
otherwise encourages, has to be weighed against these considerations.
Reported figures may be commercially sensitive to the operator of the
system under test. The configuration identifier described in
Section 9 is specified as an identifier rather than as the
configuration itself so that reproducibility and confidentiality can
coexist.
13. IANA Considerations
This document has no IANA actions.
Registration of this metric in the Performance Metrics Registry
established by [RFC8911] would be a reasonable later step, and
Appendix A shows what such an entry would contain. It is
deliberately not requested here. The registry has to date been
Nygate Expires 11 March 2027 [Page 28]
Internet-Draft Mouth-to-Ear Response Latency September 2026
populated with IP-layer path metrics [RFC8912], so eligibility for a
metric of this kind is a question for the working group rather than
an entitlement to assert in an individual submission, and a request
made before the venue question of Section 1.2 is settled would be
moot if this work moves elsewhere.
14. Future Work
Conveying MRL between endpoints in band, by means of an RTCP Extended
Report block in the manner of [RFC3611], would allow a system to
report the metric about itself rather than requiring an external
caller to measure it. That is a separate document and a different
working group.
Time to greeting, the interval from session establishment to the
onset of the unprompted audio described in Section 3.6, is a second
quantity the method already observes without defining. It is at
least as visible to a caller as MRL, and it appears to be published
by nobody. Defining it here would widen this document beyond a
single metric, so it is noted rather than specified, and it would sit
naturally in whatever document this one becomes part of.
15. References
15.1. Normative References
[RFC2119] Bradner, S., "Key words for use in RFCs to Indicate
Requirement Levels", BCP 14, RFC 2119,
DOI 10.17487/RFC2119, March 1997,
<https://www.rfc-editor.org/rfc/rfc2119>.
[RFC3550] Schulzrinne, H., Casner, S., Frederick, R., and V.
Jacobson, "RTP: A Transport Protocol for Real-Time
Applications", STD 64, RFC 3550, DOI 10.17487/RFC3550,
July 2003, <https://www.rfc-editor.org/rfc/rfc3550>.
[RFC3551] Schulzrinne, H. and S. Casner, "RTP Profile for Audio and
Video Conferences with Minimal Control", STD 65, RFC 3551,
DOI 10.17487/RFC3551, July 2003,
<https://www.rfc-editor.org/rfc/rfc3551>.
[RFC8174] Leiba, B., "Ambiguity of Uppercase vs Lowercase in RFC
2119 Key Words", BCP 14, RFC 8174, DOI 10.17487/RFC8174,
May 2017, <https://www.rfc-editor.org/rfc/rfc8174>.
Nygate Expires 11 March 2027 [Page 29]
Internet-Draft Mouth-to-Ear Response Latency September 2026
[RFC8259] Bray, T., Ed., "The JavaScript Object Notation (JSON) Data
Interchange Format", STD 90, RFC 8259,
DOI 10.17487/RFC8259, December 2017,
<https://www.rfc-editor.org/rfc/rfc8259>.
15.2. Informative References
[G114] ITU-T, "One-way transmission time", ITU-T Recommendation
G.114, 2003.
[G711] ITU-T, "Pulse code modulation (PCM) of voice frequencies",
ITU-T Recommendation G.711, 1988.
[HARNESS] Nygate, D., "voice-ai-latency-harness: an instrument for
measuring mouth-to-ear response latency",
DOI 10.5281/zenodo.22124823, 2026,
<https://doi.org/10.5281/zenodo.22124823>.
[RFC2330] Paxson, V., Almes, G., Mahdavi, J., and M. Mathis,
"Framework for IP Performance Metrics", RFC 2330,
DOI 10.17487/RFC2330, May 1998,
<https://www.rfc-editor.org/rfc/rfc2330>.
[RFC3611] Friedman, T., Ed., Caceres, R., Ed., and A. Clark, Ed.,
"RTP Control Protocol Extended Reports (RTCP XR)",
RFC 3611, DOI 10.17487/RFC3611, November 2003,
<https://www.rfc-editor.org/rfc/rfc3611>.
[RFC6076] Malas, D. and A. Morton, "Basic Telephony SIP End-to-End
Performance Metrics", RFC 6076, DOI 10.17487/RFC6076,
January 2011, <https://www.rfc-editor.org/rfc/rfc6076>.
[RFC6390] Clark, A. and B. Claise, "Guidelines for Considering New
Performance Metric Development", BCP 170, RFC 6390,
DOI 10.17487/RFC6390, October 2011,
<https://www.rfc-editor.org/rfc/rfc6390>.
[RFC7679] Almes, G., Kalidindi, S., Zekauskas, M., and A. Morton,
Ed., "A One-Way Delay Metric for IP Performance Metrics
(IPPM)", STD 81, RFC 7679, DOI 10.17487/RFC7679, January
2016, <https://www.rfc-editor.org/rfc/rfc7679>.
[RFC7799] Morton, A., "Active and Passive Metrics and Methods (with
Hybrid Types In-Between)", RFC 7799, DOI 10.17487/RFC7799,
May 2016, <https://www.rfc-editor.org/rfc/rfc7799>.
Nygate Expires 11 March 2027 [Page 30]
Internet-Draft Mouth-to-Ear Response Latency September 2026
[RFC8911] Bagnulo, M., Claise, B., Eardley, P., Morton, A., and A.
Akhter, "Registry for Performance Metrics", RFC 8911,
DOI 10.17487/RFC8911, November 2021,
<https://www.rfc-editor.org/rfc/rfc8911>.
[RFC8912] Morton, A., Bagnulo, M., Eardley, P., and K. D'Souza,
"Initial Performance Metrics Registry Entries", RFC 8912,
DOI 10.17487/RFC8912, November 2021,
<https://www.rfc-editor.org/rfc/rfc8912>.
[TTFAB] OpenBenchmarks Labs, "Voice agent latency benchmark: time
to first audio byte measured from real phone calls", 2026,
<https://openbenchmarks.com/voice-agent-latency>.
Appendix A. Draft Performance Metrics Registry Entry
[[EDITOR'S NOTE: filled in against the column structure defined in
Section 7 of [RFC8911], using the blank template in Section 11 of
that document, as far as the current text supports. Incomplete
deliberately; the gaps are the same gaps flagged in the body.
Note that Section 7.1.2 imposes a structured naming convention on
registered metrics, which the Name entry below has yet to
satisfy.]]
Identifier: TBD by IANA
Name: TBD, following the naming convention of Section 7.1.2 of
[RFC8911]
URI: TBD
Description: The interval between transmission of the final speech
sample of a caller's utterance and arrival of the first sample of
a conversational voice system's response, observed at the calling
endpoint's RTP reference point.
Change Controller: IETF
Version: 1
Reference Definition: This document, Section 3.3 through
Section 3.5.
Fixed Parameters: Onset variant thresholds as tabulated in
Section 3.5; de-jitter anchor rule as specified in Section 3.4.2.
Reference Method: This document, Method of Measurement.
Nygate Expires 11 March 2027 [Page 31]
Internet-Draft Mouth-to-Ear Response Latency September 2026
Packet Stream Generation: Prerecorded stimulus transmitted at the
nominal frame rate of the codec in use.
Traffic Filter: The RTP stream of the session under measurement.
Sampling Distribution: TBD; see the open item under Measurement
Timing.
Runtime Parameters: Codec, frame period, playout target depth,
stimulus identifier and hash, SUT configuration identifier.
Roles: Calling endpoint; system under test.
Output Type: Signed scalar, reported in two variants, with an
uncertainty term.
Metric Units: Milliseconds.
Calibration: As specified in Section 4.7. An implementation reports
its own calibrated accuracy and the conditions of calibration.
Appendix B. Reference Implementation
An open-source implementation of this method exists and is archived
at [HARNESS]. It implements the metric definition, the onset
variants, the quality control conditions and the calibration
procedure described here, and it publishes per-host calibration
results.
This appendix is informative. Conformance is to this document, and
any disagreement between this document and that implementation is a
defect in the implementation.
Appendix C. Open Questions for Review
This appendix is retained deliberately rather than removed before
submission. An individual submission exists to attract review, and
the author's own list of doubts is a more efficient way to obtain
useful review than letting each reader rediscover them.
1. Filler and non-lexical audio is now specified in Section 3.7 by a
structural discriminator rather than by content recognition, and
implemented. What remains open is whether 150 ms and 2000 ms are
the right constants. They were chosen to sit well above the
pauses inside ordinary speech, measured at around 50 ms on the
reference implementation's material, and no corpus of real filler
behaviour informed them yet. Evidence welcome.
Nygate Expires 11 March 2027 [Page 32]
Internet-Draft Mouth-to-Ear Response Latency September 2026
2. The stimulus annotation rule is now a conformance test against a
published reference signal rather than a mandated algorithm, and
the signal exists. What remains open is where the signal should
live: carried in this document, registered, or referenced in the
archived implementation as it is here.
3. The interchange record in Section 9.2, specifically whether it
warrants a registered media type. The reference_point
enumeration of Section 4.2 answers the other half of what this
question used to ask; what remains open there is whether the list
is the right list, since it was drawn from observed practice
rather than from a survey.
4. Whether Informational is the right category, or whether the
conformance language argues for Standards Track.
5. Venue. Whether IPPM's charter admits a metric of this kind, and
if not whether DISPATCH or the Independent Submission stream is
the better route. See the editor's note in the Introduction.
6. Whether registry registration is appropriate for an application-
layer metric of this kind.
7. Whether Section 7.1 draws the line in the right place. It
requires the turn index to be recorded and forbids silent
pooling, and it deliberately does not prescribe how many turns to
measure. An argument that a fixed number belongs in the
specification, for comparability, is one the author would want to
hear.
Author's Address
Daniel Nygate
Email: dnygate@outlook.com
Nygate Expires 11 March 2027 [Page 33]