Minutes IETF126: fann: Thu 09:30
minutes-126-fann-202607230930-00
| Meeting Minutes | Fast Network Notifications (fann) WG | |
|---|---|---|
| Date and time | 2026-07-23 09:30 | |
| Title | Minutes IETF126: fann: Thu 09:30 | |
| State | Active | |
| Other versions | markdown | |
| Last updated | 2026-07-28 |
FANN Working Group - Fast Network Notifications
Date: Thursday, 23 July 2026
Time: 11:30-12:30 CEST (60 min, Session II)
Room: Grand Park Hall 3
Area: Routing (RTG)
Chairs: Carlos J. Bernardos, Eddie Ruan
AD: Ketan Talaulikar
Minute taker: Mike McBride, Jie Dong
1. Chairs' introduction (10 min)
Chairs
- Note Well, agenda bash
- WG charter, milestones Scope reminder: problem space first, solutions deferred
In scope for today's session: problem statement, requirement, gap
analysis and deployment scenarios.
(from chat)
Mike McBride: What's the purpose of the problem statement, etc not being
published.
Carlos Bernardos: @Mike the purpose is to guide the framework work, and
not publishing it is just stated in the charter.
Adrian Farrel: @Mike The IESG seems to be tidal on this sort of thing.
Current thinking is along the lines of "What purpose would be served by
publication?"
Adrian Farrel: My advice. Work on the doc as though it was going to be
an RFC. When it is ready, talk to the AD and get them to read the doc
before saying yes or no.
Mike McBride: Thanks will do. If problem/gap/requirement documents are
typically not published I'll leave it alone.
Ketan Talaulikar: There is recognition that this is a fast moving area
in the industry.
2. Network notification: problem statement, requirements and gap analysis (15 min)
Jie Dong / Mike McBride
- draft-dong-fann-problem-statement-00
Wim Henderickx: We should look at security implications. What are you
authenticating, and the security aspects related. BTS (back to sender)
is similar to this and suggest to look at that work as well. May also
ask them to share information. Should not duplicate the work.
Jie: Security was discussed, we do have security considerations and
focus on single domain to relieve security concerns for now. For the
second comment, can discuss offline.
Jeff Tantsura: BTS work has started several years ago in IEEE. We should
issue liaison to IEEE on that, ESC is not a standard today. Never before
have we understood the application layer constraints on the app layer.
The transfers are stateful, It has its own telemetry at different
levels. We need to understand the timing of the feedback loop to make
sure they don't interwork in strange ways. There are basic things as you
see in marking to more complicated RTT measurement and ECN measurement
on application layer. Soo all of this needs to be well undrestaood
before we decide outside of basic encoding, like timeing feedback loop,
what and who is the consumer, and what they do with this information.
Tim Chown: Adopt the problem statement, in good initial shape. Quite
impressive to have 9 documents already submitted. I support adoption.
(from chat)
Carolina Caeiro: I have a question on what document will establish the
type of information of Fast Network Notifications - would this be
established by the problem statement, or the requirements document
further down?
Carlos Bernardos: The PS document should shed light on what is needed
(based on the gap analysis), and then the framework will define and
specify that.
Carolina Caeiro: Can you clarify what you mean by "fine-grained network
status information"? What would this entail and how is it limited to
link failures/signal degradation?
Mike McBride / Jie Dong: Fine-grained network status information refers
to quantifiable network metrics such as link utilization, queue length,
level of congestion, link or node delay, jitter, and packet loss, etc.
(see problem statement draft, Section 4.1).
Ketan Talaulikar: The document is identified as a support document for
WG decision making and guiding the framework and further work. I leave
it to the WG on how much time to spend on polishing it and when to
switch focus on the framework.
Jeffrey Haas: It's a bit peculiar that the criticism of BFD's time is
less about what the protocol can carry (it carries timers in
microseconds), and more about what implementations have chosen to do for
scale. Different applications might push the core protocol faster.
Carlos: Poll to the WG.
- Have you read or reviewed this document (or the one adopted in the
RTGWG)?
73 Total participants
Yes - 30
No - 7
no opinion - 0
- Do you think this document forms a good basis for adoption?
74 total participants
Yes - 29
No - 0
- Reflections on interconnection scenarios (10 min)
Luis M. Contreras- (no I-D)
Jeffrey Haas: The minute you say put anything in BGP is no longer fast.
But if all your are doing is fast for global repair, you are not doing
anything useful.
Luis: The action would depend on the nature of the problem, maybe you
are seeing a trend, an increasing trend of traffic. Maby you have time
for reacting. Based on the problem you identify, you could have
different tools. Need to explore the proper situation and the timer for
it.
Jeff Haas: There will be different hiararchies of speed. The control
plane mechanisms can have the second level mitigation. And it is
recognized by this WG that control plane is the wrong place to solve
these fast notification problems.
Jie: This is a useful use case. Do you consider notifications sent
between different administrations, or is it still within the same
domain?
Luis: Basically domain b is taking action for their own response to the
problem. The starting point is the action taken in the own domain for
the event perceived on the interconnection.
Jie: Do you expect action will be in the control plane or the reaction
will be directly in date plane?
Luis: I didn’t consider the solutions. Just the problem. It would be
used as a trigger for immediate actions which are faster than human
reaction.
Jeff Tantsura: Great use case. Should drive us to think about the
encoding. Ver fixed, very fast versus something more flexible.
potientially TLV based, could encode the remote link information, but
may also the attached prefix information, Perhaps some other info could
be fed into BGP, and BGP can react. Another consideration is stateless
versus stateful solution. Most solutions drafts published are stateless.
- Problems and gap analysis for DCI congestion notification (10 min)
Xiao Min
Haoyu Song: The consumer of this fast CNP is the end host, need to
confirm that is in the charter scope or not. It was assumed the consumer
is a network device. If you do this it just informs the host and becomes
a transport layer problem. Also the second question is you send the fast
CNP, the packet still continue to the receiver host, the receiver would
still process the ECN, requies update to the end-to-end ROCEv2 protocol.
The DCI case further complicates the problem, it will involve much
longer latency, the mechanisms of RoCEv2 with ECN was designed for DC
networks with very small latency. The problem for DCI would be totally
different.
Xiao Min: Think it is within the charter. In the problem statement
draft, end host as the consumer is within the charter. DCI is long
distance latency, but it is required for real AI training deployments.
Jeff Tantsura: Switches does not need to know anything about QPs. In
IBTA spec, in CNP you need to reverse the source and Destination QPs.
There is no remapping whatsoever.
Xiao Min: The format of CNP, there is a source QP. The switch needs to
know the source QP to send fast CNP.
Jeff Tantsura: The receiver parse it and reverse it to the destination
QP.
Xiao: The receiver is not involved in this case.
Jeff: You need to check the spec. And reliable communications do not
include the source, need to figure out how to deal with that.
(from chat)
Joel Halpern: I note that pfc has an explicit lifetime, which could
easily be shorter than the latency across the wide area network. Is
there an expectation that this propagation takes that into account?
Adrian Farrel: @Joel. Important question, but akin to interflap or
bandwidth variance in TE. I think trigger thresholds and damping will be
critical.
Joel Halpern: @Adrian I am concerned because the PFC generation is
structured around assumptions about network topology, and thus it seems
important in trying to propagate such information to recognize the
differences in cases and explain how to deal with them. For our existing
IP TE solutions we recognize the inherent latency of the responses in
the design of the system.
Jie Dong: Regarding the scope of recipients, my understanding is the
major target is the network nodes which are in the traffic forwarding
path and can take actions in the network layer, delivering the
notification to hosts would need coordination with the transport area
WGs.
- Use cases and requirements for flow control collaboration across
DCNs and WAN (5 min)
Zhengxin Han
Takumi Sakabe: It’s two times RTT times Peak rate per tenant. At 500
kiometers, that is about 12 MBs for 10 giga tenent. So a few hundrads
tenants means gigabytes of the buffer at the edge. That is expensive
hardware. What is the assumed limit on the distance and tenant count? I
think those numbers belong in the requirements.
Takumi: How long is the distance?
Zhengxin: We've done 300 kms in the trial. Can talk offline about the
details.
- Ability requirements for stability guarantees in per-packet load
balancing networks (5 min)
Kefei Liu (remote)
Carlos: Will take the discussion to the list. The purpos is to feed the
gap analysis, and will have WG adopiton call on the first item (problem
statement).
-- time permitting --
There was no time for following presentations.
- A framework for Fast Network Notifications
Haoyu Song