Minutes IETF125: rtgwg: Thu 01:00
minutes-125-rtgwg-202603190100-00
| Meeting Minutes | Routing Area Working Group (rtgwg) WG | |
|---|---|---|
| Date and time | 2026-03-19 01:00 | |
| Title | Minutes IETF125: rtgwg: Thu 01:00 | |
| State | Active | |
| Other versions | markdown | |
| Last updated | 2026-03-25 |
IETF 125 RTGWG Minutes
09:00-11:00 - Thursday Session I, March 19, 2026
Chairs:
Jeff Tantsura (jefftant.ietf@gmail.com)
Yingzhen Qu (yingzhen.ietf@gmail.com)
WG Page: https://datatracker.ietf.org/group/rtgwg/about/
Materials: https://datatracker.ietf.org/meeting/125/session/rtgwg
Notetaker: David 'equinox' Lamparter (apologies if some notes are a bit
wonky, the room acoustics were not great.)
-
9:00
Meeting Administrivia and WG Update
Chairs (5 mins) -
9:05 / [9:05 → 9:12]
YANG Models for Quality of Service (QoS) in IP networks
https://datatracker.ietf.org/doc/draft-ietf-rtgwg-qos-model/
Aseem Choudhary (5 mins)
- Chairs: The draft has not been LCed yet. Please take the work to
finish line.
no questions/comments
-
9:10 / [9:12 → 9:20]
Fast Network Notifications Problem Statement
https://datatracker.ietf.org/doc/html/draft-ietf-rtgwg-net-notif-ps-00Jie Dong (15 mins)
- Jeff Haas: latency introduced into the feedback mechanism by
security, inherently slows things down, what are the latency
requirements to accomplish the goal? - Jie Dong: yes, was also discussed during the meeting about how much
security we should consider, balance security and the effectiveness
of the notifications. We will update the draft to reflect that. - Jeff T: we may use the concept of limited domain that SRv6
introduced. -
Yingzhen: We talked about this during the charter discussion. Please
consider updating based on the discussion. -
Jie: yes, single administrative domain.
(+ some discussion re. limited domain)
- 9:25 / [9:21 → 9:35]
IP Fast Reroute for AI/ML Fabrics
https://datatracker.ietf.org/doc/draft-clad-rtgwg-ipfrr-aiml/
Roy Jiang (10 mins)
- Greg Mirsky: you need to achieve 100µs switchover time. (yes) what
failure detection are you planning to use? - Roy Jiang: hardware based, on NPU
- Greg Mirsky: would BFD on hardware be suitable?
- Roy Jiang: no, BFD still involves CPU for handling, CPU is source of
slowness, want NPU to react to the failure. - Greg Mirsky: how to guarantee detection of link failure if nothing
is being transmitted on the link? - Roy Jiang: another topic, would discuss on list.
- Jeff T: It's implementation related and interesting.
-
Weiqiang Cheng: tried a very similar solution a few years back, will
talk offline -
Jeff Tantsura: interesting to see technologies come back. please be
careful to separate datacenter networks from WAN, evaluate both.
-
9:35 / [9:35 → 9:49]
Efficient Remote Protection
https://datatracker.ietf.org/doc/draft-clad-rtgwg-efficient-remote-protection/Francois Clad (10 mins)
-
Weiqiang Cheng: We found the same problem and similar solutions. I
provided a pointer to the draft in chat, and let's talk offline. -
Jeff Tantsura: robust survivability framework on the host, don't
forget to take that into consideration, RTT based decisions,
hairpinning is detrimental. If QP goes over longer path, entire
collective will suffer. Usually we have hardware props implemented
from the host side that are very fast. We must make sure the in
network fast reroute doesn't interact badly with host fast reroute.
This might be where host change route and use different GPUs. This
might not be open knowledge. The network should not intervene. -
Yao Liu: in multihomed scenario, how to know the receiver?
-
Francois Clad: need a network notification mechanism to make this
known to the sender, maybe subscription. -
Jeff Tantsura (chair hat): external definitions should come in from
FANN.
- 9:45 / [9:49 → 9:59]
Fully Adaptive Routing Ethernet in Scale-Up Networks
https://datatracker.ietf.org/doc/draft-xu-rtgwg-fare-in-sun/
Xu Xiaohu (15 mins)
-
Jeff Tantsura: most of the scale-up framework mentions no-ip
encapsulation, doesn't that mean we need non-IP identifiers? the
endpoint aren't IP addresses -
Xu Xiaohu: but some GPU vendors still support IP based solution.
-
Jeff Tantsura: still, approaches with compressed ethernet headers
are being discussed, definitely L2, this wouldn't work. -
Jeff Tantsura: in a single tier, there's no route and no BGP, we may
only need fast network notification. -
Xu Xiaohu: We're introducing BGP on the host. Even Nvidia supports
running BGP on host.
-
10:00 / [9:59 → 10:09]
Use cases and Requirement for Flow Control Collaboration Across DCNs and
WAN
https://datatracker.ietf.org/doc/draft-han-rtgwg-codeployment-pfc-fgfc/Zhengxin Han (10 mins)
-
Jeff Tantsura: buffer space is required for message depending on
distance of WAN links, there's a formula, can you include a table
with that in the draft? -
Zhengxin Han: yes, we will do need that. We have seen new devices
with huge buffers. -
Yingzhen Qu: an appendix might be useful
-
10:10 / [10:09 → 10:15]
Multicast Use Cases for Large Language Model Synchronization
https://datatracker.ietf.org/doc/draft-liu-rtgwg-llmsync-multicast/Yisong Liu (10 mins)
no questions/comments -
10:20 / [10:15 → 10:27]
Requirements and Gap Analysis of Multicast in AI Data Centers
https://datatracker.ietf.org/doc/draft-zhang-rtgwg-multicast-requirements-gaps-aidc/Junye Zhang (10 mins)
-
Tony Przygienda: BIER has a solution for the sparse part, there's a
related draft, uBIER, solves it up to a reasonable size -
Junye Zhang: appreciate the input. We should also optimize to meet
the requirement in AIDCs.
- 10:30 / [10:27 → 10:37]
Symmetry-Driven Asynchronous Forwarding with Fast Reroute for LEO
Satellite Networks (SDAF)
https://datatracker.ietf.org/doc/draft-luan-rtgwg-sdaf/
Mingliang Ke (15 mins)
no questions/comments
If time permits:
- [10:37 → 10:46]
Congestion Control Based on SRv6 Path
https://datatracker.ietf.org/doc/draft-liu-rtgwg-srv6-cc/01/
Yisong Liu
- Yuxuan Weng: What if multiple downstream nodes send congestion
message simultaneously to the same head end node? - Yisong: For SRv6, the congestion control message is only sent
one-hop up.
Chat History
Yingzhen Qu
00:01:51
Good morning! 你好!
Yingzhen Qu
00:02:10
please help with minutes:
https://notes.ietf.org/notes-ietf-125-rtgwg?both
David Lamparter
00:05:34
opens pad
Jeffrey Haas
00:17:41
For the fantel stuff, there will be an ugly interaction between security
and speed. crypto slows thing down.
Zafar Ali
00:20:26
Customers will deploy it in a secure domain
Jeffrey Haas
00:20:28
limited domain only changes the hand-waving about security. for these
mechanisms, latency for the feedback mechanism - along with programming
time - is the critical component
David Black
00:20:39
@Jeff - degree of hardware support seriously affects degree of ugly.
Jeffrey Haas
00:20:40
I laugh at the idea of "limted"
Tony Przygienda
00:21:24
the "limited domain" figleaf seems to become more and more the ultimate
excuse for anything now that is insecure or doesn't scale ;-)
Jeffrey Haas
00:21:34
^^^
Jeffrey Haas
00:21:50
"What does it take to secure this" is always the question
Joel Halpern
00:22:03
I at least would like to know what the "domain" is for the fann work.
Jeffrey Haas
00:22:42
A lovely property for a subset of the FANN work is that it may indeed be
a constrained subset of the topology. But that still means the
mechanisms need to be secure within that constrained subset.
Jeffrey Haas
00:24:10
2/3 layer ; 3/5 folded...
Zafar Ali
00:24:13
The (initial) focus should be on AIDC: scale-out (within the DC) and
scale-across (across data centers (DCI)).
Zafar Ali
00:24:37
This is what industry is asking for.
Jeffrey Haas
00:24:53
That's the expected case, Zafar. If the mechanism is IP UDP, describe
how you prevent injection of state that does Bad Things.
Jeffrey Haas
00:25:12
If it's L2, it's a lot more clear.
Yingzhen Qu
00:28:01
@David Lamparter thanks for helping with the notes. really appreciate
Joel Halpern
00:31:46
If you want every node to have topology knowledge, wouldn't it make more
sense to use a routing protocol (OSPF, IS-IS, LSVR) that provides
topology visibility?
Jie Dong
00:31:50
In the updated charter text, there are descriptions of the scope of the
domain: he Working Group will initially focus on solutions for networks
that are under a single administrative control or within a closed group
of administrative control. The FANN WG will not spend energy on
solutions for large groups of domains such as the Internet.
Jeffrey Haas
00:32:23
I can move my question to chat if time is short
krishnaswamy ananthamurthy
00:33:46
Every node will send the details to Route Reflector which can aggregate
and send it to the controller
Tony Li
00:34:04
And that will only take 10,000ms
Changwang Lin
00:34:22
There is a similar draft:
https://datatracker.ietf.org/doc/draft-liu-rtgwg-path-aware-remote-protection/,
which we can discuss together.
Tony Li
00:34:24
So are we proposing processing BGP in hardware?
Zafar Ali
00:34:38
The requirement is to get it done in HW and in usecond time
krishnaswamy ananthamurthy
00:34:45
BGP LS only for topology learning not for detection
Tony Przygienda
00:35:13
well, LSVR is basically this (IGP on top of BGP) unless we talk BGP-LS
to a centralized controller to program protection ? by now we can just
as well carry bits by hand in small buckets claiming it will be very
simple and very reactive ;-)
Jeffrey Haas
00:35:18
actually, I'll skip mic. The main issue is that FRR mechanisms require
that the control planes be consistent. As an example, if BGP-LS is used
for topology - great. However it means that the backup route needs to be
computed loop-free vs. whatever the underlying signaling mechanism is.
So, if it's "normal" bgp, it means the local node MUST understand if
it's backng up normal BGP using topology hints
Tony Li
00:35:22
Then that suggests that you're back to BFD
Jeffrey Haas
00:35:51
BFD, or at least fast interface notifications is the usual requirement.
Zafar Ali
00:37:50
Hi Please take a look at
https://datatracker.ietf.org/doc/draft-camarillo-rtgwg-lsn/. It
addresses the uSec notification requirements and also addresses Jeff's
comment on security.
Weiqiang Cheng
00:39:14
Francois, please see the draft:
https://datatracker.ietf.org/doc/draft-liu-rtgwg-path-aware-remote-protection/
Weiqiang Cheng
00:40:07
Your solution is really similar to the draft.
Jeffrey Haas
00:41:48
The bit vector mapping will be messy on that.
Jeffrey Haas
00:42:54
There's a missing mapping mechanism. Is that published yet?
Jeffrey Haas
00:43:14
Compare vs. the next-nexthop-nodes proposal that provides the mapping
layer
Tony Przygienda
00:43:54
once you go more than one hop you start to invent flooding basically
AFAIS or something equivalent since in epistomological terms it's the
same thing. and in hop by hop forwarding e'one inbetween will have to
adjust weights due to traffic shift. if we commit to state in the
network then it becomes basically RSVP for MP-TE ;-)
Jeffrey Haas
00:44:09
^^
Andrew Stone
00:44:23
localized flood (?)
Andrew Stone
00:44:37
flood within radius?
Jeffrey Haas
00:44:43
for path vector work, you can distribute one hop out. past that point,
you start looking closer to flooding mechanisms, even if limited
flooding scope
Tony Przygienda
00:44:46
"limited domain" flood? her we go. But I want TM on that ;-)
Andrew Stone
00:44:51
\:D
Zafar Ali
00:44:56
Mapping could be configured or discivery based.
Tony Li
00:44:58
If you don't want to flood, then the notification needs to follow the
reverse path of the incoming traffic
Martin Horneffer
00:45:05
I wonder why people don't afford some inter-spine links to make the
topology LFA capable.
Tony Li
00:45:35
People settled on leaf-spine a long time ago.
Jeffrey Haas
00:46:08
I think we can blame very old Clos work.
Tony Li
00:46:10
They will not change topology, they will not change protocols.
Tony Przygienda
00:46:25
inter-spine links even if used only to protect completely throw off BW
balance on the fabric as in "forget bisectional"
Jeffrey Haas
00:46:35
A signfiicant number of proposals work fine with "regular" topologies.
Martin Horneffer
00:46:48
how much complexity will they add just to avoid inter-spin links? shrugs
Tony Przygienda
00:46:56
CLOS has very strong linear programming theory under it as "cheapest way
to get bisectional"
Tony Li
00:46:59
Blame RFC 7938
Tony Przygienda
00:47:19
BGP on the other hand, I roll eyes the same way @Li does ;-)
Jeffrey Haas
00:47:41
I think you can blame 1950's switching networks.
Tony Przygienda
00:47:53
the theory still holds
Tony Li
00:47:59
No, this is squarely on one single OSPF implementation.
Tony Li
00:48:15
The screwdriver didn't work, so they decided to use the hammer.
Jeffrey Haas
00:48:25
pick your favorite routing protocol flashlight to say "this link is on
in a regular toplogy"
Tony Przygienda
00:48:57
it's not "something someone made up on powerpoint". Banyan work was
direct outcome of crossbar scaling limitations (and yes, a IP
router/switch is basically a crossbar) and we fight BW rather than
trunking capacity (#calls) but in terms of linear programming it's all
the same thing
Tony Przygienda
00:49:58
this one is fun. add an extended community, call it a protocol ;-)
krishnaswamy ananthamurthy
00:50:22
Q&A can be controlled by Chairs :-)
Tony Przygienda
00:50:45
no soup for you, no Q&A for you !
Jeffrey Haas
00:51:35
The solution isn't what I'd pick. The underlying problem is when you do
layer abstraction, how do you tie together the various bw signaling.
Pierre Francois
00:51:51
@Francois you keep TI-LFA on at the point of local repair until the
remote node activates its behaviour?
Yingzhen Qu
00:52:05
we do lock the queue. however we also want to see good disucssions.
Jeffrey Haas
00:52:28
best we can do is encourage presenters to be early to enable questions
and discussion
Tony Li
00:53:23
Presenters who spend their time poorly punish themselves.
Yingzhen Qu
00:53:26
@Jeff Haas. That's the point, save some time for Q&A.
Tony Przygienda
00:54:17
depends on objective, maximize face time, minimize the opportunity for
mike observations like "that simplay ain't gonna work" is a strategy
that has its own value ;-)
Les Ginsberg
00:54:31
Generally speaking people put too much info in the slides and therefore
have no time for discussion. RTFD yourself should be the model...
Francois Clad
00:54:31
@Pierre yes. TI-LFA is still useful to handle the traffic that gets to
the PLR
Tony Przygienda
00:54:32
as long there's a subversive chat running sideways I think we're fine
;-)
Yingzhen Qu
00:54:49
Talk straight to the point, and make better use of time for discussions
krishnaswamy ananthamurthy
00:56:20
@Yingzhen, I agree that presenters have to keep time for Q&A, also 10
min to less for 0th version draft at the same time.
Yingzhen Qu
00:57:32
depends on the draft. it's ok to request a 15 or 20 mins slot
Tony Przygienda
00:58:21
if you can't explain WHAT you're dong and WHY it matters on a -00 in 5
minutes it's basically "yet another intro to very basic problem solved
long ago I disovered and like 20 minutes to ramble on"
Pierre Francois
00:59:41
@francois the requirement draft says it's not fast enough :-)
Tony Przygienda
01:00:25
yeah, those presos were good, state the problem, say why current
solutions don't meet requirements, get feedback on what needs solved for
it to be of interest
Joel Halpern
01:00:53
How do we get folks to understand that there is no loss-less
transmission and there is no congestion-free transmission. There can be
transmission with low levels or very low levels of those failures. But 0
ain't it.
Tony Przygienda
01:01:32
the "broadcast my neighbor's link state to all my neighbors" has already
proposals and as Jeff said, it needs be L2 (i.e. in silicon) to make
sense in terms of better delays
Zafar Ali
01:02:24
@Joel yes indeed.
Tony Li
01:02:34
You cannot.
Tony Li
01:02:53
The industry has already locked in on the 'lossless' requirement.
Joel Halpern
01:03:21
I canna change the laws of physics captain. And I'm not in a television
program that then lets me do so.
Adrian Farrel
01:03:27
The marketing department took over many years ago
David Black
01:03:48
Have to get close to "lossless"- RDMA retransmit mechanism is
inefficient, hence needs to be used only very rarely.
Tony Li
01:04:11
And if you stand up and say that there emperor is wearing no clothes, no
one cares.
Tony Przygienda
01:04:14
@Joel, well, there is but you end up in regular topologies and
reservation and hence some form of signalling (as long links don't fail)
but that ain't IP ;-) And of course you end up with next level of
problems like buffering depth and head on blocking ;-)
Joel Halpern
01:05:36
@David Black: On the one hand, defining what the target is would be
useful. On the other hand, claiming that a protocol designed for a
different problem and then abused requires that we change the network
architecture seems odd?
Joel Halpern
01:05:57
Even ATM assumed there was no such thing as 0 jitter.
Tony Li
01:06:15
They write the check.
Jeffrey Haas
01:06:36
delusions make for poor code SLAs
Joel Halpern
01:06:39
@Tony li - wil they get upset when they can't get what they paid for?
Tony Li
01:06:55
Of course not.
Tony Li
01:07:27
That would require them to admit that what they wanted was undoable.
Joel Halpern
01:08:09
It does seem more and mroe that we are going to reinvent ATM with large
frames on fast links.
Tony Li
01:08:39
You mean Frame Relay?
Tony Przygienda
01:09:33
yepp, and since the big frames will cause the same problems we will
invent the frame-interrupt-interleaving stuff that AFAIR cascade
introduced on FR
Joel Halpern
01:09:40
Frame Relay SVCs never had the effort put in to work out all the corner
cases.
Joel Halpern
01:10:17
They got overtaken first by ATM and then by MPLS.
Tony Li
01:10:19
RFC 1925, rule 11
Joel Halpern
01:11:00
True.
Jeffrey Haas
01:12:40
What's back-ending the downloads in question for this llm sync
conversation? TCP and QUIC will leave links under-utilized for single
streams.
Tony Przygienda
01:13:15
core of the problem is that in terms of CAPEX % networking is a midget
so when the server/app guys say jump we can only ask "how high" and
pretend we won't fail over (while making sure we get paid for the anemic
resulting effort) ;-)
Jeffrey Haas
01:13:37
... and multicast will require dealing with lost packets and impact on
the s,g stream
Tony Przygienda
01:13:56
the multicast discussion is the same "describe the problem you have
which needs multicast and then tell me you must have it solved but you
don't want multicast under any circumstances" ;-)
Jeffrey Haas
01:14:11
ah. "invent reliable multicast" and then iteate.
Tony Przygienda
01:14:35
@Haas yes, reliable multicast blackhole. That's why BIER very, very
carefully killed any attempts to make it "reliable" from day one ;-)
Adrian Farrel
01:15:05
Surely we can just set a bit in the packet to say it MUST be delivered?
Jeff Tantsura
01:15:18
"evil" bit? :)
Adrian Farrel
01:15:35
Well, it would have to be the "not evil" bit
Tony Przygienda
01:15:41
and if that doesn't work we can add another bit saying it "REALLY MUST
be delivered " ;-)
Andrew Stone
01:15:58
Adrian april 1st is coming up :)
Tony Przygienda
01:16:09
Adrian, didn't you have a "soon" draft, myabe we need a "really" draft
;-)
Adrian Farrel
01:16:54
https://datatracker.ietf.org/doc/draft-farrel-soon/
Jeffrey Haas
01:21:24
snark aside, multicasting the data can reduce things. the challenge is
to bound the workload temporally based on reconciling incomplete data in
the multicast receiver.
Tony Przygienda
01:24:05
yepp, it would solve lots of what needed (except incast really) and
BIER, allowing extremely easy stripping would allow for a new degree of
cheap architectural freedom like e.g. stripping snapshoting if storage
is too slow etc ...
Tony Przygienda
01:24:59
so that's a good envelope. BTW sparness has a draft called u-bier which
solves it up to a reasonable size
Jeff Tantsura
01:26:02
as a side note - on a multi-plane fabric, it will take host about 20
micro to detect loss and move traffic to other planes
Tony Przygienda
01:27:22
https://www.ietf.org/proceedings/120/slides/slides-120-bier-u-bier-00.pdf
Zafar Ali
01:27:31
Before discussing any solution, I am still unsure of multicast
requirement in this area.
Tony Przygienda
01:27:39
example of <5 mins, 4 viewgraphs, explain problm, suggest novel approach
Mike McBride
01:28:22
There are p2mp and mp2p ai traffic patterns that we think mcast could
help with.
Tony Przygienda
01:28:43
@Zafar, use cases were shown and they are valid IME. big requirement
that multicast cannot meet AFAIS is incast, it's really a congestion
control problem (or otherwise we end up in token reservation on
multicast of some such and with that reliable mcast again)
Martin Horneffer
01:30:45
isn't retransmission still an open issue for reliable MC transport?
Tony Przygienda
01:31:05
interesting preso, ordered FIB sounds intriguing but how do you sync up
the vector clocks across nodes to make sure node 1 order is same as node
2 order to get a consistent cut ?
Tony Przygienda
01:33:24
hmm, that is interesting and looks like token ring with every node
having a token (so no congestion control)
David Lamparter
01:34:15
"extremely low" complexity is a bold claim
Tony Li
01:34:40
The reference to MPLS-TP is triggering my PTSD
Tony Przygienda
01:35:28
it is low complexity since it does not really involve computaton. In
coarse terms, you need a ring and yu send things twice (counter rotating
token ring was doing something "similar" AFAIR but with a token budget)
David Lamparter
01:35:57
hm.
Tony Li
01:36:02
Counter-rotating is not the shortest path. Traverse to the next torus
would be.
Tony Li
01:36:21
s/torus/ring/
Tony Przygienda
01:36:28
yeah, torus is another kettle of fish. never thought through that one
Jeffrey Haas
01:39:28
... it's almost as if we're inventing many of the fun things rsvp had in
srv6
Tony Przygienda
01:40:01
noo, Jeff. we do signalling w/o calling it signalling since SRv6 does
NOT have signalling ;-)
Joel Halpern
01:40:33
Even RSVP had the sense not to reroute dynamically based on congestion
signals. If we assume adding resource reservations to SRv6 (or SR-MPLS),
then we are definitely back to reinventing old work.
Jeffrey Haas
01:41:06
perhaps less core state, but the head end thundering herd...
Tony Przygienda
01:41:23
who could have predicted THAT (except people like Yakov after first
presos and couple other decently experienced and clever people) ;-)
Tony Przygienda
01:43:00
I wait with bathed breath for the "congestion SID" that is surely to
come now ;-)
Zafar Ali
01:43:19
I feel it is the other way around. RSVP-TE was always HBH without any
ECMP/ uCMP. SRTE introduced ECMP/ UCMP for TE ~15 years ago. Now
Multi-path TE is reinventing the wheel with ECMP/ UCMP in RSVP-TE (which
it was not designed for).
David Lamparter
01:44:03
"SyncE is just STM for GenZ"
Zafar Ali
01:44:05
I do not agree with the proposal being presented
Jeffrey Haas
01:44:08
ECMP can always be formed at a node based on available paths.
Tony Przygienda
01:46:09
ECMP groups over MPLS have been there about forever incl protection (as
is bin packing which is kind of much harder problem)
Zafar Ali
01:46:43
Can you explain how RSVP-TE mid point handles ECMP/ UCMP?
Tony Przygienda
01:47:10
MP-TE is a bit of a different beast, it's kind of unicast multicast ;-)
almost. RSPV is just convienent signalling to mod to support it
Andrew Stone
01:47:25
Tony - also note there's SR MPTE if you haven't seen :)
Zafar Ali
01:48:02
I have seen, SR MPTE does what you call MPTE with ZERO change to SR.
Tony Li
01:48:04
It might amuse you to recall that RSVP was originally designed for
mcast.
Jeffrey Haas
01:48:19
^^
David Lamparter
01:48:21
Apologies if the notes are a bit wonky
Tony Przygienda
01:48:29
^^