nanog mailing list archives

Re: IS-IS Flat Network Size


From: Pedro Prado via NANOG <nanog () lists nanog org>
Date: Tue, 28 Jul 2026 09:04:13 +0100

I'd say, just get know what you need to monitor in case things start to
exceed the current configuration (perhaps copp) or the supported scale
(maybe process health,memory).

*Pedro Martins Prado*
pedro.prado () gmail com / +353 83 036 1875 (FaceTime & WhatsApp)


On Tue, 28 Jul 2026 at 08:44, Saku Ytti via NANOG <nanog () lists nanog org>
wrote:

This is absolutely the norm. People framing it as some engineering
challenge that needs consideration are unnecessarily creating
complexity and concern where there is none.

Tier1s run global flat level2 at this scale, and have since forever.

Outage information cannot propagate faster than light, no matter how
you. bake it.

The justification should be the opposite, you should have a strong
reason not to run flat IGP, if you don't have it, run flat.

You will struggle to justify any of what is proposed here, regarding
SPF time, regarding convergence time, that these can be improved by
adding complexity. SPF and convergence don't even matter in any modern
design, because when you converge, you converge for all single
failures too, that is, on link-down, you immediately forward around
the problem, instead of waiting for SPF, since waiting for SPF takes a
long time in any scale, even intraDC.


On Tue, 28 Jul 2026 at 06:48, Matthew Petach via NANOG
<nanog () lists nanog org> wrote:

On Mon, Jul 27, 2026 at 6:14 PM Dan Snyder via NANOG <
nanog () lists nanog org>
wrote:

It really depends on what hardware you are using, how large of a fault
domain you want if there is an issue in the L2 area, and how quickly
you
expect it to converge if there is a failure.


To add a bit to Dan's excellent reply

When I'm working on a network design, one of the questions first and
foremost in my mind is
"what's the blast radius for the common failures that will happen, and
where does it make
sense to put blast doors in?"
Failures come from many sources; bugs in automation tools, humans making
typos in inputs
to automation tools, hardware failures at L1, L2, L3, fiber cuts, etc.
The duration of the outage also comes into play; undersea, transoceanic
fiber failures can have
months-long repair cycles, whereas a failed optic in a datacenter
aggregation router can often
be swapped out in minutes.

While it can often be tempting to simply keep scaling up a flat network
as
you grow larger and
larger, what you'll generally find is that your fragility increases as
you
do that.  What's more is
that the increase in fragility follows a power law that grows faster than
the diameter of your flat
network, and that there are stepwise jumps in the fragility of your
network
as you hit certain
boundary points.

Recognize that restoration times are long for subsea links; that costs
for
them are much higher,
so putting in N+M or 2N pathways subsea almost never pencils out for the
finance team paying f
or them; that trying to find sufficient diversity that you can do N+M
for a
small value of M while
ensuring that there's no shared fate between any of N and M so that you
can
survive a subsea
cut without having to do higher-order traffic engineering is an
exponentially increasing problem.
And then stare long and hard at the beautiful diagram of your single flat
L2 network that
circles the planet, and think to yourself "when 2/3 of my transpacific
capacity goes down
due to an undersea earthquake and landslide in the strait of taiwan, how
am
I going to
tell a flat L2 network that I don't want *all* my traffic still flowing
across the few links that
remain?"[0]

That is, Mark Tinka and Tom Beecher are spot-on; IS-IS will handle the
number of nodes you're
talking about plus another order of magnitude without blinking; and with
a
few knobs, you can
even keep your SPF timing reasonable with round-the-planet latencies.
But just because it *can* doesn't mean that's a good network design.
^_^;

So, I would encourage you to think about the size and scope of your
network, not just today
but where you envision it being 3  years from now, 5 years from now, 10
years from now,
and start thinking "how can I start putting blast doors into my network
to
limit the impacts
of failure, and give me higher-level control points that will let me
traffic engineer in ways
that go beyond what I can do with a simple flat L2 IS-IS network?"

I'll admit--I don't know the first thing about your network.  It may be
that you have a tightly
geographically constrained network, so your edge-to-edge latency is low
double digits,
you have plentiful access to diverse fiber everywhere you need to go so
that you can
overengineer your Quantity Of Service to a degree that you can leave
everything to
IS-IS to route around failures without ever needing to traffic engineer
your traffic, and
you have trucks ready to roll at a moment's notice to swap out any failed
hardware
in less time than it takes a sleeping network engineer to wake up, find
their glasses,
find their phone for 2FA to log into the network tooling to figure out
what
broke, why
they got paged, and how to route around it.  If so, that's awesome!  In
that case,
you can rest easy knowing that a flat IS-IS L2 topology will scale up
just
fine for
you, even at 10x your current size.  :)

But...on the off-chance that's not what your network looks like, I've
found
that
thinking about failure modes and how to limit the impact of them goes a
long
way towards developing a robust network design, and that the network
design
you arrive at is likely to not be a simple single flat L2 topology.

Thanks for asking the question!  :)

Matt


[0]There's a reason many networks do continental IS-IS networks, and then
use BGP plus their favorite flavor of traffic engineering overlay across
transoceanic links.
BGP provides very, very strong blast doors to limit the blast radius of
IGP
goofs, as well
as an easy way to ensure your IGP won't continue to try passing all of
your
traffic across a
small number of remaining transoceanic links during a major outage.
_______________________________________________
NANOG mailing list

https://lists.nanog.org/archives/list/nanog () lists nanog org/message/2ESUIWF3PFXIH4XZ3FSBV4LGOLEIBMU2/



--
  ++ytti
_______________________________________________
NANOG mailing list

https://lists.nanog.org/archives/list/nanog () lists nanog org/message/QFDLIZ5IPDNPYCD2J5WRNXJGCY24PDNA/
_______________________________________________
NANOG mailing list 
https://lists.nanog.org/archives/list/nanog () lists nanog org/message/N2HLYQN2JPTZ53ATXJNB46KVVOCES355/

Current thread: