← Back to list
Networking
#STP#RSTP#MSTP#루프방지#L2스위칭
Last updated · 2026-10-09

Spanning Tree Protocol (STP·RSTP·MSTP)

1. Overview

A. Definition

Spanning Tree Protocol (STP) is the IEEE 802.1D standard protocol that, in an L2 Ethernet switch network containing physical loops, logically blocks selected ports so as to automatically compute and maintain a single loop-free tree topology. Applying the graph-theory notion of a spanning tree to networking, it builds a set of paths that keeps every switch connected yet acyclic, fundamentally preventing the broadcast floods caused by L2 loops.

B. Background and necessity

For availability, Ethernet connects switches with doubled or tripled links, yet an L2 frame has no lifetime field like the IP packet's TTL. Consequently, if a loop exists, broadcast and unknown-unicast frames circulate endlessly and multiply exponentially, producing a broadcast storm. At the same time, the same frame returns via multiple paths and the MAC address table thrashes — MAC flapping — paralyzing the switch's CPU and bandwidth within seconds.

Indeed, after repeated incidents where a single miswired cable took down an entire enterprise network, the STP devised by Radia Perlman in 1990 was standardized as IEEE 802.1D and became the de facto prerequisite for Ethernet L2 redundancy. For example, even an ordinary redundant design that cross-connects two access switches to two distribution switches forms, without STP, a closed circuit of four links that is itself the seed of a storm. In such a topology STP automatically puts one of the four links to sleep, guaranteeing a loop-free state in normal operation.

The evolution of the standard is also worth noting. The original 802.1D (1990) defined basic STP; RSTP for fast convergence was separately standardized as 802.1w (2001) and later merged back into 802.1D-2004; and MSTP for per-VLAN multiple trees appeared as 802.1s (2002) before being absorbed into 802.1Q. Thus "802.1D" today usually refers to the 2004 edition including RSTP, and new equipment supports RSTP behavior or better by default.

The essence of STP is reconciling the conflicting demands of physical redundancy (availability) and logical loop-freedom (stability). Lines are laid redundantly to prepare for failure, but in normal times some of them are blocked to maintain the tree; when an active path fails, a previously blocked port is reawakened and switched to a detour route. In other words, STP is a self-healing L2 topology management mechanism that "keeps spare paths asleep until needed," and the history of improving its convergence speed and bandwidth efficiency is precisely the STP → RSTP → MSTP evolution.

2. Overall structure — bridge topology and root-centered tree

flowchart TB
  R["Root Bridge<br/>(lowest Bridge ID)"]
  B2["Switch B"]
  B3["Switch C"]
  B4["Switch D"]
  R ---|"Designated Port(DP)"| B2
  R ---|"Designated Port(DP)"| B3
  B2 ---|"Root Port(RP)"| R
  B3 ---|"Root Port(RP)"| R
  B2 --- B4
  B3 --- B4
  B4 ---|"Root Port(RP)"| B2
  B4 -. "Blocking Port(Blocking)<br/>loop removed" .- B3
  style R fill:#fef3e8,stroke:#ed8f2f,stroke-width:3px
  style B4 fill:#e8f0fe,stroke:#2f6fed

The reference point from which STP builds the tree is the Root Bridge. All switches exchange a control frame called a BPDU (Bridge Protocol Data Unit) every 2 seconds (Hello), compare the Bridge ID it carries (a 2-byte priority + 6-byte MAC address, 8 bytes total), and elect the switch with the smallest value as the single root of the whole network. The default priority is 32768; on a tie, the lower MAC address wins.

Because the root is the apex of the tree and the origin of all path calculations, in practice one manually sets a low priority (e.g., 4096) on a stable switch in the core/distribution layer so that the intended switch is guaranteed to become root. If the root is not designated, aging equipment that happens to have the lowest factory MAC may accidentally become root, commonly causing the inefficiency of traffic being drawn toward a slow corner switch.

Once the root is fixed, each remaining non-root switch selects as its Root Port the single port with the lowest cumulative Path Cost to the root. Path cost is conceptually the inverse of link speed; by the IEEE standard values (short method), 10Mbps=100, 100Mbps=19, 1Gbps=4, 10Gbps=2, so slower lines have higher cost and are more easily excluded from the tree. As 10G and above became common, 16 bits could no longer distinguish them, so a 32-bit long method using a wider range such as 1Gbps=20000 and 10Gbps=2000 was standardized in parallel.

On each LAN segment, the switch with the lowest cost to the root holds the Designated Port that represents that segment, while any port that is neither a root port nor a designated port is blocked and does not forward frames. It is exactly this blocking port that breaks the physical loop and completes the tree. Importantly, a blocking port is not dead: it keeps receiving BPDUs and watches the tree, and the instant an active path fails it is promoted at once to become a detour.

The main components are summarized below.

Element Role Note
BPDU Frame exchanging topology info between switches Configuration BPDU·TCN BPDU, Hello 2s
Bridge ID Basis for root election (priority+MAC) smaller wins, default priority 32768
Root Bridge Apex of the tree, origin of path calculation one per network
Root Port (RP) Lowest-cost port toward the root one per non-root switch
Designated Port (DP) Segment representative port (forwards frames) one per segment
Blocking Port Stops forwarding to remove loops reactivated on active-path failure

BPDUs come in two kinds by role. The Configuration BPDU, which the root propagates downward every Hello interval in normal operation, carries the root ID, path cost, sending Bridge ID, and timer values to maintain the tree. By contrast, when some switch's port state changes (link down/up), that switch sends a TCN (Topology Change Notification) BPDU upward toward the root, and the root responds by sending a Configuration BPDU with the TC flag set down to the whole network.

This topology-change notification matters because of the consistency of MAC learning information. Switches that receive a TC shorten the aging timer of the MAC address table from the default 300 seconds to the Forward Delay (15 seconds), so that stale MAC learning tied to the old path is flushed quickly and relearned for the new topology. The wider the TC propagation range, the greater the temporary flooding due to relearning, so keeping the L2 domain small is key to operational stability.

3. Convergence process and port state transitions

stateDiagram-v2
  [*] --> Blocking: port UP
  Blocking --> Listening: selected as designated/root port
  Listening --> Learning: Forward Delay 15s
  Learning --> Forwarding: Forward Delay 15s
  Forwarding --> Blocking: topology change or loop detected
  note right of Listening: processes BPDUs only, no frame/MAC learning
  note right of Learning: begins MAC learning, frame forwarding not yet

STP convergence passes through four decisions.

A. Root bridge election. When a port comes up, every switch emits BPDUs assuming itself to be root, but on hearing a smaller Bridge ID it acknowledges that side as root and stops propagating its own BPDUs. In a few exchanges the whole network agrees on a single root; since an initial priority-setting mistake is the most common cause of incidents in this process, the priorities of the root and backup root must be explicitly assigned at design time. The backup root is given a priority one step lower than the root (e.g., 8192) so that a predictable switch succeeds on root failure.

B. Root port election. Each non-root switch adds its own port cost to the received BPDU to compute the cumulative cost to the root, and designates the lowest-cost port as the root port. On a tie, the sender's Bridge ID decides, then the port ID (priority+port number). For example, if an access switch has both a 1Gbps direct link (cost 4) and a 100Mbps detour (cost 19), the 1Gbps side becomes the root port and the 100Mbps side is pushed to standby.

C. Designated port election. On a segment where two switches meet, the side with lower cost to the root holds the designated port, and on a tie the one with the smaller Bridge ID wins. A port that becomes neither a root port nor a designated port is blocked and excluded from the tree. As a result, the correct converged state is one where every segment has exactly one designated port and every non-root switch has exactly one root port.

D. Port state transitions and timers. Traditional STP transitions a port in the order Blocking → Listening → Learning → Forwarding, where Listening and Learning each take a Forward Delay of 15 seconds, and a blocking port takes Max Age 20 seconds to detect a failure, so in the worst case a 30–50 second service interruption occurs. This "near-minute-scale" convergence delay is fatal for VoIP and real-time services, so RSTP (802.1w) emerged — reducing the port states to the three Discarding, Learning, and Forwarding, and switching instantly via a Proposal/Agreement handshake — achieving convergence within 1 second.

The default timers used by traditional STP and the resulting convergence delays are summarized below. These timers are propagated by the root via BPDUs, so the whole network follows the root's values; arbitrarily reducing them induces false detections and instability in networks with large propagation delay, so as a rule they should not be tampered with.

Timer Default Meaning Impact
Hello Time 2s BPDU transmission interval failure-detection sensitivity
Max Age 20s wait before discarding info on BPDU loss indirect failure detection
Forward Delay 15s dwell in each of Listening·Learning main cause of transition delay
Convergence (direct failure) ~30s Listening 15 + Learning 15 two-stage transition
Convergence (indirect failure) ~50s Max Age 20 + 30 blocking-port promotion

RSTP's fast transition rests not on mere timer reduction but on finer-grained port roles. Beyond the existing RP/DP, RSTP explicitly maintains, as immediately promotable standby roles, an Alternate port (an alternate path to the root) and a Backup port (a spare on the same segment). Thus, when the root port is cut, it raises the Alternate straight to Forwarding without waiting for a timer, realizing what is practically a near-hitless transition. Comparing the correspondence of port roles and states with traditional STP:

Aspect Traditional STP RSTP
Number of states 5 (incl. Disabled) 3
Active states Blocking/Listening/Learning/Forwarding Discarding/Learning/Forwarding
Standby roles none (lumped as Blocking) Alternate·Backup explicit
Transition method timer expiry Proposal/Agreement handshake
Edge port separate feature (PortFast) standardized built-in (Edge Port)

4. STP·RSTP·MSTP comparison

The differences among the three standards are not a mere version bump but the result of solving different problems along two axes, "convergence speed" and "VLAN scalability." RSTP targets transition speed; MSTP targets the control load of large-scale VLAN environments.

Aspect STP (802.1D) RSTP (802.1w) MSTP (802.1s)
Convergence time 30–50s within 1s within 1s (RSTP-based)
Port states 5 stages 3 (Discarding/Learning/Forwarding) 3
Port roles RP/DP RP/DP/Alternate/Backup RP/DP/Alternate/Backup
VLAN handling single tree for all (CST) single tree many VLANs→few instances (MSTI)
Main use legacy small/medium scale large campus with hundreds of VLANs

Traditional STP keeps only one tree (CST) across the whole network, so all VLANs use the same path and a blocked redundant link is never used in normal operation, wasting nearly half the bandwidth — a structural limitation. Cisco's PVST+ distributed load by running a separate tree per VLAN, but with hundreds of VLANs the BPDU and CPU burden exploded accordingly.

MSTP (802.1s) bundles multiple VLANs into a few MST instances (MSTI) — e.g., VLANs 1–500 to instance 1 (root = switch A) and 501–1000 to instance 2 (root = switch B) — to activate both sides of a redundant link with two trees while capping the control overhead at the number of instances. With roles thus divided as "speed from RSTP, scalability from MSTP," practice generally adopts Rapid-PVST+ or MSTP by default. One pitfall, however, is that to belong to the same MST region the region name, revision, and VLAN-instance mapping must match exactly on all switches, so a mapping mismatch splits the region and silently falls back to a single tree against intent.

As a concrete example, consider a university campus that dual-connects 10 buildings each to 2 distribution switches and operates 600 VLANs. Using only traditional STP (CST), one of each building's redundant links (10 links total) is permanently blocked, putting tens of Gbps of redundant bandwidth to sleep. Switching to MSTP and splitting VLANs 1–300 to instance 1 (root = distribution A) and 301–600 to instance 2 (root = distribution B), the link blocked in instance 1 becomes an active path in instance 2, so both uplinks are used. As a result the usable bandwidth effectively doubles, and even if one distribution switch dies, the traffic of the other instance keeps flowing. By contrast, running 600 separate trees with PVST+ would have multiplied the BPDU and compute load 600-fold and hit the switch CPU's limit. This is precisely the practical reason MSTP is chosen on large campuses.

For stable operation one must also understand the protection features. An access port where an endpoint attaches is set to PortFast to skip Listening/Learning and go to Forwarding immediately, but if such a port receives a BPDU (a situation where a switch is wrongly connected), BPDU Guard instantly blocks the port (err-disable) to prevent a loop. Representative others are Root Guard, which prevents a root takeover by a lower-Bridge-ID BPDU arriving from outside, and Loop Guard·UDLD, which prevent a blocking port from wrongly going Forwarding when it fails to receive BPDUs due to a unidirectional link fault. An L2 network operated without these protections is effectively a time bomb.

5. In depth — the decline of STP in data centers and the shift to L3 fabrics·VXLAN-EVPN

STP is still active in campus and branch networks, but in modern data centers dominated by East-West traffic it is being rapidly displaced. The root causes are that ① it wastes bandwidth by blocking half the redundant links, ② its single-root-centered structure concentrates traffic at the core, and ③ there is the blast radius problem of service shaking during convergence on topology change. The larger a single L2 domain grows, the more a single topology change propagates to the whole, and the difficulty of containing the failure impact scope was the greatest operational burden.

In response, the industry tried, as transitional measures, TRILL (RFC 6325) and SPB (802.1aq) to solve L2 multipath within L2, but ultimately converged on the direction of "confine L2 to be small and expand via L3." TRILL and SPB overcame STP's single-tree limitation with IS-IS-based shortest-path multipathing, but failed to become mainstream due to legacy-equipment compatibility and a lack of ecosystem.

Today's standard design is to build a leaf-spine (Clos) fabric purely in L3 and activate all links simultaneously with ECMP, then provide the L2 connectivity tenants need via a VXLAN overlay with its control supplied by the BGP EVPN control plane. In this structure the redundant paths STP used to block are all used for forwarding, so usable bandwidth doubles, and failure convergence is handled by the routing protocol (tens to hundreds of ms).

Server dual-homing, too, instead of STP blocking, uses MLAG/vPC to bind two upstream switches into one and uses both links. Even in an MLAG environment, however, STP (or EVPN's duplication-prevention mechanism) still runs between the two peer switches as a last-resort safety net, preventing the worst case where a loop forms when the peers are split (split-brain) by a control-plane failure. That is, even in modern design STP survives with its role changed from "primary path controller" to "final safeguard against loops."

In sum, STP has not disappeared; rather its scope of application has shrunk to "small L2 domains and campuses," and the data center's multipath challenge has been taken over by L3 routing and overlays. In a professional-engineer answer, describing this "L2 STP → L3 fabric + overlay" transition context, together with the point that STP nonetheless remains a safety net, effectively reflects the latest trends.

6. Considerations and implications (professional engineer's perspective)

  • Explicit fixing of root design: The root and backup root priorities must be manually assigned (e.g., 4096/8192), and PortFast·BPDU Guard must be made the standard for access ports. Left to automatic election, aging switches with the lowest MAC become root, commonly causing incidents where traffic gets tangled.
  • Trade-off between convergence speed and service level: If real-time services exist, traditional STP (30–50s) is unsuitable, so make RSTP/Rapid-PVST+ the default — while also considering that, without protection features (Loop Guard·UDLD), fast convergence can instead spread a loop rapidly.
  • Bandwidth efficiency vs. operational complexity: Reviving redundant links with MSTP·PVST+ increases bandwidth but complicates VLAN-instance mapping and region-boundary configuration. Single trees and multiple instances should be applied selectively according to the number of VLANs and the organization's scale.
  • Architecture transition strategy: New data centers should review L3 leaf-spine + VXLAN-EVPN as the default rather than STP-dependent L2 expansion, and existing L2 networks are best transitioned gradually by partitioning the domain (small L2 + L3 separation) to reduce the blast radius.
  • Responding to security threats: If an attacker injects a forged BPDU with a low Bridge ID to hijack the root, they can draw traffic to themselves and eavesdrop (MITM); therefore enforcing Root Guard·BPDU Guard·BPDU Filter on boundary ports and fundamentally blocking BPDU reception on user ports should be the security baseline.
  • Related-technology perspective: Understanding STP is the premise for designing VXLAN·EVPN·MLAG·SDN (centralized L2 control based on OpenFlow)·microsegmentation, so it should be learned not as a single protocol but as the starting point of L2 availability design.

References


In one line: STP (802.1D) is a protocol that, to prevent L2 Ethernet loops, automatically computes a loop-free root-bridge-centered tree and blocks spare ports; its slow convergence (30–50s) and bandwidth waste were improved by RSTP (802.1w, within 1s) and MSTP (802.1s, multiple instances), and in modern East-West-centric data centers its role is being reshaped by L3 leaf-spine fabrics and VXLAN-EVPN·MLAG.