A Beginner’s Guide to NVIDIA Spectrum and Spectrum-X Networking

Spectrum vs ordinary switching, fabric thinking, history, and AI motivation
  1. Part 1 What NVIDIA Spectrum Is
  2. Part 2 Why NVIDIA Spectrum
  3. Part 3 — AI Factory Fabric Architecture
  4. Part 4 — RDMA and ROCE
  5. Part 5 — NOS, IaC, and Operations
  6. Part 6 — HA, Physical Layer Design, and Scaling
  7. Part 7 — Multi-Cloud, Security, and Lifecycle Management
  8. Part 8 — Design Scenarios, Glosary, and References ← You are here

Spectrum Networking Design Walkthrough

Scenario

A fictional organization is building an AI training and distributed-inference environment with 512 GPU servers across 16 racks. Each server has one 800G-class AI-network attachment represented here as two 400G ports for simple cabling arithmetic, plus separate service and OOB connectivity. Storage is logically isolated and may be physically separated as growth requires.

1. Requirements

  • 512 GPU servers across 16 racks, 32 servers per rack.
  • High east-west bandwidth for NCCL collectives and distributed inference.
  • Near-non-blocking normal operation with defined degraded-state headroom.
  • RoCEv2 with validated Spectrum-X QoS/congestion profiles.
  • BGP/ECMP routed underlay; no EVPN/VXLAN on the dedicated AI back-end unless a future multi-tenant requirement appears.
  • Separate OOB management and separate service network.
  • Central telemetry with NetQ and external metrics/log integration.
  • Git/CI/DSX Air change workflow.

2. Topology

Use a two-tier leaf-spine design. Each rack receives redundant leaf connectivity or independent rails according to the selected Spectrum-X reference architecture. Spines provide equal-cost paths between every rack. If the selected generation/radix cannot meet growth targets in two tiers, evaluate Spectrum-X Multiplane before adding a third tier.

Figure 30. End-to-end fictional AI factory

3. Capacity assumptions

For each rack, calculate total server injection bandwidth and match leaf uplink bandwidth to the target oversubscription ratio. For training-heavy workloads, begin with a non-blocking or close-to-non-blocking target, then validate with application models. Reserve headroom so loss of one spine or one rail does not immediately create persistent queues.

4. Host connectivity

Map each GPU server’s SuperNIC ports to independent leaf/plane paths. Confirm PCIe and GPU locality so the NIC serving a GPU is not unnecessarily crossing CPU sockets or constrained PCIe paths. Use the NVIDIA-validated NIC firmware and Spectrum-X profile for the selected reference architecture.

5. Routing

Run an L3 underlay with eBGP leaf-to-spine adjacencies and ECMP. Use loopbacks for stable identities and BFD where the failure-detection requirement justifies it. Keep the AI fabric routing table intentionally simple: host/rail prefixes and infrastructure reachability, not enterprise service routes.

6. Congestion strategy

Use the validated Spectrum-X RoCE profile. ECN and endpoint congestion control should prevent persistent queues; PFC, if the validated mode uses it, is limited to the required priority. Enable supported adaptive routing on the eligible links and observe path-balance telemetry. Do not hand-tune thresholds before establishing a repeatable workload baseline.

7. Resilience

Model link, leaf, spine, and rack failures. Confirm both logical reachability and remaining bandwidth. Define a maintenance drain procedure so a spine or leaf can be removed from service without surprising the application team.

8. Operations and 9. Observability

All configuration originates from source control. CI validates addressing, BGP peers, MTU, QoS profiles, and topology. DSX Air validates control-plane and automation behaviour. NetQ and external telemetry monitor path imbalance, queues, RoCE, physical health, and routing. Application dashboards track NCCL bandwidth and GPU utilization so network events can be correlated to business-relevant performance.

10. Future scaling

When rack count or GPU count approaches the two-tier radix limit, evaluate larger Spectrum generations, additional planes, or pod expansion. The decision should be based on effective bandwidth, failure domains, optics/power, and operational complexity rather than a desire to keep the topology visually simple.

Source note: See references [1], [3], [16], [27], [31].

Spectrum Networking vs Traditional Enterprise Networking

DimensionTraditional enterprise networkSpectrum / AI-scale fabric
Traffic patternUser/application, north-south heavyMachine-to-machine, east-west heavy
Endpoint speed1/10/25G common100/200/400/800G-class server/uplink links
TopologyAccess/distribution/coreClos leaf-spine, multi-rail/multiplane at scale
OversubscriptionOften high and acceptableTraining may require low oversubscription
Latency sensitivityApplication dependentCollective phases can amplify tail latency
Congestion sensitivityTCP usually masks moderate congestionCongestion can directly idle expensive GPUs
RoutingOSPF/BGP/static mix; L2 common at accessL3 BGP/ECMP underlay common
OverlayCampus segmentation or DC virtualizationEVPN/VXLAN when multi-tenancy needs it; optional for AI back-end
AutomationOften partialEssential at large scale
ObservabilityDevice/interface monitoringQueue/path/RoCE/application correlation
Failure handlingRedundant devices and protocolsParallel paths plus capacity headroom
RDMARareCommon for high-performance compute/storage

Glossary

TermBeginner-friendly definition
ACLAccess Control List; rules that permit, deny, or classify traffic.
ASICApplication-Specific Integrated Circuit; the switch silicon that forwards packets at line rate.
BFDBidirectional Forwarding Detection; a fast failure-detection protocol commonly paired with routing.
BGPBorder Gateway Protocol; scalable routing protocol widely used in data-centre leaf-spine fabrics.
Bisection bandwidthAggregate bandwidth available between two halves of a network; useful for assessing distributed-workload capacity.
ClosMulti-stage network topology that provides many parallel paths. Leaf-spine is a common two-stage Clos form.
CNPCongestion Notification Packet used in RoCE congestion-control feedback.
Cumulus LinuxNVIDIA Debian-based network operating system for Spectrum switches.
DCQCNData Center Quantized Congestion Notification; a common RoCEv2 congestion-control algorithm using ECN feedback.
DPUData Processing Unit; programmable infrastructure processor for networking, storage, and security offload.
DSX AirNVIDIA cloud-hosted data-centre simulation/digital-twin platform.
ECMPEqual-Cost Multi-Path; forwarding across multiple routes with equal routing cost.
ECNExplicit Congestion Notification; marks packets to signal congestion without dropping them.
EVPNEthernet VPN; BGP-based control plane often used with VXLAN to distribute endpoint/tenant reachability.
GPUDirect RDMATechnology enabling RDMA-capable adapters to transfer data directly to/from GPU memory.
LeafSwitch tier connected to servers or endpoints; each leaf connects upward to all spines in a classic leaf-spine fabric.
MLAGMulti-Chassis Link Aggregation; two switches present an active-active LAG to attached devices.
NCCLNVIDIA Collective Communications Library; GPU collective communications library used by distributed AI applications.
NetQNVIDIA network operations and telemetry platform for Cumulus/Spectrum environments.
NVUENVIDIA User Experience; structured configuration and operational interface used by Cumulus Linux.
PFCPriority Flow Control; Ethernet mechanism that pauses selected traffic priorities hop by hop.
RDMARemote Direct Memory Access; direct memory-to-memory data transfer with reduced CPU involvement.
RoCERDMA over Converged Ethernet.
RoCEv2Routable RoCE transport over UDP/IP.
SpineSwitch tier that interconnects all leaf switches in a leaf-spine fabric.
SpectrumNVIDIA family of high-performance Ethernet switching ASICs and switch systems.
Spectrum-XNVIDIA AI-optimized Ethernet platform combining Spectrum switches, SuperNICs, software, telemetry, and validated tuning.
Spectrum-XGSSpectrum-X scale-across technology for connecting distributed data centres into a larger AI factory.
SuperNICNVIDIA term for a network accelerator optimized for network-intensive AI workloads.
ToRTop of Rack; a switch physically located in or associated with a server rack, often acting as a leaf.
VNIVXLAN Network Identifier; identifies a logical VXLAN segment.
VRFVirtual Routing and Forwarding instance; creates separate routing tables for isolation.
VTEPVXLAN Tunnel Endpoint; device that encapsulates/decapsulates VXLAN traffic.
VXLANVirtual Extensible LAN; UDP-based overlay encapsulation used to carry tenant networks over an IP underlay.

References and Further Reading

Primary technical references are NVIDIA sources because product positioning, supported combinations, and release qualifications change quickly. Standards references should be added when turning individual sections into deeply cited blog posts.

[1] NVIDIA, “NVIDIA Spectrum-X Ethernet Networking Platform,” current product overview, accessed September 2026. https://www.nvidia.com/en-us/networking/spectrumx/

[2] NVIDIA, “Ethernet Switching for AI and the Cloud,” Spectrum Ethernet switch portfolio, accessed September 2026. https://www.nvidia.com/en-us/networking/ethernet-switching/

[3] NVIDIA Docs, “NVIDIA Spectrum-X Ethernet Networking Platform” (Kubernetes/Network Operator documentation), release 26.7 context. https://docs.nvidia.com/networking/display/kubernetes2670/spectrum-x/spectrum-x.html

[4] NVIDIA, “Networking Solutions for the Era of AI,” Ethernet, InfiniBand, and BlueField portfolio overview. https://www.nvidia.com/en-us/networking/

[5] NVIDIA Docs, “RDMA over Converged Ethernet – RoCE,” Cumulus Linux documentation. https://docs.nvidia.com/networking-ethernet-software/cumulus-linux-515/Layer-1-and-Switch-Ports/Quality-of-Service/RDMA-over-Converged-Ethernet-RoCE/

[6] NVIDIA Docs, “Quality of Service,” Cumulus Linux 5.18, including PFC considerations. https://docs.nvidia.com/networking-ethernet-software/cumulus-linux/Layer-1-and-Switch-Ports/Quality-of-Service/

[7] NVIDIA Newsroom, “NVIDIA Completes Acquisition of Mellanox,” 27 April 2020. https://nvidianews.nvidia.com/news/nvidia-completes-acquisition-of-mellanox-creating-major-force-driving-next-gen-data-centers

[8] NVIDIA Newsroom, “NVIDIA Announces Spectrum High-Performance Data Center Networking Infrastructure Platform,” 22 March 2022. https://nvidianews.nvidia.com/news/nvidia-announces-spectrum-high-performance-data-center-networking-infrastructure-platform

[9] NVIDIA Blog, “Built for Vera Rubin, NVIDIA Spectrum-6 Arrives in Gigascale AI Factories,” 21 July 2026. https://blogs.nvidia.com/blog/nvidia-spectrum-six-arrives-in-gigascale-ai-factories/

[10] NVIDIA Docs, “NVIDIA Networking Documentation,” Cumulus Linux, NetQ, and DSX Air documentation hub. https://docs.nvidia.com/networking-ethernet-software/

[11] NVIDIA Docs, “Cumulus Linux Release Versioning and Support Policy,” accessed September 2026. https://docs.nvidia.com/networking-ethernet-software/knowledge-base/Support/Support-Offerings/Cumulus-Linux-Release-Versioning-and-Support-Policy/

[12] NVIDIA Docs, “Cumulus Linux 5.15 User Guide” and current 5.x documentation. https://docs.nvidia.com/networking-ethernet-software/cumulus-linux-515/

[13] NVIDIA, “BlueField Networking Platform,” BlueField-3 and BlueField-4 overview. https://www.nvidia.com/en-us/networking/products/data-processing-unit/

[14] NVIDIA Docs, “NVIDIA DSX Air User Guide.” https://docs.nvidia.com/networking-ethernet-software/nvidia-air/

[15] NVIDIA Newsroom, “NVIDIA Introduces Spectrum-XGS Ethernet,” 22 August 2025. https://nvidianews.nvidia.com/news/nvidia-introduces-spectrum-xgs-ethernet-to-connect-distributed-data-centers-into-giga-scale-ai-super-factories

[16] NVIDIA, “Spectrum-X Validated Solution Stack,” August 2026 v2.1.6 and prior validated combinations. https://networking-docs.nvidia.com/software/spectrumx-solution-stack

[17] NVIDIA, “NVIDIA InfiniBand Adapters” and current InfiniBand platform material. https://www.nvidia.com/en-us/networking/infiniband-adapters/

[18] NVIDIA Technical Blog, “How to Connect Distributed Data Centers Into Large AI Factories with Scale-Across Networking,” September 2025. https://developer.nvidia.com/blog/how-to-connect-distributed-data-centers-into-large-ai-factories-with-scale-across-networking/

[19] NVIDIA NCCL documentation, current release documentation. https://docs.nvidia.com/deeplearning/nccl/

[20] NVIDIA Docs, Spectrum-X NIC Configuration, Network Operator 26.7. https://docs.nvidia.com/networking/display/kubernetes2670/spectrum-x/spectrum-x-configuration.html

[21] NVIDIA Docs, Adaptive Routing, NVUE 5.x. https://docs.nvidia.com/networking-ethernet-software/nvue-reference/Set-and-Unset-Commands/Adaptive-Routing/

[22] NVIDIA Docs, Cumulus Linux EVPN/VXLAN active-active mode. https://docs.nvidia.com/networking-ethernet-software/cumulus-linux/Network-Virtualization/VXLAN-Active-Active-Mode/

[23] NVIDIA Docs, Cumulus Linux EVPN inter-subnet routing. https://docs.nvidia.com/networking-ethernet-software/cumulus-linux-515/Network-Virtualization/Ethernet-Virtual-Private-Network-EVPN/Inter-subnet-Routing/

[24] NVIDIA Docs, NVUE CLI, Cumulus Linux 5.x. https://docs.nvidia.com/networking-ethernet-software/cumulus-linux-515/System-Configuration/NVIDIA-User-Experience-NVUE/NVUE-CLI/

[25] NVIDIA, “Onyx for Next-Generation Data Centers,” product page. https://www.nvidia.com/en-au/networking/ethernet-switching/onyx/

[26] NVIDIA Spectrum-6 SN6000 Ethernet Switch Systems Hardware User Manual, 2026. https://docs.nvidia.com/nvidia-spectrum-6-sn6000-ethernet-switch-systems-hardware-user-manual.pdf

[27] NVIDIA Docs, DSX Air Custom Topology and Quick Start. https://docs.nvidia.com/networking-ethernet-software/nvidia-air/Custom-Topology/

[28] NVIDIA Docs, NVIDIA NetQ 5.3 User Guide. https://docs.nvidia.com/networking-ethernet-software/cumulus-netq-53/

[29] NVIDIA Docs, Cumulus Linux configuration guidance for Ethernet Storage Fabrics, including MLAG/active-active design. https://docs.nvidia.com/networking-ethernet-software/guides/esf-generic-config-guide/

[30] NVIDIA Spectrum switch hardware manuals and Ethernet switching product tables, current generations. https://www.nvidia.com/en-us/networking/ethernet-switching/

[31] NVIDIA Docs, Spectrum-X Launch Kit / Network Operator profiles, current release documentation. https://docs.nvidia.com/networking/display/kubernetes2670/k8s-launch-kit/profiles/spectrum-x.html

[32] NVIDIA Docs, Cumulus Linux authentication, authorization, user accounts, and system configuration documentation. https://docs.nvidia.com/networking-ethernet-software/cumulus-linux/System-Configuration/Authentication-Authorization-and-Accounting/User-Accounts/

[33] NVIDIA Docs, NetQ release/version support policy and current NetQ 5.3 documentation. https://docs.nvidia.com/networking-ethernet-software/knowledge-base/Support/Support-Offerings/Cumulus-NetQ-Release-Versioning-and-Support-Policy/

Standards and complementary references

  • IEEE 802.1Q family: VLANs, Priority Flow Control, and related Ethernet bridging/QoS behaviour.
  • IETF RFC 3168: Explicit Congestion Notification (ECN).
  • IETF BGP and EVPN RFCs for routed underlays and overlays.
  • RoCE specifications from the InfiniBand Trade Association.
  • Open Compute Project and SONiC documentation for open network operating-system context.