inifiniband-rdma-fundamentals

A beginner-friendly guide to high-performance scale-out networking

HPC and AI Environments
  1. Part 1 – Why HPC and AI Need InfiniBand and RDMA
  2. Part 2 – Inside InfiniBand and RDMA: Fabric Architecture, Memory and Data Movement
  3. Part 3 – How InfiniBand Fabrics Work: Addressing, Flow Control, Topology and Routing ← You are here
  4. Part 4 – The HPC and AI Networking Stack: MPI, UCX, NCCL, GPUDirect RDMA and Collectives
  5. Part 5 – Operating a Reliable InfiniBand Fabric: Resiliency, Troubleshooting and Performance
  6. Part 6 – InfiniBand, RoCE, NVLink, and Design Options
  7. Part 7 – From Four Nodes to Thousands of GPUs: Scaling and Learning InfiniBand
  8. Part 8 – InfiniBand & RDMA Quick Reference, Glossary, FAQ

Addressing, Subnets and the Subnet Manager

GUID, GID and LID

Figure 9. Identity hierarchy

Table 9. GUID versus GID versus LID

IdentifierScope/purposeBeginner mental model
GUIDGlobally unique hardware/port identityWho is this object?
GID128-bit global identifier associated with a port/contextHow can this endpoint be globally identified?
LIDLocal Identifier assigned within an IB subnetWhich local destination should the switches forward toward?

A Local Identifier is assigned by the Subnet Manager and is used for forwarding inside the subnet. NVIDIA’s terminology reference defines the LID as unique within the subnet and used to direct packets locally [R14]. A GID is 128 bits and incorporates a subnet prefix plus interface identity. These constructs are not simply MAC and IP addresses with different names.

Partitions and P_Keys

InfiniBand partitions use P_Keys to control which endpoints may communicate within logical groups. Partitions are useful for administrative separation and workload isolation, but they should not be confused with encryption or a complete zero-trust security model.

Why a Subnet Manager is required

The SM discovers fabric devices, assigns local identifiers, calculates routes and programs configuration needed for traffic to flow. NVIDIA describes the SM as a central entity that discovers and configures the fabric, including routing, QoS and partitioning [R9]. OpenSM is the widely used open-source implementation; UFM integrates management around these functions.

Figure 10. Fabric bring-up and discovery

What if the Subnet Manager goes away?

A stable, already-programmed fabric can often continue forwarding existing traffic if the active SM disappears because switches already hold forwarding state. The risk is loss of management and reconfiguration capability: new devices, failed links or topology changes may not be handled correctly until an SM is available again. Production designs therefore commonly use master/standby SM arrangements. UFM documents SM priority and handover behavior for this purpose [R10].

DO NOT ASSUME: The Subnet Manager is not a central forwarding appliance. Application packets normally travel directly through the programmed switches.

Key Takeaways

  • LID is the key local forwarding identity inside an InfiniBand subnet.
  • P_Keys create partitions but do not encrypt traffic.
  • The SM is essential for discovery and route programming; redundancy matters because failures still require reconfiguration.

How InfiniBand Packets Flow: Virtual Lanes, Credits and QoS

Hop-by-hop forwarding

Figure 11. Simplified hop-by-hop packet flow

Within a subnet, switches use forwarding information associated with destination identifiers and routing policy programmed by the SM. The exact packet headers and routing modes are more detailed than the diagram, but the important mental model is distributed forwarding: each switch makes a local output-port decision using fabric state.

Credit-based link flow control

InfiniBand uses receiver-based credit mechanisms so a sender does not transmit into a virtual lane unless the next hop has advertised enough buffering. This is fundamentally different from relying on routine packet drops as the main congestion signal. Backpressure propagates when downstream buffers fill.

Figure 12. Credit-based flow control

PERFORMANCE NOTE: Lossless link behaviour prevents a class of packet-loss problems; it does not remove queueing, contention, backpressure or congestion. A lossless fabric can still perform badly.

Service Levels and Virtual Lanes

A Service Level (SL) is a traffic-classification concept carried end to end; the fabric maps SLs to Virtual Lanes (VLs) on individual links. VLs provide separate buffering/arbitration resources that can reduce interference between traffic classes. NVIDIA’s NCCL documentation notes that InfiniBand QoS uses SLs mapped to VLs, with behavior defined by the subnet-manager configuration [R15].

Table 10. Service Level versus Virtual Lane

ConceptScopePurpose
Service Level (SL)End-to-end traffic classificationExpress desired QoS class/path treatment
Virtual Lane (VL)Per-link logical lane/buffer/arbitration resourceProvide traffic separation and scheduling on a physical link

Head-of-line blocking

If unrelated traffic shares the same buffers and one destination becomes congested, packets behind it can be delayed even when their own path is free. VL separation and careful routing/QoS design help limit this effect, but too many traffic classes create operational complexity.

Key Takeaways

  • InfiniBand switches forward packets hop by hop using programmed fabric state.
  • Credits protect receiver buffers and create backpressure instead of normal-drop behavior.
  • SLs classify traffic; VLs provide link-local separation and arbitration.

Fabric Topology, Routing and Congestion

From a small switch to a Clos/fat-tree

Figure 13. Small single-switch fabric

Figure 14. Two-tier non-blocking-style fat tree / Clos

A Clos or fat-tree fabric uses multiple parallel paths through upper-tier switches. “Non-blocking” means the topology has enough cross-sectional capacity to support a defined traffic model without intentional oversubscription. It does not guarantee perfect application scaling because endpoints, routing and workload synchronization still matter.

Oversubscription and bisection bandwidth

Table 11. Blocking versus non-blocking design

DesignMeaningBenefitCost/risk
Non-blocking / full bisectionUplink capacity broadly matches downstream demand for the design traffic modelPredictable all-to-all performanceMore switch ports, cables, optics and power
OversubscribedDownstream endpoint bandwidth exceeds uplink capacityLower cost and complexityContention under simultaneous east-west load

A 32-node conceptual example

Suppose 32 nodes each have one 400 Gb/s fabric port. The endpoint edge represents 12.8 Tb/s of nominal injection bandwidth. A conceptually non-blocking design must provide enough leaf uplink and spine capacity that any balanced partition of the 32 nodes can exchange traffic without an intentional bottleneck. The exact number of leaf/spine ports depends on switch radix and how ports are broken out, but the design principle is simple: every downlink unit needs a corresponding path budget toward the opposite half of the fabric.

Multi-rail and rail-optimized fabrics

Figure 15. Simplified dual-rail GPU cluster

Multi-rail designs give a node more than one independent or semi-independent network path. They can increase aggregate bandwidth, improve path diversity and map communication more closely to GPU/NUMA topology. They also multiply cabling and troubleshooting complexity. UCX supports multi-rail selection and provides controls for choosing devices and adaptive routing [R16].

Routing and adaptive routing

Deterministic routing selects paths from a precomputed rule. Adaptive routing can choose among allowed alternatives based on fabric conditions, helping avoid hot spots. The routing algorithm must preserve correctness properties while balancing load; vendor implementations add their own telemetry and congestion mechanisms. Quantum-X800 advertises adaptive routing and telemetry-based congestion control as NVIDIA-specific platform features [R2].

Figure 16. Congestion hot spot

Key Takeaways

  • Topology determines the physical ceiling for cluster scaling.
  • Oversubscription is a deliberate capacity trade-off, not a protocol property.
  • Multi-rail and adaptive routing can improve bandwidth and resilience but add operational complexity.