inifiniband-rdma-fundamentals

A beginner-friendly guide to high-performance scale-out networking

HPC and AI Environments
  1. Part 1 – Why HPC and AI Need InfiniBand and RDMA
  2. Part 2 – Inside InfiniBand and RDMA: Fabric Architecture, Memory and Data Movement
  3. Part 3 – How InfiniBand Fabrics Work: Addressing, Flow Control, Topology and Routing
  4. Part 4 – The HPC and AI Networking Stack: MPI, UCX, NCCL, GPUDirect RDMA and Collectives
  5. Part 5 – Operating a Reliable InfiniBand Fabric: Resiliency, Troubleshooting and Performance
  6. Part 6 – InfiniBand, RoCE, NVLink, and Design Options
  7. Part 7 – From Four Nodes to Thousands of GPUs: Scaling and Learning InfiniBand ← You are here
  8. Part 8 – InfiniBand & RDMA Quick Reference, Glossary, FAQ

Use Cases, a Small Cluster Walkthrough and Scaling Up

Common HPC workloads

Table 20. HPC workload examples

WorkloadWhy the fabric matters
Computational fluid dynamicsDomain decomposition exchanges boundary state every iteration
Weather/climate modellingLarge distributed grids require frequent neighbor/global communication
Molecular dynamicsMany particles/forces exchanged across domain boundaries
GenomicsLarge data movement plus distributed search/assembly stages
Seismic processingBandwidth-heavy distributed transforms and reductions
Computational chemistryTightly coupled numerical kernels and collective operations
EDALarge distributed simulation/verification workloads
Physics simulationsBarrier-heavy multi-rank computation and reductions

Common AI workloads

Table 21. AI workload examples

WorkloadNetwork sensitivity
Large-language-model trainingVery high collective traffic; model/data/tensor parallelism can make fabric critical
Multimodal trainingLarge tensors and synchronization across GPU groups
Recommendation trainingCan combine huge embeddings, all-to-all patterns and storage traffic
Distributed inferenceSensitivity varies; tensor/expert parallel inference can be network-intensive
AI factoriesMany simultaneous jobs require performance isolation, telemetry and capacity planning

A four-node walkthrough

Figure 25. Four-node dual-rail training/HPC lab

  1. The server boots and enumerates the HCA on PCIe.
  2. The Linux HCA/RDMA driver exposes verbs devices.
  3. Physical InfiniBand links train and enter an operational state.
  4. The active Subnet Manager discovers nodes and switches, assigns/validates identifiers and programs routes.
  5. MPI, UCX or NCCL starts and discovers available transports/devices.
  6. The runtime allocates and registers communication buffers.
  7. Queue Pairs and related resources are created; endpoints exchange connection/memory metadata as required.
  8. Work Requests are posted. The HCA executes DMA and sends InfiniBand packets through the programmed fabric.
  9. Remote HCA(s) place or return data according to the operation.
  10. Completion Queue entries tell software which work has completed.
  11. The application proceeds to the next compute/communication phase.

From four nodes to thousands of GPUs

Scaling the node count does not change the basic RDMA objects, but it changes everything around them: switch radix, cable plant, path count, failure frequency, telemetry volume, routing computation, job placement and the cost of asymmetry. NVIDIA’s current Quantum-X800 Q3400 switch exposes 144 ports at 800 Gb/s and is designed for very large two-level fabrics, illustrating how high radix is used to keep tier count low [R12].

  • Topology symmetry becomes an operational requirement, not an aesthetic preference.
  • A single mis-cabled link can create asymmetric path length or reduced capacity.
  • Firmware/driver versions must be qualified as a cluster combination.
  • Schedulers increasingly need topology awareness so large jobs are placed on nearby healthy resources.
  • Telemetry must be correlated with job IDs and time windows, not viewed only as device health.
  • Change management matters because a “small” routing or firmware change can affect thousands of accelerators.

Key Takeaways

  • HPC and AI use different algorithms, but both benefit when communication is fast and predictable.
  • The four-node workflow already contains the same core objects found at large scale.
  • At massive scale, operations, topology and change control become as important as link speed.

Hands-on learning roadmap

Table 24. Staged learning path

StageFocusPractical outcome
1 – ConceptsFabric, RDMA, HCA, QP, MR, LIDExplain an end-to-end data path
2 – Linux RDMA stackDrivers, rdma-core, verbs devicesIdentify hardware and ports
3 – Verbs/perftestSend/Read/Write latency and bandwidthMeasure a controlled path
4 – MPI/UCXRanks, transport selection, multi-railRun a small distributed job
5 – NCCL/GPUCollectives, topology, GPUDirectMeasure multi-GPU communication
6 – Fabric topologySM, routes, Clos/fat tree, railsRead and validate a topology
7 – TroubleshootingCounters, errors, congestion, localityDiagnose a degraded link/job
8 – Large-scale designCapacity, resiliency, operationsProduce an architecture decision record

Without RDMA-capable hardware you can still learn the software concepts, inspect APIs and understand MPI/NCCL collectives, but you cannot faithfully reproduce HCA DMA behavior, link-level credits, real verbs latency or InfiniBand switch routing. A two-node lab with supported HCAs and one small switch is enough to make the abstractions tangible.