inifiniband-rdma-fundamentals

A beginner-friendly guide to high-performance scale-out networking

HPC and AI Environments
  1. Part 1 – Why HPC and AI Need InfiniBand and RDMA
  2. Part 2 – Inside InfiniBand and RDMA: Fabric Architecture, Memory and Data Movement
  3. Part 3 – How InfiniBand Fabrics Work: Addressing, Flow Control, Topology and Routing
  4. Part 4 – The HPC and AI Networking Stack: MPI, UCX, NCCL, GPUDirect RDMA and Collectives ← You are here
  5. Part 5 – Operating a Reliable InfiniBand Fabric: Resiliency, Troubleshooting and Performance
  6. Part 6 – InfiniBand, RoCE, NVLink, and Design Options
  7. Part 7 – From Four Nodes to Thousands of GPUs: Scaling and Learning InfiniBand
  8. Part 8 – InfiniBand & RDMA Quick Reference, Glossary, FAQ

The HPC Software Stack: MPI, UCX, Verbs and OFED

From application to fabric

Figure 17. Typical HPC communication software stack

MPI (Message Passing Interface) is the application-facing programming model used by many HPC codes. Implementations such as Open MPI or MPICH can use Unified Communication X (UCX) and other transports underneath. UCX provides a higher-level communication framework that automatically selects among RDMA, shared memory, GPU transports and TCP, hiding much of the low-level verbs complexity [R5].

rdma-core and libibverbs

On mainstream Linux, rdma-core provides the core userspace RDMA libraries and tools, including libibverbs. The kernel exposes userspace-verbs interfaces, while device-specific drivers implement hardware support. This split is why an application can use a common verbs API across different RDMA-capable devices even though hardware backends differ [R6].

OFED, MLNX_OFED and DOCA-OFED

OFED historically refers to OpenFabrics Enterprise Distribution packaging around the RDMA ecosystem. In modern Linux, many components are upstream through kernel and rdma-core packages. NVIDIA also packages vendor-qualified driver stacks. In 2026 NVIDIA documentation describes DOCA-OFED as the driver-only networking profile equivalent to its MLNX_OFED driver/tool set [R17].

Table 12. MPI, UCX, NCCL and verbs

Layer/toolPrimary audienceWhat it abstracts
Message Passing Interface (MPI)HPC application developerRanks, messages, collectives
Unified Communication X (UCX)Communication middleware / MPI runtimeTransport selection, memory types, RDMA/shared-memory paths
NVIDIA Collective Communications Library (NCCL)GPU/AI framework/runtimeGPU collectives and point-to-point communication
libibverbsLow-level systems programmerPDs, MRs, QPs, CQs, WRs and device capabilities

Containers do not erase RDMA requirements

A container still needs access to RDMA device nodes, drivers, libraries and adequate locked-memory limits. UCX documentation explicitly calls out /dev/infiniband devices and memory-lock considerations for containerized RDMA [R18]. Kubernetes environments often use device plugins/operators to make these resources manageable rather than handing devices to containers manually.

Key Takeaways

  • MPI is usually the application API; UCX is often the transport-selection middleware; verbs is the low-level RDMA API.
  • The Linux RDMA stack spans userspace libraries, kernel interfaces and device drivers.
  • Containerization changes packaging and device exposure, not the fundamental RDMA hardware requirements.

AI Networking: NCCL, NVLink, InfiniBand and GPUDirect RDMA

Scale-up versus scale-out

Modern GPU servers use multiple interconnect domains. Inside a system, PCIe, NVLink and NVSwitch connect GPUs, CPUs and I/O devices. Between systems, InfiniBand or an Ethernet/RoCE fabric provides scale-out connectivity. The two layers cooperate; one does not automatically replace the other.

Figure 18. Scale-up versus scale-out networking

What NCCL does

NCCL provides high-performance GPU collectives such as AllReduce, AllGather and ReduceScatter as well as point-to-point primitives. It discovers available topology and network plugins so frameworks such as PyTorch can communicate efficiently without implementing the fabric details themselves.

Why GPUDirect RDMA matters

Without a direct GPU-to-I/O path, network transfers may require staging through host memory, adding copies and CPU/memory-system activity. NVIDIA GPUDirect RDMA enables supported peer devices such as network adapters to exchange data directly with GPU memory over PCIe, subject to platform and topology constraints [R3].

Figure 19. GPU communication without GPUDirect RDMA

Figure 20. GPU communication with GPUDirect RDMA

Locality still matters

A direct path is not automatically a short path. GPU, HCA and CPU socket placement determine which PCIe switches or inter-socket links a transfer crosses. NCCL exposes controls describing whether GPUDirect RDMA is permitted at different GPU-to-NIC distance classes; this is a practical reminder that PHB/PIX/PXB/SYS topology can change performance [R19].

Figure 21. GPU/HCA locality example

COMMON MISTAKE: GPUDirect RDMA does not mean the remote GPUs are directly cabled to each other. The HCA and switched network remain in the path.

Key Takeaways

  • NVLink/NVSwitch are mainly scale-up interconnects; InfiniBand is a scale-out fabric.
  • NCCL turns topology-aware GPU communication into usable collective operations for AI frameworks.
  • GPUDirect RDMA can avoid host-memory staging, but PCIe/NUMA placement still matters.

Collective Communication and In-Network Computing

Why collectives dominate many parallel workloads

A collective is a communication operation involving a group rather than one sender and one receiver. Distributed training repeatedly aggregates gradients or redistributes tensors; HPC algorithms perform similar global exchanges. The algorithm used for a collective determines how many network steps and bytes cross each link.

Figure 22. Collective communication patterns

AllReduce, ReduceScatter and AllGather

AllReduce combines values from all ranks and returns the result to all ranks. Efficient implementations frequently decompose this work into ReduceScatter followed by AllGather, using ring, tree or topology-aware algorithms. The best strategy depends on message size, rank count, link topology and accelerator generation.

NVIDIA SHARP as a vendor-specific example

Some modern InfiniBand platforms can perform portions of collective reduction inside the network. NVIDIA Scalable Hierarchical Aggregation and Reduction Protocol (SHARP) is a vendor-specific technology that offloads supported collective operations into the network hierarchy. NVIDIA documents current SHARP support for XDR topologies and integration with Open MPI and NCCL [R20][R21].

Figure 23. Traditional endpoint reduction versus in-network reduction

VENDOR-SPECIFIC: SHARP, UFM and NVIDIA adaptive-routing implementations are NVIDIA technologies. The guide uses them as concrete examples, not as definitions of base InfiniBand.

Key Takeaways

  • Collectives are first-class performance events in HPC and AI, not just many independent point-to-point transfers.
  • Algorithm and topology interact: rings, trees and hierarchical methods behave differently.
  • In-network reduction can reduce endpoint/network work for supported operations but does not accelerate every workload.