


Figure 17. Typical HPC communication software stack
MPI (Message Passing Interface) is the application-facing programming model used by many HPC codes. Implementations such as Open MPI or MPICH can use Unified Communication X (UCX) and other transports underneath. UCX provides a higher-level communication framework that automatically selects among RDMA, shared memory, GPU transports and TCP, hiding much of the low-level verbs complexity [R5].
On mainstream Linux, rdma-core provides the core userspace RDMA libraries and tools, including libibverbs. The kernel exposes userspace-verbs interfaces, while device-specific drivers implement hardware support. This split is why an application can use a common verbs API across different RDMA-capable devices even though hardware backends differ [R6].
OFED historically refers to OpenFabrics Enterprise Distribution packaging around the RDMA ecosystem. In modern Linux, many components are upstream through kernel and rdma-core packages. NVIDIA also packages vendor-qualified driver stacks. In 2026 NVIDIA documentation describes DOCA-OFED as the driver-only networking profile equivalent to its MLNX_OFED driver/tool set [R17].
Table 12. MPI, UCX, NCCL and verbs
| Layer/tool | Primary audience | What it abstracts |
| Message Passing Interface (MPI) | HPC application developer | Ranks, messages, collectives |
| Unified Communication X (UCX) | Communication middleware / MPI runtime | Transport selection, memory types, RDMA/shared-memory paths |
| NVIDIA Collective Communications Library (NCCL) | GPU/AI framework/runtime | GPU collectives and point-to-point communication |
| libibverbs | Low-level systems programmer | PDs, MRs, QPs, CQs, WRs and device capabilities |
A container still needs access to RDMA device nodes, drivers, libraries and adequate locked-memory limits. UCX documentation explicitly calls out /dev/infiniband devices and memory-lock considerations for containerized RDMA [R18]. Kubernetes environments often use device plugins/operators to make these resources manageable rather than handing devices to containers manually.
Modern GPU servers use multiple interconnect domains. Inside a system, PCIe, NVLink and NVSwitch connect GPUs, CPUs and I/O devices. Between systems, InfiniBand or an Ethernet/RoCE fabric provides scale-out connectivity. The two layers cooperate; one does not automatically replace the other.

Figure 18. Scale-up versus scale-out networking
NCCL provides high-performance GPU collectives such as AllReduce, AllGather and ReduceScatter as well as point-to-point primitives. It discovers available topology and network plugins so frameworks such as PyTorch can communicate efficiently without implementing the fabric details themselves.
Without a direct GPU-to-I/O path, network transfers may require staging through host memory, adding copies and CPU/memory-system activity. NVIDIA GPUDirect RDMA enables supported peer devices such as network adapters to exchange data directly with GPU memory over PCIe, subject to platform and topology constraints [R3].

Figure 19. GPU communication without GPUDirect RDMA

Figure 20. GPU communication with GPUDirect RDMA
A direct path is not automatically a short path. GPU, HCA and CPU socket placement determine which PCIe switches or inter-socket links a transfer crosses. NCCL exposes controls describing whether GPUDirect RDMA is permitted at different GPU-to-NIC distance classes; this is a practical reminder that PHB/PIX/PXB/SYS topology can change performance [R19].

Figure 21. GPU/HCA locality example
| COMMON MISTAKE: GPUDirect RDMA does not mean the remote GPUs are directly cabled to each other. The HCA and switched network remain in the path. |
A collective is a communication operation involving a group rather than one sender and one receiver. Distributed training repeatedly aggregates gradients or redistributes tensors; HPC algorithms perform similar global exchanges. The algorithm used for a collective determines how many network steps and bytes cross each link.

Figure 22. Collective communication patterns
AllReduce combines values from all ranks and returns the result to all ranks. Efficient implementations frequently decompose this work into ReduceScatter followed by AllGather, using ring, tree or topology-aware algorithms. The best strategy depends on message size, rank count, link topology and accelerator generation.
Some modern InfiniBand platforms can perform portions of collective reduction inside the network. NVIDIA Scalable Hierarchical Aggregation and Reduction Protocol (SHARP) is a vendor-specific technology that offloads supported collective operations into the network hierarchy. NVIDIA documents current SHARP support for XDR topologies and integration with Open MPI and NCCL [R20][R21].

Figure 23. Traditional endpoint reduction versus in-network reduction
| VENDOR-SPECIFIC: SHARP, UFM and NVIDIA adaptive-routing implementations are NVIDIA technologies. The guide uses them as concrete examples, not as definitions of base InfiniBand. |