inifiniband-rdma-fundamentals

A beginner-friendly guide to high-performance scale-out networking

HPC and AI Environments
  1. Part 1 – Why HPC and AI Need InfiniBand and RDMA
  2. Part 2 – Inside InfiniBand and RDMA: Fabric Architecture, Memory and Data Movement
  3. Part 3 – How InfiniBand Fabrics Work: Addressing, Flow Control, Topology and Routing
  4. Part 4 – The HPC and AI Networking Stack: MPI, UCX, NCCL, GPUDirect RDMA and Collectives
  5. Part 5 – Operating a Reliable InfiniBand Fabric: Resiliency, Troubleshooting and Performance
  6. Part 6 – InfiniBand, RoCE, NVLink, and Design Options
  7. Part 7 – From Four Nodes to Thousands of GPUs: Scaling and Learning InfiniBand
  8. Part 8 – InfiniBand & RDMA Quick Reference, Glossary, FAQ ← You are here

Quick Reference, Glossary, FAQ

Quick reference cheat sheet

Table 25. InfiniBand/RDMA cheat sheet

TermMeaning
IBInfiniBand
RDMARemote Direct Memory Access
HCAHost Channel Adapter
QPQueue Pair
SQSend Queue
RQReceive Queue
CQCompletion Queue
WRWork Request
WQEWork Queue Element
MRMemory Region
PDProtection Domain
L_KeyLocal memory access key
R_KeyRemote memory access key
GUIDGlobally Unique Identifier
GIDGlobal Identifier
LIDLocal Identifier
P_KeyPartition Key
SLService Level
VLVirtual Lane
SMSubnet Manager
MPIMessage Passing Interface
UCXUnified Communication X
NCCLNVIDIA Collective Communications Library
GPUDirect RDMADirect peer data path between supported GPU memory and PCIe devices such as NICs
SHARPNVIDIA Scalable Hierarchical Aggregation and Reduction Protocol
OFEDOpenFabrics Enterprise Distribution family/packaging concept
ClosMulti-stage network topology with parallel paths
Fat treeCommon Clos-like topology sized for high bisection bandwidth
RailIndependent/semi-independent fabric path per node
OversubscriptionMore endpoint bandwidth than upstream capacity
Bisection bandwidthCapacity crossing a logical split of the fabric

Table 26. Layer to example technology

LayerExamples
AI/HPC applicationPyTorch, scientific simulation
Collective/message APINCCL, MPI
Communication middlewareUCX, libfabric
RDMA APIlibibverbs / rdma-core
Driver/kernelmlx5 driver, Linux RDMA subsystem
Endpoint hardwareConnectX HCA/SuperNIC
FabricInfiniBand switches, cables, optics

Glossary

Adaptive routing: Selecting among permitted paths based on changing fabric conditions.

AllGather: Collective in which every rank receives the data contributed by every rank.

AllReduce: Collective that combines values from all ranks and returns the result to all ranks.

Backpressure: Upstream slowing caused by lack of downstream buffer/credit availability.

Bisection bandwidth: Aggregate capacity available across a logical split of a topology.

Completion Queue (CQ): Queue through which software learns that posted work has completed.

Credit: Receiver-advertised buffer availability used by link flow control.

GID: 128-bit global identifier associated with an InfiniBand port/context.

GPUDirect RDMA: NVIDIA technology enabling supported PCIe devices to access GPU memory directly.

GUID: Globally unique identifier used to identify InfiniBand objects/ports.

HCA: Host Channel Adapter; endpoint device implementing InfiniBand transport functions.

IPoIB: IP over InfiniBand; an IP network interface carried over the IB fabric.

LID: Local Identifier used for routing within an InfiniBand subnet.

Memory Region: Memory range registered for HCA/RDMA access.

MPI: Portable programming interface for distributed-memory message passing.

NCCL: NVIDIA library for GPU collectives and point-to-point communication.

NDR: 400 Gb/s-class InfiniBand generation.

Oversubscription: Condition where offered endpoint bandwidth can exceed upstream topology capacity.

P_Key: Partition key used to control InfiniBand partition membership.

Protection Domain: Container that associates RDMA resources for protection/isolation.

Queue Pair: Send and Receive queues plus transport state used by verbs.

R_Key: Key granting permitted remote access to a registered memory region.

RDMA: Remote Direct Memory Access; direct memory-oriented communication with low CPU involvement on the fast path.

ReduceScatter: Collective that reduces data and distributes distinct result segments across ranks.

Service Level: End-to-end InfiniBand traffic classification used for QoS/path treatment.

SHARP: NVIDIA in-network collective acceleration technology.

Subnet Manager: Control entity that discovers/configures an InfiniBand subnet and programs routing.

UCX: Communication framework that selects among RDMA, GPU, shared-memory and TCP transports.

Virtual Lane: Per-link logical traffic lane with its own buffering/arbitration treatment.

XDR: 800 Gb/s-class current InfiniBand generation in NVIDIA Quantum-X800 deployments.

Frequently Asked Questions

Is InfiniBand Layer 2 or Layer 3?

Neither analogy is perfect. InfiniBand defines its own layered architecture, including link, network/routing and transport concepts. It is better to learn its native model than force it directly into Ethernet/IP OSI labels.

Does InfiniBand use MAC addresses?

Native InfiniBand uses GUIDs, GIDs and LIDs rather than Ethernet MAC addressing.

Can InfiniBand carry IP traffic?

Yes. IP over InfiniBand (IPoIB) provides an IP interface over the fabric.

Is IPoIB the same as RDMA?

No. IPoIB carries IP packets. RDMA uses verbs/transports to access registered memory efficiently.

Do I need a Subnet Manager?

Yes for a functioning InfiniBand subnet. It discovers and configures the fabric and programs forwarding.

Can I run InfiniBand without NVIDIA hardware?

InfiniBand is an industry standard, not an NVIDIA-owned protocol. Hardware availability is a market/ecosystem question, although NVIDIA is the dominant modern supplier.

What is the difference between an HCA and a normal NIC?

An HCA natively implements InfiniBand transport/RDMA resources such as QPs and registered-memory operations.

Why does memory need to be registered?

The HCA needs stable mappings and permission to DMA into the process buffer.

What is a Queue Pair?

A Send Queue and Receive Queue with transport state; it is the core verbs endpoint object.

Is RDMA secure?

RDMA can be protected by memory keys, process permissions and fabric isolation, but security depends on the complete host/fabric design.

What is the difference between InfiniBand and RoCE?

Both can provide RDMA. InfiniBand uses an InfiniBand fabric and control model; RoCE carries RDMA over Ethernet.

What is GPUDirect RDMA?

A direct supported PCIe peer path between GPU memory and network/storage devices, reducing host-memory staging.

What is SHARP?

NVIDIA technology that moves supported collective reduction work into the InfiniBand network hierarchy.

Why do AI nodes use multiple HCAs?

To increase aggregate bandwidth, improve locality to GPU groups and provide multiple rails/paths.

Why do cable topology and GPU/NIC placement matter?

They determine hop count, PCIe/NUMA locality, bisection capacity and where congestion appears.

Is InfiniBand useful for inference?

It can be, especially for distributed or expert/tensor-parallel inference. Single-node or embarrassingly parallel inference may not need the same fabric intensity.

Can the fabric keep working if the SM fails?

Existing forwarding may continue in a stable topology, but reconfiguration is at risk; production fabrics use SM redundancy.

Is 800 Gb/s XDR automatically twice as fast as 400 Gb/s NDR for my job?

No. The workload must be network-limited and the endpoints, PCIe, topology and software must sustain the higher rate.

The final big idea

Figure 26. End-to-end mental model

InfiniBand is not simply a faster cable or a faster Ethernet replacement. It is a complete high-performance communication fabric whose endpoint hardware, transport model, memory semantics, routing, flow control and software ecosystem are designed around efficient communication between distributed compute resources.

If you remember only these things…

  • InfiniBand is the fabric; RDMA is the memory-oriented communication capability.
  • The HCA is an active transport engine connected to host memory through PCIe.
  • RDMA fast paths depend on registered memory, Queue Pairs and Completion Queues.
  • The Subnet Manager discovers/configures routes; it is not normally in the data path.
  • Credit-based flow control reduces drops but does not eliminate congestion.
  • Topology and bisection bandwidth matter as much as port speed.
  • MPI/UCX serve HPC; NCCL serves GPU collectives; all ultimately rely on the hardware path beneath them.
  • GPUDirect RDMA avoids host-memory staging when supported, but locality still matters.
  • NVLink is mainly scale-up; InfiniBand is scale-out.
  • At large scale, operations, telemetry and change management are part of network performance.

References and Further Reading

The guide refers to architecture standards, current vendor documentation and upstream Linux/RDMA projects as of September 2026.

[R1] InfiniBand Trade Association – InfiniBand Architecture Specification FAQ. https://infinibandta.org/ibta-specification/

[R2] NVIDIA – Quantum-X800 InfiniBand Platform. https://www.nvidia.com/en-au/networking/products/infiniband/quantum-x800/

[R3] NVIDIA CUDA – GPUDirect RDMA documentation. https://docs.nvidia.com/cuda/gpudirect-rdma/

[R4] NVIDIA – UFM Enterprise Appliance release/documentation (2026). https://docs.nvidia.com/networking/display/ufmenterpriseapplianceswv1152/release-notes

[R5] OpenUCX documentation. https://openucx.readthedocs.io/en/master/

[R6] Linux kernel documentation – Userspace verbs access. https://docs.kernel.org/infiniband/user_verbs.html

[R7] Linux libibverbs manual – ibv_create_qp. https://man7.org/linux/man-pages/man3/ibv_create_qp.3.html

[R8] NVIDIA MLNX_OFED – IP over InfiniBand introduction. https://docs.nvidia.com/networking/display/MLNXOFEDv494080/Introduction

[R9] NVIDIA MLNX-OS – Subnet Manager. https://docs.nvidia.com/networking/display/NVIDIAMLNXOSUserManualv3112200LTS/Subnet%2BManager

[R10] NVIDIA UFM – Subnet Manager handover configuration. https://docs.nvidia.com/networking/display/ufmsdnappumv4142/subnet%2Bmanager%2Btab

[R12] NVIDIA Quantum-3 XDR switch systems user manual / Quantum-X800 platform. https://docs.nvidia.com/networking/display/nvidia-q32xx-and-q34xx-xdr-800gb-s-infiniband-switch-systems-user-manual.pdf

[R13] NVIDIA DGX SuperPOD design guide – InfiniBand cables primer. https://docs.nvidia.com/dgx-superpod/design-guide-cabling-data-centers/latest/infiniband-overview.html

[R14] NVIDIA MLNX_OFED – Common abbreviations and glossary. https://docs.nvidia.com/networking/display/mlnxofedv590560107/common%2Babbreviations%2Band%2Brelated%2Bdocuments

[R15] NVIDIA NCCL – communicator Quality of Service documentation. https://docs.nvidia.com/deeplearning/nccl/user-guide/docs/usage/communicators.html

[R16] OpenUCX FAQ – multi-rail and adaptive routing. https://openucx.readthedocs.io/en/master/faq.html

[R17] NVIDIA DOCA SDK – DOCA profiles / DOCA-OFED. https://docs.nvidia.com/doca/sdk/doca-profiles/

[R18] OpenUCX – Running UCX / container RDMA requirements. https://openucx.readthedocs.io/en/master/running.html

[R19] NVIDIA NCCL – GPUDirect RDMA topology environment variables. https://docs.nvidia.com/deeplearning/nccl/user-guide/docs/env.html

[R20] NVIDIA SHARP – current release documentation. https://docs.nvidia.com/networking/display/sharpv3103/changes-and-new-features

[R21] NVIDIA SHARP – monitoring and NCCL/Open MPI integration. https://docs.nvidia.com/networking/display/sharpv3130/SHARP-Monitoring

[R22] NVIDIA NCCL – InfiniBand networking troubleshooting. https://docs.nvidia.com/deeplearning/nccl/user-guide/docs/troubleshooting/networking_troubleshooting.html

[R23] linux-rdma perftest – InfiniBand verbs performance tests. https://github.com/linux-rdma/perftest