inifiniband-rdma-fundamentals

A beginner-friendly guide to high-performance scale-out networking

HPC and AI Environments
  1. Part 1 – Why HPC and AI Need InfiniBand and RDMA
  2. Part 2 – Inside InfiniBand and RDMA: Fabric Architecture, Memory and Data Movement ← You are here
  3. Part 3 – How InfiniBand Fabrics Work: Addressing, Flow Control, Topology and Routing
  4. Part 4 – The HPC and AI Networking Stack: MPI, UCX, NCCL, GPUDirect RDMA and Collectives
  5. Part 5 – Operating a Reliable InfiniBand Fabric: Resiliency, Troubleshooting and Performance
  6. Part 6 – InfiniBand, RoCE, NVLink, and Design Options
  7. Part 7 – From Four Nodes to Thousands of GPUs: Scaling and Learning InfiniBand
  8. Part 8 – InfiniBand & RDMA Quick Reference, Glossary, FAQ

Inside an InfiniBand Fabric: Architecture, Components and Speeds

High-level architecture

Figure 3. Basic two-tier InfiniBand fabric

Compute nodes attach through Host Channel Adapters (HCAs). Switches forward packets across paths programmed by the subnet-management function. The Subnet Manager (SM) discovers the topology, assigns local identifiers, calculates routes and configures fabric policy. NVIDIA Unified Fabric Manager (UFM) can incorporate an SM and adds monitoring, telemetry and operational tooling; it is a vendor-specific management product rather than part of the base InfiniBand standard [R4].

What the major components do

Table 3. Fabric components

ComponentRoleAnalogyImportant note
HCATerminates InfiniBand links and executes transport/RDMA functionsHigh-performance network adapter plus transport engineUsually connected through PCIe; locality to CPU/GPU matters
InfiniBand switchForwards IB packets between portsFabric junctionHigh radix (port count) reduces the number of tiers needed
Switch ASICImplements high-speed forwarding, buffering and often advanced functionsTraffic engine inside the switchVendor capabilities vary
PCIe connectionConnects HCA to host I/O hierarchyOn-ramp from server to fabricCan bottleneck or add NUMA penalties
Direct Attach Copper (DAC)/Active Optical Cable (AOC)/opticsPhysical media between portsCabling plantDistance, thermals and signal integrity matter
Subnet ManagerDiscovers/configures subnet and routesFabric controller / routing engineNot normally in the data path
UFM / management toolingMonitors and operates the fabricNOC/operations platformVendor-specific, optional but useful at scale

HCA versus a conventional NIC

Table 4. HCA versus conventional Ethernet NIC

AspectHCA (InfiniBand context)Conventional Ethernet NIC
Native transport modelInfiniBand transports and RDMA resourcesEthernet frames; often TCP/IP via host stack
Memory registrationCore RDMA conceptNot required for ordinary socket traffic
Queue pairsCentral verbs abstractionTX/RX rings exist but exposed differently
AddressingGUID/GID/LID and subnet conceptsMAC/IP addresses
Kernel bypass potentialCommon on fast path through verbs librariesPossible with specialized frameworks, not normal sockets default
Primary optimizationLow-latency high-throughput endpoint communicationGeneral-purpose interoperability

Link generations and naming

InfiniBand names link generations rather than using only Ethernet-style speed labels. Historically the sequence has included SDR, DDR, QDR, FDR, EDR, HDR, NDR and now XDR. Modern NVIDIA Quantum-X800 systems provide 800 Gb/s XDR-class connectivity, while NDR systems provide 400 Gb/s-class ports [R2][R12].

Table 5. InfiniBand generations (rounded nominal 4-lane link class)

GenerationNominal classContext
Single Data Rate (SDR)10 Gb/s raw classOriginal generation; historical
Double Data Rate (DDR)20 Gb/s raw classHistorical
Quad Data Rate (QDR)40 Gb/s classHistorical but still encountered
Fourteen Data Rate (FDR)56 Gb/s classIntroduced 64b/66b era
Enhanced Data Rate (EDR)100 Gb/s classWidely deployed HPC generation
High Data Rate (HDR)200 Gb/s classModern HPC/early large AI era
Next Data Rate (NDR)400 Gb/s classCurrent high-performance deployments
eXtended Data Rate (XDR)800 Gb/s class per 4x link (and up to 1.6 Tb/s per 8x link)Current 2026 high-end generation; NVIDIA Quantum-X800
DO NOT ASSUME: The name on the port is not the payload bandwidth seen by an MPI rank or GPU. Also check lane width, PCIe generation, cable type, protocol overhead and the workload’s message sizes.

Cables and lane widths

InfiniBand uses multiple serial lanes to build a link. Common generations have used 1X, 2X and 4X arrangements depending on product and speed. NVIDIA’s cabling guidance describes this multi-lane design and the evolution from copper CX4 through Quad Small Form-factor Pluggable (QSFP)-family and OSFP form factors [R13]. At current speeds, cabling is part of the architecture: reach, connector density, optics power and serviceability affect rack and row design.

Key Takeaways

  • An HCA is an active transport endpoint, not merely a passive cable adapter.
  • The Subnet Manager configures the fabric but normally does not forward application data.
  • InfiniBand speed names are generations; application throughput is always lower and workload-dependent.

RDMA from First Principles: Memory, Protection and Queues

Why memory must be registered

RDMA hardware needs permission and stable address translations for memory it will access. Software therefore registers a memory region (MR), establishing which virtual-address range the HCA may use and what operations are permitted. Linux documents that direct userspace I/O requires target memory to remain resident and that the kernel accounts for pinned memory and process limits [R6]. Modern implementations may use optimizations such as on-demand paging, registration caches or DMA-BUF integration, but the beginner model remains: the HCA cannot DMA into arbitrary process memory.

Figure 4. Registered memory and protection relationship

The core objects

Table 6. Core RDMA objects

ObjectMeaningWhy it exists
Protection Domain (PD)Container tying related resources into a protection boundaryPrevents unrelated QPs/MRs from being mixed accidentally
Memory Region (MR)Registered memory rangeDefines DMA-accessible memory and permissions
L_KeyLocal access key for a registered memory regionAuthorizes local HCA access
R_KeyRemote access key exposed for permitted one-sided operationsAuthorizes remote RDMA access to a region
Queue Pair (QP)Send Queue plus Receive Queue and transport stateMain endpoint abstraction for posting work
Work Request (WR)Software request describing an operationTells the HCA what to do
Work Queue Element (WQE)Hardware-consumable representation of posted workEntry processed by the HCA
Completion Queue (CQ)Queue of completion notificationsLets software learn that work completed

Queue Pair anatomy

Figure 5. Queue Pair and Completion Queue

The Send Queue is used for operations initiated by the local endpoint. The Receive Queue supplies buffers for two-sided receives. The Completion Queue reports finished work. A QP is created inside a protection domain; Linux libibverbs documents this association explicitly [R7].

COMMON MISTAKE: A “Queue Pair” is not a pair of hosts. It is a pair of work queues associated with an endpoint transport context.

Key Takeaways

  • Registered memory defines what the HCA may DMA to or from.
  • Keys are permissions/capabilities associated with registered memory, not encryption keys.
  • Queue Pairs and Completion Queues form the execution model behind verbs operations.

Moving Data: Send/Receive, RDMA Read, RDMA Write and Transports

Two-sided Send/Receive

Send/Receive is called two-sided because both endpoints participate explicitly: the receiver posts receive buffers, and the sender posts a Send. This resembles message passing and is often a natural fit for control traffic or protocols where both sides need clear message boundaries.

Figure 6. Two-sided Send/Receive

RDMA Write

RDMA Write places local data into a remote registered memory region. The remote CPU does not need to execute a receive operation for every transfer. The remote side must still have arranged and authorized the target memory beforehand, normally exchanging address and R_Key information through a control protocol.

Figure 7. One-sided RDMA Write

RDMA Read

RDMA Read pulls bytes from a remote registered region into local registered memory. The requesting endpoint initiates the operation and gets a completion when its local transfer is complete. As with Write, the remote side must previously have authorized access.

Figure 8. One-sided RDMA Read

Atomics

RDMA atomic operations provide hardware-assisted read-modify-write primitives such as compare-and-swap or fetch-and-add on supported transports and memory. They are powerful building blocks for distributed coordination, but they are not a replacement for an application-level consistency design.

Table 7. Send/Receive versus one-sided operations

OperationRemote receive required per transfer?Common strengthImportant caution
Send/ReceiveYesMessage-oriented protocols and explicit peer participationReceiver must keep receive buffers posted
RDMA WriteNoPush bulk data with low remote CPU involvementRemote memory and R_Key must be exchanged safely
RDMA ReadNoPull data on demandRequester controls timing and may create read pressure
AtomicsNo explicit receiveSmall synchronization primitivesHardware/support/scaling characteristics vary

Transport services

Table 8. Common InfiniBand transport services

TransportConnection modelReliabilityBeginner use
RC – Reliable ConnectedConnection-oriented QP to QPReliable, orderedMost common conceptual model for one-sided RDMA
UC – Unreliable ConnectedConnection-orientedNo end-to-end retransmission guaranteeSpecialized; less commonly discussed in beginner deployments
UD – Unreliable DatagramConnectionless datagramsUnreliableScalable datagrams, management-style patterns; IPoIB commonly uses UD

Reliable Connected (RC) is the easiest transport for a beginner to associate with RDMA Read/Write because it provides a reliable connected relationship. Unreliable Datagram (UD) is useful where connectionless scaling matters. IP over InfiniBand (IPoIB) commonly uses UD by default and can provide a normal-looking IP interface over the fabric [R8].

Key Takeaways

  • Two-sided operations require explicit receive participation; one-sided operations act on pre-authorized remote memory.
  • “One-sided” does not mean “no protocol” or “no remote software.” Setup and memory exchange still matter.
  • RC and UD solve different scaling and semantics problems.