


Figure 3. Basic two-tier InfiniBand fabric
Compute nodes attach through Host Channel Adapters (HCAs). Switches forward packets across paths programmed by the subnet-management function. The Subnet Manager (SM) discovers the topology, assigns local identifiers, calculates routes and configures fabric policy. NVIDIA Unified Fabric Manager (UFM) can incorporate an SM and adds monitoring, telemetry and operational tooling; it is a vendor-specific management product rather than part of the base InfiniBand standard [R4].
Table 3. Fabric components
| Component | Role | Analogy | Important note |
| HCA | Terminates InfiniBand links and executes transport/RDMA functions | High-performance network adapter plus transport engine | Usually connected through PCIe; locality to CPU/GPU matters |
| InfiniBand switch | Forwards IB packets between ports | Fabric junction | High radix (port count) reduces the number of tiers needed |
| Switch ASIC | Implements high-speed forwarding, buffering and often advanced functions | Traffic engine inside the switch | Vendor capabilities vary |
| PCIe connection | Connects HCA to host I/O hierarchy | On-ramp from server to fabric | Can bottleneck or add NUMA penalties |
| Direct Attach Copper (DAC)/Active Optical Cable (AOC)/optics | Physical media between ports | Cabling plant | Distance, thermals and signal integrity matter |
| Subnet Manager | Discovers/configures subnet and routes | Fabric controller / routing engine | Not normally in the data path |
| UFM / management tooling | Monitors and operates the fabric | NOC/operations platform | Vendor-specific, optional but useful at scale |
Table 4. HCA versus conventional Ethernet NIC
| Aspect | HCA (InfiniBand context) | Conventional Ethernet NIC |
| Native transport model | InfiniBand transports and RDMA resources | Ethernet frames; often TCP/IP via host stack |
| Memory registration | Core RDMA concept | Not required for ordinary socket traffic |
| Queue pairs | Central verbs abstraction | TX/RX rings exist but exposed differently |
| Addressing | GUID/GID/LID and subnet concepts | MAC/IP addresses |
| Kernel bypass potential | Common on fast path through verbs libraries | Possible with specialized frameworks, not normal sockets default |
| Primary optimization | Low-latency high-throughput endpoint communication | General-purpose interoperability |
InfiniBand names link generations rather than using only Ethernet-style speed labels. Historically the sequence has included SDR, DDR, QDR, FDR, EDR, HDR, NDR and now XDR. Modern NVIDIA Quantum-X800 systems provide 800 Gb/s XDR-class connectivity, while NDR systems provide 400 Gb/s-class ports [R2][R12].
Table 5. InfiniBand generations (rounded nominal 4-lane link class)
| Generation | Nominal class | Context |
| Single Data Rate (SDR) | 10 Gb/s raw class | Original generation; historical |
| Double Data Rate (DDR) | 20 Gb/s raw class | Historical |
| Quad Data Rate (QDR) | 40 Gb/s class | Historical but still encountered |
| Fourteen Data Rate (FDR) | 56 Gb/s class | Introduced 64b/66b era |
| Enhanced Data Rate (EDR) | 100 Gb/s class | Widely deployed HPC generation |
| High Data Rate (HDR) | 200 Gb/s class | Modern HPC/early large AI era |
| Next Data Rate (NDR) | 400 Gb/s class | Current high-performance deployments |
| eXtended Data Rate (XDR) | 800 Gb/s class per 4x link (and up to 1.6 Tb/s per 8x link) | Current 2026 high-end generation; NVIDIA Quantum-X800 |
| DO NOT ASSUME: The name on the port is not the payload bandwidth seen by an MPI rank or GPU. Also check lane width, PCIe generation, cable type, protocol overhead and the workload’s message sizes. |
InfiniBand uses multiple serial lanes to build a link. Common generations have used 1X, 2X and 4X arrangements depending on product and speed. NVIDIA’s cabling guidance describes this multi-lane design and the evolution from copper CX4 through Quad Small Form-factor Pluggable (QSFP)-family and OSFP form factors [R13]. At current speeds, cabling is part of the architecture: reach, connector density, optics power and serviceability affect rack and row design.
RDMA hardware needs permission and stable address translations for memory it will access. Software therefore registers a memory region (MR), establishing which virtual-address range the HCA may use and what operations are permitted. Linux documents that direct userspace I/O requires target memory to remain resident and that the kernel accounts for pinned memory and process limits [R6]. Modern implementations may use optimizations such as on-demand paging, registration caches or DMA-BUF integration, but the beginner model remains: the HCA cannot DMA into arbitrary process memory.

Figure 4. Registered memory and protection relationship
Table 6. Core RDMA objects
| Object | Meaning | Why it exists |
| Protection Domain (PD) | Container tying related resources into a protection boundary | Prevents unrelated QPs/MRs from being mixed accidentally |
| Memory Region (MR) | Registered memory range | Defines DMA-accessible memory and permissions |
| L_Key | Local access key for a registered memory region | Authorizes local HCA access |
| R_Key | Remote access key exposed for permitted one-sided operations | Authorizes remote RDMA access to a region |
| Queue Pair (QP) | Send Queue plus Receive Queue and transport state | Main endpoint abstraction for posting work |
| Work Request (WR) | Software request describing an operation | Tells the HCA what to do |
| Work Queue Element (WQE) | Hardware-consumable representation of posted work | Entry processed by the HCA |
| Completion Queue (CQ) | Queue of completion notifications | Lets software learn that work completed |

Figure 5. Queue Pair and Completion Queue
The Send Queue is used for operations initiated by the local endpoint. The Receive Queue supplies buffers for two-sided receives. The Completion Queue reports finished work. A QP is created inside a protection domain; Linux libibverbs documents this association explicitly [R7].
| COMMON MISTAKE: A “Queue Pair” is not a pair of hosts. It is a pair of work queues associated with an endpoint transport context. |
Send/Receive is called two-sided because both endpoints participate explicitly: the receiver posts receive buffers, and the sender posts a Send. This resembles message passing and is often a natural fit for control traffic or protocols where both sides need clear message boundaries.

Figure 6. Two-sided Send/Receive
RDMA Write places local data into a remote registered memory region. The remote CPU does not need to execute a receive operation for every transfer. The remote side must still have arranged and authorized the target memory beforehand, normally exchanging address and R_Key information through a control protocol.

Figure 7. One-sided RDMA Write
RDMA Read pulls bytes from a remote registered region into local registered memory. The requesting endpoint initiates the operation and gets a completion when its local transfer is complete. As with Write, the remote side must previously have authorized access.

Figure 8. One-sided RDMA Read
RDMA atomic operations provide hardware-assisted read-modify-write primitives such as compare-and-swap or fetch-and-add on supported transports and memory. They are powerful building blocks for distributed coordination, but they are not a replacement for an application-level consistency design.
Table 7. Send/Receive versus one-sided operations
| Operation | Remote receive required per transfer? | Common strength | Important caution |
| Send/Receive | Yes | Message-oriented protocols and explicit peer participation | Receiver must keep receive buffers posted |
| RDMA Write | No | Push bulk data with low remote CPU involvement | Remote memory and R_Key must be exchanged safely |
| RDMA Read | No | Pull data on demand | Requester controls timing and may create read pressure |
| Atomics | No explicit receive | Small synchronization primitives | Hardware/support/scaling characteristics vary |
Table 8. Common InfiniBand transport services
| Transport | Connection model | Reliability | Beginner use |
| RC – Reliable Connected | Connection-oriented QP to QP | Reliable, ordered | Most common conceptual model for one-sided RDMA |
| UC – Unreliable Connected | Connection-oriented | No end-to-end retransmission guarantee | Specialized; less commonly discussed in beginner deployments |
| UD – Unreliable Datagram | Connectionless datagrams | Unreliable | Scalable datagrams, management-style patterns; IPoIB commonly uses UD |
Reliable Connected (RC) is the easiest transport for a beginner to associate with RDMA Read/Write because it provides a reliable connected relationship. Unreliable Datagram (UD) is useful where connectionless scaling matters. IP over InfiniBand (IPoIB) commonly uses UD by default and can provide a normal-looking IP interface over the fabric [R8].