

Table 25. InfiniBand/RDMA cheat sheet
| Term | Meaning |
| IB | InfiniBand |
| RDMA | Remote Direct Memory Access |
| HCA | Host Channel Adapter |
| QP | Queue Pair |
| SQ | Send Queue |
| RQ | Receive Queue |
| CQ | Completion Queue |
| WR | Work Request |
| WQE | Work Queue Element |
| MR | Memory Region |
| PD | Protection Domain |
| L_Key | Local memory access key |
| R_Key | Remote memory access key |
| GUID | Globally Unique Identifier |
| GID | Global Identifier |
| LID | Local Identifier |
| P_Key | Partition Key |
| SL | Service Level |
| VL | Virtual Lane |
| SM | Subnet Manager |
| MPI | Message Passing Interface |
| UCX | Unified Communication X |
| NCCL | NVIDIA Collective Communications Library |
| GPUDirect RDMA | Direct peer data path between supported GPU memory and PCIe devices such as NICs |
| SHARP | NVIDIA Scalable Hierarchical Aggregation and Reduction Protocol |
| OFED | OpenFabrics Enterprise Distribution family/packaging concept |
| Clos | Multi-stage network topology with parallel paths |
| Fat tree | Common Clos-like topology sized for high bisection bandwidth |
| Rail | Independent/semi-independent fabric path per node |
| Oversubscription | More endpoint bandwidth than upstream capacity |
| Bisection bandwidth | Capacity crossing a logical split of the fabric |
Table 26. Layer to example technology
| Layer | Examples |
| AI/HPC application | PyTorch, scientific simulation |
| Collective/message API | NCCL, MPI |
| Communication middleware | UCX, libfabric |
| RDMA API | libibverbs / rdma-core |
| Driver/kernel | mlx5 driver, Linux RDMA subsystem |
| Endpoint hardware | ConnectX HCA/SuperNIC |
| Fabric | InfiniBand switches, cables, optics |
Adaptive routing: Selecting among permitted paths based on changing fabric conditions.
AllGather: Collective in which every rank receives the data contributed by every rank.
AllReduce: Collective that combines values from all ranks and returns the result to all ranks.
Backpressure: Upstream slowing caused by lack of downstream buffer/credit availability.
Bisection bandwidth: Aggregate capacity available across a logical split of a topology.
Completion Queue (CQ): Queue through which software learns that posted work has completed.
Credit: Receiver-advertised buffer availability used by link flow control.
GID: 128-bit global identifier associated with an InfiniBand port/context.
GPUDirect RDMA: NVIDIA technology enabling supported PCIe devices to access GPU memory directly.
GUID: Globally unique identifier used to identify InfiniBand objects/ports.
HCA: Host Channel Adapter; endpoint device implementing InfiniBand transport functions.
IPoIB: IP over InfiniBand; an IP network interface carried over the IB fabric.
LID: Local Identifier used for routing within an InfiniBand subnet.
Memory Region: Memory range registered for HCA/RDMA access.
MPI: Portable programming interface for distributed-memory message passing.
NCCL: NVIDIA library for GPU collectives and point-to-point communication.
NDR: 400 Gb/s-class InfiniBand generation.
Oversubscription: Condition where offered endpoint bandwidth can exceed upstream topology capacity.
P_Key: Partition key used to control InfiniBand partition membership.
Protection Domain: Container that associates RDMA resources for protection/isolation.
Queue Pair: Send and Receive queues plus transport state used by verbs.
R_Key: Key granting permitted remote access to a registered memory region.
RDMA: Remote Direct Memory Access; direct memory-oriented communication with low CPU involvement on the fast path.
ReduceScatter: Collective that reduces data and distributes distinct result segments across ranks.
Service Level: End-to-end InfiniBand traffic classification used for QoS/path treatment.
SHARP: NVIDIA in-network collective acceleration technology.
Subnet Manager: Control entity that discovers/configures an InfiniBand subnet and programs routing.
UCX: Communication framework that selects among RDMA, GPU, shared-memory and TCP transports.
Virtual Lane: Per-link logical traffic lane with its own buffering/arbitration treatment.
XDR: 800 Gb/s-class current InfiniBand generation in NVIDIA Quantum-X800 deployments.
Is InfiniBand Layer 2 or Layer 3?
Neither analogy is perfect. InfiniBand defines its own layered architecture, including link, network/routing and transport concepts. It is better to learn its native model than force it directly into Ethernet/IP OSI labels.
Does InfiniBand use MAC addresses?
Native InfiniBand uses GUIDs, GIDs and LIDs rather than Ethernet MAC addressing.
Can InfiniBand carry IP traffic?
Yes. IP over InfiniBand (IPoIB) provides an IP interface over the fabric.
Is IPoIB the same as RDMA?
No. IPoIB carries IP packets. RDMA uses verbs/transports to access registered memory efficiently.
Do I need a Subnet Manager?
Yes for a functioning InfiniBand subnet. It discovers and configures the fabric and programs forwarding.
Can I run InfiniBand without NVIDIA hardware?
InfiniBand is an industry standard, not an NVIDIA-owned protocol. Hardware availability is a market/ecosystem question, although NVIDIA is the dominant modern supplier.
What is the difference between an HCA and a normal NIC?
An HCA natively implements InfiniBand transport/RDMA resources such as QPs and registered-memory operations.
Why does memory need to be registered?
The HCA needs stable mappings and permission to DMA into the process buffer.
What is a Queue Pair?
A Send Queue and Receive Queue with transport state; it is the core verbs endpoint object.
Is RDMA secure?
RDMA can be protected by memory keys, process permissions and fabric isolation, but security depends on the complete host/fabric design.
What is the difference between InfiniBand and RoCE?
Both can provide RDMA. InfiniBand uses an InfiniBand fabric and control model; RoCE carries RDMA over Ethernet.
What is GPUDirect RDMA?
A direct supported PCIe peer path between GPU memory and network/storage devices, reducing host-memory staging.
What is SHARP?
NVIDIA technology that moves supported collective reduction work into the InfiniBand network hierarchy.
Why do AI nodes use multiple HCAs?
To increase aggregate bandwidth, improve locality to GPU groups and provide multiple rails/paths.
Why do cable topology and GPU/NIC placement matter?
They determine hop count, PCIe/NUMA locality, bisection capacity and where congestion appears.
Is InfiniBand useful for inference?
It can be, especially for distributed or expert/tensor-parallel inference. Single-node or embarrassingly parallel inference may not need the same fabric intensity.
Can the fabric keep working if the SM fails?
Existing forwarding may continue in a stable topology, but reconfiguration is at risk; production fabrics use SM redundancy.
Is 800 Gb/s XDR automatically twice as fast as 400 Gb/s NDR for my job?
No. The workload must be network-limited and the endpoints, PCIe, topology and software must sustain the higher rate.

Figure 26. End-to-end mental model
InfiniBand is not simply a faster cable or a faster Ethernet replacement. It is a complete high-performance communication fabric whose endpoint hardware, transport model, memory semantics, routing, flow control and software ecosystem are designed around efficient communication between distributed compute resources.
The guide refers to architecture standards, current vendor documentation and upstream Linux/RDMA projects as of September 2026.
[R1] InfiniBand Trade Association – InfiniBand Architecture Specification FAQ. https://infinibandta.org/ibta-specification/
[R2] NVIDIA – Quantum-X800 InfiniBand Platform. https://www.nvidia.com/en-au/networking/products/infiniband/quantum-x800/
[R3] NVIDIA CUDA – GPUDirect RDMA documentation. https://docs.nvidia.com/cuda/gpudirect-rdma/
[R4] NVIDIA – UFM Enterprise Appliance release/documentation (2026). https://docs.nvidia.com/networking/display/ufmenterpriseapplianceswv1152/release-notes
[R5] OpenUCX documentation. https://openucx.readthedocs.io/en/master/
[R6] Linux kernel documentation – Userspace verbs access. https://docs.kernel.org/infiniband/user_verbs.html
[R7] Linux libibverbs manual – ibv_create_qp. https://man7.org/linux/man-pages/man3/ibv_create_qp.3.html
[R8] NVIDIA MLNX_OFED – IP over InfiniBand introduction. https://docs.nvidia.com/networking/display/MLNXOFEDv494080/Introduction
[R9] NVIDIA MLNX-OS – Subnet Manager. https://docs.nvidia.com/networking/display/NVIDIAMLNXOSUserManualv3112200LTS/Subnet%2BManager
[R10] NVIDIA UFM – Subnet Manager handover configuration. https://docs.nvidia.com/networking/display/ufmsdnappumv4142/subnet%2Bmanager%2Btab
[R12] NVIDIA Quantum-3 XDR switch systems user manual / Quantum-X800 platform. https://docs.nvidia.com/networking/display/nvidia-q32xx-and-q34xx-xdr-800gb-s-infiniband-switch-systems-user-manual.pdf
[R13] NVIDIA DGX SuperPOD design guide – InfiniBand cables primer. https://docs.nvidia.com/dgx-superpod/design-guide-cabling-data-centers/latest/infiniband-overview.html
[R14] NVIDIA MLNX_OFED – Common abbreviations and glossary. https://docs.nvidia.com/networking/display/mlnxofedv590560107/common%2Babbreviations%2Band%2Brelated%2Bdocuments
[R15] NVIDIA NCCL – communicator Quality of Service documentation. https://docs.nvidia.com/deeplearning/nccl/user-guide/docs/usage/communicators.html
[R16] OpenUCX FAQ – multi-rail and adaptive routing. https://openucx.readthedocs.io/en/master/faq.html
[R17] NVIDIA DOCA SDK – DOCA profiles / DOCA-OFED. https://docs.nvidia.com/doca/sdk/doca-profiles/
[R18] OpenUCX – Running UCX / container RDMA requirements. https://openucx.readthedocs.io/en/master/running.html
[R19] NVIDIA NCCL – GPUDirect RDMA topology environment variables. https://docs.nvidia.com/deeplearning/nccl/user-guide/docs/env.html
[R20] NVIDIA SHARP – current release documentation. https://docs.nvidia.com/networking/display/sharpv3103/changes-and-new-features
[R21] NVIDIA SHARP – monitoring and NCCL/Open MPI integration. https://docs.nvidia.com/networking/display/sharpv3130/SHARP-Monitoring
[R22] NVIDIA NCCL – InfiniBand networking troubleshooting. https://docs.nvidia.com/deeplearning/nccl/user-guide/docs/troubleshooting/networking_troubleshooting.html
[R23] linux-rdma perftest – InfiniBand verbs performance tests. https://github.com/linux-rdma/perftest