inifiniband-rdma-fundamentals

A beginner-friendly guide to high-performance scale-out networking

HPC and AI Environments
  1. Part 1 – Why HPC and AI Need InfiniBand and RDMA ← You are here
  2. Part 2 – Inside InfiniBand and RDMA: Fabric Architecture, Memory and Data Movement
  3. Part 3 – How InfiniBand Fabrics Work: Addressing, Flow Control, Topology and Routing
  4. Part 4 – The HPC and AI Networking Stack: MPI, UCX, NCCL, GPUDirect RDMA and Collectives
  5. Part 5 – Operating a Reliable InfiniBand Fabric: Resiliency, Troubleshooting and Performance
  6. Part 6 – InfiniBand, RoCE, NVLink, and Design Options
  7. Part 7 – From Four Nodes to Thousands of GPUs: Scaling and Learning InfiniBand
  8. Part 8 – InfiniBand & RDMA Quick Reference, Glossary, FAQ

The one-sentence summary

InfiniBand is a specialized switched fabric for moving data between distributed compute resources with very low latency and high throughput; RDMA lets applications arrange much of that movement directly between registered memory regions with far less CPU and kernel involvement on the fast path.

Why HPC and AI Need a Different Kind of Network

The one-sentence mental model

InfiniBand is a specialized high-performance switched fabric designed for efficient communication among CPUs, GPUs, storage systems and other devices. Remote Direct Memory Access (RDMA) is the communication model that allows software to arrange transfers directly between registered memory regions, reducing repeated kernel crossings, CPU work and intermediate copies on the fast path.

BEGINNER SHORTCUT: Think of InfiniBand as the road system and RDMA as a way of moving cargo from an approved loading bay at one endpoint to an approved loading bay at another without making every packet stop at the normal operating-system networking desk.

Figure 1. Conventional socket path versus an RDMA-oriented fast path

The drawing is deliberately simplified. RDMA does not make the operating system magically disappear: drivers, device permissions, memory registration, resource creation, connection setup and teardown still require software and often kernel involvement. Linux documents this aspect directly: resource-management operations travel through the userspace-verbs interface, while many fast-path operations can interact with mapped hardware resources without a system call on each operation [R6].

Why ordinary enterprise networking is not always enough

Many enterprise applications are loosely coupled. A web server can wait a millisecond for a database response and still deliver a perfectly acceptable user experience. A tightly coupled simulation or distributed training job is different: thousands of workers repeatedly exchange state and then wait for one another. A small delay multiplied across many synchronization points becomes lost accelerator time.

AI training creates a particularly visible version of the problem. A single server may contain eight or more GPUs connected by an internal accelerator fabric. Once a model spans multiple servers, gradients, parameters or activation data must cross the scale-out network. The expensive GPUs can only remain busy if communication completes quickly enough to overlap with computation.

Table 1. Enterprise application network versus HPC/AI compute fabric

CharacteristicEnterprise application networkHPC/AI compute fabric
Primary traffic patternClient/server, north-south plus mixed east-westHeavy east-west, often all-to-all or collective
Latency sensitivityUsually milliseconds are acceptable for many appsMicroseconds can materially affect scaling
SynchronizationOften asynchronous or loosely coupledFrequent barriers and collectives
CPU overheadUsually acceptable within normal sockets stackOften worth aggressively reducing
Message sizesMixed; often request/responseTiny control messages through very large tensor/data transfers
Performance goalGood aggregate service throughputPredictable low latency plus sustained per-node bandwidth
Failure modelRetries at application/service layers commonA single slow or failed rank can stall a parallel job

Bandwidth is necessary, but not sufficient

A 400 Gb/s or 800 Gb/s link sounds enormous, but link rate alone does not answer the important question: can the application move the right messages, at the right size, with low enough latency, while many peers communicate simultaneously? Small-message rate, tail latency, PCIe placement, switch topology, routing and congestion can matter as much as nominal bandwidth.

PERFORMANCE NOTE: Theoretical link rate is not application throughput. Encoding, protocol headers, flow control, PCIe effects, memory subsystem limits, message size and software overhead all reduce or reshape the number an application observes.

Key Takeaways

  • HPC and distributed AI are tightly coupled systems; communication delay directly consumes compute time.
  • RDMA reduces fast-path CPU/kernel work, but does not eliminate the OS or control software.
  • Low latency, message rate and predictable behaviour matter alongside raw bandwidth.

History, Vocabulary and the Big Picture

A short history

InfiniBand emerged around the turn of the millennium from an industry effort to create a high-performance, channel-based switched interconnect rather than continuing to stretch shared-bus I/O designs. The InfiniBand Trade Association (IBTA) defines the architecture and still describes InfiniBand as an industry-standard, channel-based switched fabric for server and storage connectivity [R1].

Figure 2. InfiniBand and RDMA evolution in context

Early adoption was strongest in HPC because scientific workloads were already limited by message passing and cluster I/O. Over time, faster generations, mature verbs libraries, MPI integration and accelerator support expanded the role. Mellanox became the most visible commercial InfiniBand supplier and was acquired by NVIDIA in 2020. In today’s AI infrastructure, InfiniBand is widely used for scale-out GPU communication; IBTA explicitly points to large distributed AI training as a major modern use case [R1].

InfiniBand, RDMA, verbs, RoCE and iWARP

Table 2. InfiniBand and related terms

TermWhat it isBeginner mental model
InfiniBand (IB)A complete switched interconnect architecture and link/transport ecosystemThe specialized fabric
RDMAA memory-oriented communication capability/modelHow endpoints can move data efficiently
VerbsProgramming operations and APIs for RDMA resources and work requestsThe low-level control vocabulary
RoCEv2RDMA carried over routable Ethernet/IP/UDPRDMA on an Ethernet fabric
iWARPRDMA over a TCP/IP transportRDMA using TCP semantics
Ethernet/TCP-IPGeneral-purpose networking family and conventional transport stackThe normal data-center network

InfiniBand and RDMA are closely associated, but they are not synonyms. InfiniBand natively provides RDMA transports. RoCE provides RDMA semantics over Ethernet. Applications usually consume these capabilities through libraries such as Message Passing Interface (MPI), Unified Communication X (UCX) or NVIDIA Collective Communications Library (NCCL) rather than issuing raw verbs directly.

COMMON MISTAKE: “RDMA fabric” describes a capability, not necessarily an InfiniBand fabric. Always ask what link/network technology carries the RDMA traffic.

Key Takeaways

  • InfiniBand is a network/fabric architecture; RDMA is a communication capability.
  • Verbs expose low-level RDMA resources and operations.
  • RoCE and iWARP show that RDMA is not exclusive to InfiniBand.