

The one-sentence summary
InfiniBand is a specialized switched fabric for moving data between distributed compute resources with very low latency and high throughput; RDMA lets applications arrange much of that movement directly between registered memory regions with far less CPU and kernel involvement on the fast path.
InfiniBand is a specialized high-performance switched fabric designed for efficient communication among CPUs, GPUs, storage systems and other devices. Remote Direct Memory Access (RDMA) is the communication model that allows software to arrange transfers directly between registered memory regions, reducing repeated kernel crossings, CPU work and intermediate copies on the fast path.
| BEGINNER SHORTCUT: Think of InfiniBand as the road system and RDMA as a way of moving cargo from an approved loading bay at one endpoint to an approved loading bay at another without making every packet stop at the normal operating-system networking desk. |

Figure 1. Conventional socket path versus an RDMA-oriented fast path
The drawing is deliberately simplified. RDMA does not make the operating system magically disappear: drivers, device permissions, memory registration, resource creation, connection setup and teardown still require software and often kernel involvement. Linux documents this aspect directly: resource-management operations travel through the userspace-verbs interface, while many fast-path operations can interact with mapped hardware resources without a system call on each operation [R6].
Many enterprise applications are loosely coupled. A web server can wait a millisecond for a database response and still deliver a perfectly acceptable user experience. A tightly coupled simulation or distributed training job is different: thousands of workers repeatedly exchange state and then wait for one another. A small delay multiplied across many synchronization points becomes lost accelerator time.
AI training creates a particularly visible version of the problem. A single server may contain eight or more GPUs connected by an internal accelerator fabric. Once a model spans multiple servers, gradients, parameters or activation data must cross the scale-out network. The expensive GPUs can only remain busy if communication completes quickly enough to overlap with computation.
Table 1. Enterprise application network versus HPC/AI compute fabric
| Characteristic | Enterprise application network | HPC/AI compute fabric |
| Primary traffic pattern | Client/server, north-south plus mixed east-west | Heavy east-west, often all-to-all or collective |
| Latency sensitivity | Usually milliseconds are acceptable for many apps | Microseconds can materially affect scaling |
| Synchronization | Often asynchronous or loosely coupled | Frequent barriers and collectives |
| CPU overhead | Usually acceptable within normal sockets stack | Often worth aggressively reducing |
| Message sizes | Mixed; often request/response | Tiny control messages through very large tensor/data transfers |
| Performance goal | Good aggregate service throughput | Predictable low latency plus sustained per-node bandwidth |
| Failure model | Retries at application/service layers common | A single slow or failed rank can stall a parallel job |
A 400 Gb/s or 800 Gb/s link sounds enormous, but link rate alone does not answer the important question: can the application move the right messages, at the right size, with low enough latency, while many peers communicate simultaneously? Small-message rate, tail latency, PCIe placement, switch topology, routing and congestion can matter as much as nominal bandwidth.
| PERFORMANCE NOTE: Theoretical link rate is not application throughput. Encoding, protocol headers, flow control, PCIe effects, memory subsystem limits, message size and software overhead all reduce or reshape the number an application observes. |
InfiniBand emerged around the turn of the millennium from an industry effort to create a high-performance, channel-based switched interconnect rather than continuing to stretch shared-bus I/O designs. The InfiniBand Trade Association (IBTA) defines the architecture and still describes InfiniBand as an industry-standard, channel-based switched fabric for server and storage connectivity [R1].

Figure 2. InfiniBand and RDMA evolution in context
Early adoption was strongest in HPC because scientific workloads were already limited by message passing and cluster I/O. Over time, faster generations, mature verbs libraries, MPI integration and accelerator support expanded the role. Mellanox became the most visible commercial InfiniBand supplier and was acquired by NVIDIA in 2020. In today’s AI infrastructure, InfiniBand is widely used for scale-out GPU communication; IBTA explicitly points to large distributed AI training as a major modern use case [R1].
Table 2. InfiniBand and related terms
| Term | What it is | Beginner mental model |
| InfiniBand (IB) | A complete switched interconnect architecture and link/transport ecosystem | The specialized fabric |
| RDMA | A memory-oriented communication capability/model | How endpoints can move data efficiently |
| Verbs | Programming operations and APIs for RDMA resources and work requests | The low-level control vocabulary |
| RoCEv2 | RDMA carried over routable Ethernet/IP/UDP | RDMA on an Ethernet fabric |
| iWARP | RDMA over a TCP/IP transport | RDMA using TCP semantics |
| Ethernet/TCP-IP | General-purpose networking family and conventional transport stack | The normal data-center network |
InfiniBand and RDMA are closely associated, but they are not synonyms. InfiniBand natively provides RDMA transports. RoCE provides RDMA semantics over Ethernet. Applications usually consume these capabilities through libraries such as Message Passing Interface (MPI), Unified Communication X (UCX) or NVIDIA Collective Communications Library (NCCL) rather than issuing raw verbs directly.
| COMMON MISTAKE: “RDMA fabric” describes a capability, not necessarily an InfiniBand fabric. Always ask what link/network technology carries the RDMA traffic. |