inifiniband-rdma-fundamentals

A beginner-friendly guide to high-performance scale-out networking

HPC and AI Environments
  1. Part 1 – Why HPC and AI Need InfiniBand and RDMA
  2. Part 2 – Inside InfiniBand and RDMA: Fabric Architecture, Memory and Data Movement
  3. Part 3 – How InfiniBand Fabrics Work: Addressing, Flow Control, Topology and Routing
  4. Part 4 – The HPC and AI Networking Stack: MPI, UCX, NCCL, GPUDirect RDMA and Collectives
  5. Part 5 – Operating a Reliable InfiniBand Fabric: Resiliency, Troubleshooting and Performance ← You are here
  6. Part 6 – InfiniBand, RoCE, NVLink, and Design Options
  7. Part 7 – From Four Nodes to Thousands of GPUs: Scaling and Learning InfiniBand
  8. Part 8 – InfiniBand & RDMA Quick Reference, Glossary, FAQ

Resiliency, Failure Behavior and Security

Network redundancy is not application fault tolerance

A fabric can provide multiple physical paths and still lose an application when a node or rank fails. Redundant links help the network survive certain link/switch failures; distributed applications must separately decide whether to retry, shrink, checkpoint or terminate.

Table 13. Failure does not necessarily mean…

FailureDoes NOT necessarily meanWhat can happen instead
One link failsEntire fabric is downSM/routing can move traffic to alternate paths if topology supports it
One SM failsPackets instantly stopExisting forwarding may continue while standby SM takes over
One HCA port failsNode is unreachableA second port/rail may remain available
One MPI rank failsNetwork failedApplication/runtime policy may abort the job
One rail degradesAll traffic stopsSoftware may use another rail, possibly at reduced performance

Figure 24. Dual-path resiliency concept

Security assumptions

Traditional HPC clusters often operate as controlled environments with trusted hosts and dedicated fabrics. P_Key partitions can limit communication groups and management keys can protect control operations, but the architecture should not be mistaken for encrypted zero-trust networking. Physical access, firmware, driver integrity, management-plane authentication and workload isolation all remain security responsibilities.

Table 14. Security/isolation mechanism versus what it provides

MechanismProvidesDoes not automatically provide
P_Key partitionLogical membership/isolation of IB communication groupsPayload encryption
Management keys / admin controlsProtection of management operationsEndpoint identity governance for every application
Dedicated fabricReduced exposure to unrelated networksProtection from compromised cluster nodes
Host OS permissions/IOMMULocal device and DMA controlsEnd-to-end application authorization
DO NOT ASSUME: A “private” HPC fabric is not automatically a secure fabric. Dedicated cabling reduces exposure but does not replace host, firmware, identity and management controls.

Key Takeaways

  • Redundant network paths reduce some failure impact, but applications need their own resilience strategy.
  • SM redundancy protects the control plane; dual rails can protect data paths and add bandwidth.
  • Partitions are isolation primitives, not encryption.

Operating and Troubleshooting an InfiniBand Fabric

What operators actually watch

  • Physical link state, negotiated speed and width
  • Port error counters and link recovery/down events
  • Topology and routing consistency
  • Cable/transceiver health and temperature
  • Congestion and port-wait/backpressure indicators
  • HCA/driver/firmware compatibility
  • Application-level bandwidth, latency and collective scaling
  • GPU-to-HCA locality and NUMA placement
  • Subnet Manager health and fabric sweeps

Common tools

Table 15. Operational tools and the question they answer

ToolQuestion it answersNotes
ibstat / ibstatusIs the HCA port present, active and at the expected state?Package/availability varies by distribution
ibv_devinfoWhat verbs devices, ports and capabilities does userspace see?Good first endpoint check
iblinkinfoHow are ports connected and what states/speeds are visible?Useful topology/link view
ibnetdiscoverWhat nodes and links are in the subnet?Topology discovery snapshot
perfqueryWhat are the port counters?Use trends, not only single values
ibdiagnetIs the fabric topology/configuration healthy?NVIDIA diagnostic suite; capabilities vary by release
ibpingCan two IB management endpoints exchange test packets?Requires server/listener arrangement
ib_write_bw / ib_read_bwWhat microbenchmark bandwidth does a verbs operation achieve?From linux-rdma perftest
ib_send_lat / ib_write_latWhat microbenchmark latency does a selected operation achieve?Synthetic benchmark, not application latency
UFMWhat is happening across the entire managed fabric?Vendor-specific telemetry/operations platform

NCCL’s current troubleshooting documentation explicitly recommends checking the Subnet Manager, port counters and ibping when diagnosing InfiniBand connectivity [R22]. The linux-rdma perftest project provides the common ib_*_bw and ib_*_lat microbenchmarks and cautions that synthetic results do not predict every real application [R23].

A beginner troubleshooting matrix

Table 16. Common symptoms, causes and first checks

SymptomLikely causesWhat to checkUseful tools
Port downCable, optics, disabled port, speed mismatchPhysical state, peer, supported speedibstat, iblinkinfo, switch CLI
Lower link speedCable/module or negotiation limitsActive vs supported speed/widthiblinkinfo, UFM, switch CLI
HCA not detectedDriver, PCIe, firmwarelspci, driver logs, verbs devicelspci, dmesg, ibv_devinfo
No LIDNo active SM or port not ACTIVESM status, subnet discoverysminfo, ibstat
Cannot communicateRouting/partition/address issueLID/path/P_Key, SM healthibping, ibnetdiscover
Low bandwidthPCIe/NUMA, message size, congestionLocality, counters, benchmark parametersib_write_bw, numactl, UFM
High latencyCPU affinity, topology hops, congestionSmall-message test, localityib_*_lat, topology tools
Poor NCCL scalingGPU/HCA affinity, rail choice, collective topologyNCCL debug/topology, fabric countersNCCL logs, nvidia-smi topo, UFM
High error countersBad cable/optic, signal integrityCounter growth and peer correlationperfquery, UFM
Routing imbalancePath algorithm/hot spot/job placementPer-port utilization and route distributionUFM, ibdiagnet
OPERATIONAL NOTE: A link being ACTIVE proves connectivity, not end-to-end health. A clean fabric must also have sensible routes, low error growth, adequate headroom and expected application performance.

A disciplined workflow

  1. Verify PCIe device and driver visibility on the endpoint.
  2. Verify HCA port state, link speed and width.
  3. Verify an active Subnet Manager and valid LID assignment.
  4. Verify topology and path reachability.
  5. Check physical/error counters and compare both ends of a link.
  6. Run small controlled verbs microbenchmarks.
  7. Check CPU/NUMA and GPU/HCA locality.
  8. Run application-level MPI/NCCL tests and correlate with fabric telemetry.

Key Takeaways

  • Troubleshoot from physical/driver state upward; do not start with application tuning when the link is unhealthy.
  • Counter trends and topology context are more useful than a single “up/down” state.
  • Microbenchmarks isolate components; application benchmarks validate the system.

Performance: What the Numbers Really Mean

The important metrics

Table 17. Performance metrics

MetricWhat it tells youWhat it does NOT tell you
One-way latencyTime from source to destination under a defined testCollective scaling or throughput
Round-trip latencyRequest/response path delayUnidirectional application behavior
Message rateHow many small operations per second can be sustainedLarge-transfer throughput
BandwidthSustained data transfer rate for a message sizeTail latency or fairness
Tail latencyWorst/near-worst completion behaviourAverage throughput
Bisection bandwidthCross-fabric capacity under partitioned trafficEndpoint PCIe capability
Collective efficiencyHow well group operations scaleRaw link rate alone
Scaling efficiencyHow much extra compute performance additional nodes deliverWhether the network is the only bottleneck

Small messages versus large messages

Small messages are dominated by per-operation overhead, queue processing and latency. Large messages are dominated by sustained DMA, PCIe, link and memory bandwidth. A fabric can be excellent at one and merely adequate at the other, so benchmark the message-size distribution that resembles the workload.

NUMA and PCIe placement

A dual-socket server is itself a small topology. If the application thread allocates memory on one NUMA node while the HCA sits under the other socket, every transfer may cross the CPU interconnect. GPU systems add another layer because GPU-to-HCA distance affects GPUDirect performance. Modern perftest versions explicitly include NUMA-aware binding features for this reason [R23].

Communication/computation overlap

High-performance software tries to overlap communication with useful computation so the network is not exposed as pure idle time. This is why bandwidth, latency and GPU kernel scheduling must be evaluated together. A theoretically faster link may deliver little benefit if the workload already hides communication behind computation; conversely, a barrier-heavy job exposes every microsecond.

PERFORMANCE NOTE: Do not optimize a synthetic peak number in isolation. The target is application time-to-solution or training throughput with stable tail behavior.

Key Takeaways

  • Different message sizes stress different parts of the stack.
  • NUMA/PCIe/GPU locality can erase the advantage of an otherwise fast fabric.
  • Application scaling efficiency is the final performance metric that matters.