

Table 20. HPC workload examples
| Workload | Why the fabric matters |
| Computational fluid dynamics | Domain decomposition exchanges boundary state every iteration |
| Weather/climate modelling | Large distributed grids require frequent neighbor/global communication |
| Molecular dynamics | Many particles/forces exchanged across domain boundaries |
| Genomics | Large data movement plus distributed search/assembly stages |
| Seismic processing | Bandwidth-heavy distributed transforms and reductions |
| Computational chemistry | Tightly coupled numerical kernels and collective operations |
| EDA | Large distributed simulation/verification workloads |
| Physics simulations | Barrier-heavy multi-rank computation and reductions |
Table 21. AI workload examples
| Workload | Network sensitivity |
| Large-language-model training | Very high collective traffic; model/data/tensor parallelism can make fabric critical |
| Multimodal training | Large tensors and synchronization across GPU groups |
| Recommendation training | Can combine huge embeddings, all-to-all patterns and storage traffic |
| Distributed inference | Sensitivity varies; tensor/expert parallel inference can be network-intensive |
| AI factories | Many simultaneous jobs require performance isolation, telemetry and capacity planning |

Figure 25. Four-node dual-rail training/HPC lab
Scaling the node count does not change the basic RDMA objects, but it changes everything around them: switch radix, cable plant, path count, failure frequency, telemetry volume, routing computation, job placement and the cost of asymmetry. NVIDIA’s current Quantum-X800 Q3400 switch exposes 144 ports at 800 Gb/s and is designed for very large two-level fabrics, illustrating how high radix is used to keep tier count low [R12].
Table 24. Staged learning path
| Stage | Focus | Practical outcome |
| 1 – Concepts | Fabric, RDMA, HCA, QP, MR, LID | Explain an end-to-end data path |
| 2 – Linux RDMA stack | Drivers, rdma-core, verbs devices | Identify hardware and ports |
| 3 – Verbs/perftest | Send/Read/Write latency and bandwidth | Measure a controlled path |
| 4 – MPI/UCX | Ranks, transport selection, multi-rail | Run a small distributed job |
| 5 – NCCL/GPU | Collectives, topology, GPUDirect | Measure multi-GPU communication |
| 6 – Fabric topology | SM, routes, Clos/fat tree, rails | Read and validate a topology |
| 7 – Troubleshooting | Counters, errors, congestion, locality | Diagnose a degraded link/job |
| 8 – Large-scale design | Capacity, resiliency, operations | Produce an architecture decision record |
Without RDMA-capable hardware you can still learn the software concepts, inspect APIs and understand MPI/NCCL collectives, but you cannot faithfully reproduce HCA DMA behavior, link-level credits, real verbs latency or InfiniBand switch routing. A two-node lab with supported HCAs and one small switch is enough to make the abstractions tangible.