More GPUs. Same wall clock.
You doubled the allocation. The job did not get twice as fast. Nsight shows fat all-reduce bars, hollow SM gaps, and a data loader that still owns every step boundary. The cluster is fine. The schedule and the serial fraction are not.
HPC performance is not one trick. It is three questions in order: what is the theoretical ceiling, are you scaling the problem or just the hardware, and which roof are you actually hitting — memory, compute, or the network.
Two numbers
1. Amdahl · max speedup = 1/s
If 10% of the work never parallelizes, you cap at 10× even with infinite ranks. Cutting serial work from 10% → 5% doubles the ceiling. Buying another rack does not.
2. Overlap · compute should stay dense while collectives run
Blocking all-reduce puts the interconnect on the critical path. Async collectives (DDP buckets, MPI_Iallreduce) stack communication under the next layer’s backward pass — same idea as CUDA stream pipelines, one level up.
Flip the instruments until those feel mechanical.
The serial ceiling
Fixed problem size. More processors. How high can speedup go?
Strong vs weak scaling
Same cluster. Two different questions.
- Strong — fixed global problem. Wall time should fall toward 1/p.
- Weak — fixed work per rank. Wall time should stay flat.
Hide the network
Distributed training dies on communication latency long before single-kernel FLOPs run out. The backward pass is friendly: finish layer N grads, start all-reduce, keep computing layer N−1.
Layer N backward: [==compute==] Layer N all-reduce: [===comm===] Layer N-1 backward: [==compute==] ← overlaps
# PyTorch DDP does bucketing + async all-reduce for you. # Manual MPI sketch of the same idea: req = comm.Iallreduce(grads_n) # non-blocking grads_nm1 = layer_nm1.backward(...) # compute while network works req.Wait()
Ideal: collective time fully hidden. That only happens when per-layer compute ≥ all-reduce time for that layer’s gradients.
Which roof are you under?
The roofline puts a kernel between peak bandwidth and peak FLOPs using operational intensity (FLOPs / byte). Most DL non-GEMM ops are memory-bound. Large matmuls sit under the compute roof. Multi-node jobs add a third roof: collectives.
For an interactive A100/H100 ridge plot, see HBM memory.
Load imbalance is a fake serial fraction
Even perfect math dies if ranks finish at different times:
- Variable-length sequences without bucketing
- Mixed GPU SKUs (fast ranks wait at every barrier)
- Pipeline bubbles at fill and drain
Static partitions work when cost is uniform. Otherwise you need dynamic scheduling, sequence packing, or topology-aware placement — treat stragglers as serial work you invented.
Profile first
| Tool | Job |
|---|---|
| Nsight Systems | Timeline: gaps, H2D, kernels, NCCL |
| Nsight Compute | One kernel: occupancy, stalls, roofline |
| VTune / perf | Host: loader, threads, cache |
nsys profile --trace=cuda,nvtx,osrt,nccl python train.py ncu --target-processes all -k regex:gemm python train.py
Five minutes of Systems often beats five days of kernel micro-opts on the wrong 2% bar.
What to do
- Measure the serial + sync fraction — if s is large, more GPUs are a tax.
- Pick strong or weak on purpose — report efficiency at the rank count you will run.
- Overlap collectives — async all-reduce / DDP buckets; do not block every layer.
- Name the roof — memory → fuse & cut traffic; compute → tensor cores & precision; network → overlap & topology.
- Profile before tuning — Systems for gaps, Compute for why, one change at a time.
# Cheap first pass: is the GPU even busy? nvidia-smi dmon -s u # Then the timeline that tells the truth nsys profile -o report --trace=cuda,nvtx,osrt,nccl python train.py
Further reading
- Roofline model (Berkeley Lab)
- NVIDIA Nsight Systems
- Scalability — But at What COST? — McSherry et al.
- HPC Wiki — Performance profiling
Related concepts
Master GPU memory hierarchy from registers to global memory, understand coalescing patterns, bank conflicts, and optimization strategies for maximum performance
OpenMP parallel programming: fork-join model, scheduling, data races, false sharing, NUMA thread affinity, and GPU offloading.
Why DataLoader num_workers matters: processes hide load latency behind GPU work, how to find the sweet spot, and the memory/GIL pitfalls that come with the pool.
Learn a profiler-first Python optimization workflow: measure bottlenecks, choose the right lever, and verify performance changes.
Explore CPU pipeline stages, instruction-level parallelism, pipeline hazards, and branch prediction through interactive visualizations.
Master pipeline hazards through interactive visualizations of data dependencies, control hazards, structural conflicts, and advanced detection mechanisms.
