Skip to main content

HPC Performance Optimization: Scaling, Profiling, and Tuning

Summary
Amdahl and Gustafson ceilings, strong vs weak scaling, roofline bounds, and hiding all-reduce behind compute — the levers that decide whether more GPUs actually buy science.

More GPUs. Same wall clock.

You doubled the allocation. The job did not get twice as fast. Nsight shows fat all-reduce bars, hollow SM gaps, and a data loader that still owns every step boundary. The cluster is fine. The schedule and the serial fraction are not.

HPC performance is not one trick. It is three questions in order: what is the theoretical ceiling, are you scaling the problem or just the hardware, and which roof are you actually hitting — memory, compute, or the network.

Two numbers

1. Amdahl · max speedup = 1/s
If 10% of the work never parallelizes, you cap at 10× even with infinite ranks. Cutting serial work from 10% → 5% doubles the ceiling. Buying another rack does not.

2. Overlap · compute should stay dense while collectives run
Blocking all-reduce puts the interconnect on the critical path. Async collectives (DDP buckets, MPI_Iallreduce) stack communication under the next layer’s backward pass — same idea as CUDA stream pipelines, one level up.

Flip the instruments until those feel mechanical.

The serial ceiling

Fixed problem size. More processors. How high can speedup go?

S(p) = 1s + 1-sp

Strong vs weak scaling

Same cluster. Two different questions.

  • Strong — fixed global problem. Wall time should fall toward 1/p.
  • Weak — fixed work per rank. Wall time should stay flat.

Hide the network

Distributed training dies on communication latency long before single-kernel FLOPs run out. The backward pass is friendly: finish layer N grads, start all-reduce, keep computing layer N−1.

code
Layer N backward: [==compute==] Layer N all-reduce: [===comm===] Layer N-1 backward: [==compute==] ← overlaps
code
# PyTorch DDP does bucketing + async all-reduce for you. # Manual MPI sketch of the same idea: req = comm.Iallreduce(grads_n) # non-blocking grads_nm1 = layer_nm1.backward(...) # compute while network works req.Wait()

Ideal: collective time fully hidden. That only happens when per-layer compute ≥ all-reduce time for that layer’s gradients.

Which roof are you under?

The roofline puts a kernel between peak bandwidth and peak FLOPs using operational intensity (FLOPs / byte). Most DL non-GEMM ops are memory-bound. Large matmuls sit under the compute roof. Multi-node jobs add a third roof: collectives.

For an interactive A100/H100 ridge plot, see HBM memory.

Load imbalance is a fake serial fraction

Even perfect math dies if ranks finish at different times:

  • Variable-length sequences without bucketing
  • Mixed GPU SKUs (fast ranks wait at every barrier)
  • Pipeline bubbles at fill and drain

Static partitions work when cost is uniform. Otherwise you need dynamic scheduling, sequence packing, or topology-aware placement — treat stragglers as serial work you invented.

Profile first

ToolJob
Nsight SystemsTimeline: gaps, H2D, kernels, NCCL
Nsight ComputeOne kernel: occupancy, stalls, roofline
VTune / perfHost: loader, threads, cache
code
nsys profile --trace=cuda,nvtx,osrt,nccl python train.py ncu --target-processes all -k regex:gemm python train.py

Five minutes of Systems often beats five days of kernel micro-opts on the wrong 2% bar.

What to do

  1. Measure the serial + sync fraction — if s is large, more GPUs are a tax.
  2. Pick strong or weak on purpose — report efficiency at the rank count you will run.
  3. Overlap collectives — async all-reduce / DDP buckets; do not block every layer.
  4. Name the roof — memory → fuse & cut traffic; compute → tensor cores & precision; network → overlap & topology.
  5. Profile before tuning — Systems for gaps, Compute for why, one change at a time.
code
# Cheap first pass: is the GPU even busy? nvidia-smi dmon -s u # Then the timeline that tells the truth nsys profile -o report --trace=cuda,nvtx,osrt,nccl python train.py

Further reading

If you found this explanation helpful, consider sharing it with others.

Mastodon