The matmul is fine. The util is not.
You enabled AMP. Loss looks stable. Step time barely moved. The profiler still shows a fat GEMM — but not the Tensor Core kernel you expected. Dimensions are “almost” multiples of 8. Or someone called model.half() six months ago and nobody re-checked.
Tensor Cores are not “CUDA cores but faster.” They are a separate unit that eats warp-level matrix multiply-accumulate when precision, shape, layout, and library kernel all agree. Miss one gate and you pay CUDA-core rates while telling yourself mixed precision is on.
This page is those gates — not a catalog of every WMMA fragment shape.
Two numbers
1. ~16× dense peak (FP16/BF16 vs FP32 on A100)
Marketing math for aligned, compute-bound GEMMs. Your layer may be memory-bound and never touch it.
2. Alignment is a cliff
If M, N, or K miss the precision’s multiple (often 8 for FP16/BF16, 16 for INT8), many kernels fall back hard. K=4095 vs K=4096 is the classic trap.
Same GEMM, two engines
CUDA cores do scalar and vector MACs. Tensor Cores do tile MMA: one warp, one fragment product, far fewer cycles for the same math when the path is open.
Precision is a dial
Lower width → higher peak math rate, tighter range or fewer mantissa bits. TF32 keeps FP32 storage but truncates mantissa in the multiply. BF16 keeps FP32 exponent range (training-friendly). INT8/FP8 are inference (or specialized training) after calibration.
Alignment gates the speedup
AMP does not forgive bad shapes. Flip a dimension off the multiple and watch effective rate fall toward the FP32 CUDA-core baseline in this teaching model.
Inside one Tensor Core op
Public APIs expose warp-collective steps (load → mma_sync → store). Fragment sizes vary by architecture and precision; the story is the same: 32 lanes cooperate on one tile.
In production you rarely write WMMA by hand — cuBLAS, cuDNN, CUTLASS, and framework kernels do. The film is so you know what “Tensor Core path” means when the profiler names it.
How you actually turn them on
You do not flip a “Tensor Cores” switch. You put eligible math on a dtype and shape the hardware accelerates. In PyTorch that is almost always AMP, not a global .half().
Enable mixed precision when:
- You train or fine-tune large models on Volta or newer.
- Matmul shapes can hit the alignment multiples (or padding is measured-cheap).
- You can use BF16 on Ampere/Hopper — skip the GradScaler dance when numerics allow.
Stay in FP32 when:
- The model is tiny and launch overhead dominates.
- Numerics are still broken in full precision — fix that first.
- You need bitwise reproducibility (disable TF32 / TC-sensitive paths for tests).
- Hardware is pre-Volta.
import torch from torch.amp import autocast, GradScaler device = "cuda" model = model.to(device) # BF16 path — common Ampere/Hopper default for batch in dataloader: optimizer.zero_grad(set_to_none=True) with autocast("cuda", dtype=torch.bfloat16): loss = criterion(model(batch["input"].to(device)), batch["target"].to(device)) loss.backward() optimizer.step() # FP16 needs GradScaler scaler = GradScaler("cuda") with autocast("cuda", dtype=torch.float16): loss = criterion(model(x), y) scaler.scale(loss).backward() scaler.step(optimizer) scaler.update()
Traps
Same four failure modes show up in every “AMP didn’t help” review.
What to do
- Start with BF16 autocast on Ampere/Hopper; use FP16 + GradScaler when you must.
- Align M, N, K (and channels) to the precision’s multiple — pad only with profiler evidence.
- Prefer library kernels (cuBLAS / cuDNN / framework) over hand WMMA unless you are writing a kernel library.
- Profile Tensor Core utilization — Nsight Compute or the PyTorch profiler — not just “AMP is imported.”
- Validate numerics after every precision or quantization change; never publish peak TFLOPS as a product SLA.
Further reading
- NVIDIA A100 Tensor Core GPU
- NVIDIA H100 Tensor Core GPU
- CUDA C++ Programming Guide: WMMA
- Automatic Mixed Precision (PyTorch)
- Mixed Precision Training
Related concepts
NVIDIA vs AMD for deep learning compared at both layers: the CUDA vs ROCm software moat, the microarchitecture (warp vs wavefront, SM vs CU, Tensor vs Matrix Cores), and the datacenter accelerators (H100/H200/B200 vs MI300X/MI325X).
Deep dive into the CUDA context object: control vs data plane, inventory (memory, modules, streams, events, graphs), push/pop/setCurrent stacks, primary retain/release, flags and limits, isolation, cost, and traps.
Decision map for CUDA: a context is per-process GPU state, a stream is an in-order queue inside a context, and MPS shares one context across processes. Pick the layer that matches the problem.
Why exclusive CUDA contexts leave SMs idle under multi-process load, how MPS multiplexes clients through a shared context, thread percentage caps, and when to pick exclusive, MPS, or MIG.
A CUDA stream is an in-order queue of GPU ops. Overlap H→D, kernels, and D→H across streams — and avoid the default-stream trap that serializes everything.
A fine kernel can still profile poorly: warp divergence taxes throughput ~1/N; alone on a long scoreboard wait, SM issue util can hit 0%. Interactive demos.
