Skip to main content

Tensor Cores: Mixed Precision & Matrix Acceleration

Summary
How NVIDIA Tensor Cores accelerate GEMM: warp-level MMA, precision dials (TF32/FP16/BF16/INT8/FP8), shape alignment cliffs, AMP vs model.half(), and when the speedup actually shows up.

The matmul is fine. The util is not.

You enabled AMP. Loss looks stable. Step time barely moved. The profiler still shows a fat GEMM — but not the Tensor Core kernel you expected. Dimensions are “almost” multiples of 8. Or someone called model.half() six months ago and nobody re-checked.

Tensor Cores are not “CUDA cores but faster.” They are a separate unit that eats warp-level matrix multiply-accumulate when precision, shape, layout, and library kernel all agree. Miss one gate and you pay CUDA-core rates while telling yourself mixed precision is on.

This page is those gates — not a catalog of every WMMA fragment shape.

Two numbers

1. ~16× dense peak (FP16/BF16 vs FP32 on A100)
Marketing math for aligned, compute-bound GEMMs. Your layer may be memory-bound and never touch it.

2. Alignment is a cliff
If M, N, or K miss the precision’s multiple (often 8 for FP16/BF16, 16 for INT8), many kernels fall back hard. K=4095 vs K=4096 is the classic trap.

Same GEMM, two engines

CUDA cores do scalar and vector MACs. Tensor Cores do tile MMA: one warp, one fragment product, far fewer cycles for the same math when the path is open.

Precision is a dial

Lower width → higher peak math rate, tighter range or fewer mantissa bits. TF32 keeps FP32 storage but truncates mantissa in the multiply. BF16 keeps FP32 exponent range (training-friendly). INT8/FP8 are inference (or specialized training) after calibration.

Alignment gates the speedup

AMP does not forgive bad shapes. Flip a dimension off the multiple and watch effective rate fall toward the FP32 CUDA-core baseline in this teaching model.

Inside one Tensor Core op

Public APIs expose warp-collective steps (load → mma_sync → store). Fragment sizes vary by architecture and precision; the story is the same: 32 lanes cooperate on one tile.

In production you rarely write WMMA by hand — cuBLAS, cuDNN, CUTLASS, and framework kernels do. The film is so you know what “Tensor Core path” means when the profiler names it.

How you actually turn them on

You do not flip a “Tensor Cores” switch. You put eligible math on a dtype and shape the hardware accelerates. In PyTorch that is almost always AMP, not a global .half().

Enable mixed precision when:

  • You train or fine-tune large models on Volta or newer.
  • Matmul shapes can hit the alignment multiples (or padding is measured-cheap).
  • You can use BF16 on Ampere/Hopper — skip the GradScaler dance when numerics allow.

Stay in FP32 when:

  • The model is tiny and launch overhead dominates.
  • Numerics are still broken in full precision — fix that first.
  • You need bitwise reproducibility (disable TF32 / TC-sensitive paths for tests).
  • Hardware is pre-Volta.
code
import torch from torch.amp import autocast, GradScaler device = "cuda" model = model.to(device) # BF16 path — common Ampere/Hopper default for batch in dataloader: optimizer.zero_grad(set_to_none=True) with autocast("cuda", dtype=torch.bfloat16): loss = criterion(model(batch["input"].to(device)), batch["target"].to(device)) loss.backward() optimizer.step() # FP16 needs GradScaler scaler = GradScaler("cuda") with autocast("cuda", dtype=torch.float16): loss = criterion(model(x), y) scaler.scale(loss).backward() scaler.step(optimizer) scaler.update()

Traps

Same four failure modes show up in every “AMP didn’t help” review.

What to do

  1. Start with BF16 autocast on Ampere/Hopper; use FP16 + GradScaler when you must.
  2. Align M, N, K (and channels) to the precision’s multiple — pad only with profiler evidence.
  3. Prefer library kernels (cuBLAS / cuDNN / framework) over hand WMMA unless you are writing a kernel library.
  4. Profile Tensor Core utilization — Nsight Compute or the PyTorch profiler — not just “AMP is imported.”
  5. Validate numerics after every precision or quantization change; never publish peak TFLOPS as a product SLA.

Further reading

GPU & High-Performance Computing
NVIDIA vs AMD for Deep Learning: CUDA vs ROCm and the Datacenter Accelerators

NVIDIA vs AMD for deep learning compared at both layers: the CUDA vs ROCm software moat, the microarchitecture (warp vs wavefront, SM vs CU, Tensor vs Matrix Cores), and the datacenter accelerators (H100/H200/B200 vs MI300X/MI325X).

GPU & High-Performance Computing
CUDA Contexts: Ownership, Current Stack, Isolation

Deep dive into the CUDA context object: control vs data plane, inventory (memory, modules, streams, events, graphs), push/pop/setCurrent stacks, primary retain/release, flags and limits, isolation, cost, and traps.

GPU & High-Performance Computing
CUDA Context vs Streams vs MPS: Which Layer Fixes What

Decision map for CUDA: a context is per-process GPU state, a stream is an in-order queue inside a context, and MPS shares one context across processes. Pick the layer that matches the problem.

GPU & High-Performance Computing
CUDA Multi-Process Service (MPS): Sharing One GPU Context

Why exclusive CUDA contexts leave SMs idle under multi-process load, how MPS multiplexes clients through a shared context, thread percentage caps, and when to pick exclusive, MPS, or MIG.

GPU & High-Performance Computing
CUDA Streams: Asynchronous Execution and Concurrency

A CUDA stream is an in-order queue of GPU ops. Overlap H→D, kernels, and D→H across streams — and avoid the default-stream trap that serializes everything.

GPU & High-Performance Computing
GPU Pipeline Hazards: Scoreboard Stalls and Warp Divergence

A fine kernel can still profile poorly: warp divergence taxes throughput ~1/N; alone on a long scoreboard wait, SM issue util can hit 0%. Interactive demos.

If you found this explanation helpful, consider sharing it with others.

Mastodon