Skip to main content

GPU Streaming Multiprocessor (SM)

Summary
How a single SM actually runs work: warps of 32, cutaway of cores and memory, divergence tax, latency-hiding occupancy, and coalesced loads — instruments, not a catalog.

You launched 100k threads. They still run 32 at a time.

An A100 has 108 Streaming Multiprocessors. Your kernel grid is a sea of blocks. Every block that is currently live sits entirely on one SM — with that SM’s register file, shared memory, and four warp schedulers.

None of that looks like a CPU core. The SM does not speculate for one thread. It keeps dozens of warps resident and issues whoever is ready. Miss the warp size, the occupancy math, or the memory pattern and the hardware is mostly waiting — while the launch still “succeeds.”

This page is one SM: what is inside it, how work is packed, and the three taxes that show up in every slow kernel review.

Two numbers

1. 32
A warp is 32 threads with one instruction stream. Divergence and coalescing are both warp-shaped problems.

2. ~500 cycles
Round-trip to HBM when caches miss. The SM does not make memory faster — it hides that gap by switching warps. No spare warps → hollow idle cycles.

What’s on the floor

CUDA cores, Tensor Cores, schedulers, register file, shared/L1 — all in one workshop. Click a bay for the role that matters to software.

Threads, warps, blocks — one SM

A block never splits across SMs. Warps of 32 are the issue unit. Dial block size and how many blocks you try to park here; the stage recomposes the resident set.

Hierarchy in one line: thread → warp (32) → block (≤1024 threads, one SM) → grid (whole launch). Shared memory and __syncthreads are block-scoped. Different blocks on the same SM do not share a barrier.

Divergence is a throughput tax

If lanes in a warp take different branches, the hardware serializes paths. Idle lanes mask off. Worst case: 32-way switch ≈ 1/32 of peak for that region.

Structure data so threads in a warp stay on the same path when you can. When branching is unavoidable, push the split across warps, not inside one.

Latency is hidden by spare warps

Issue → memory stall → need another ready warp. With too few residents, the issue tape shows amber gaps. With enough, the stall disappears into other work.

Real stalls are hundreds of cycles, not four — which is why occupancy (warps waiting in the wings) is a first-class tuning knob, not a vanity metric.

Occupancy is three knobs, one SM

Registers per thread, shared memory per block, and threads per block all compete. One of them is usually the limiter. High occupancy helps memory-bound code hide latency; compute-bound kernels can live lower if each thread stays busy.

Loads that walk together

When 32 lanes touch 32 consecutive addresses, the memory system can fold them into few transactions. Scatter multiplies traffic for the same 32 values. Shared-memory tiling exists so you pay global once and reuse on-chip.

LevelScopeLatency (order of)Role
Registersper thread~0live locals
Shared / L1per block / SM~tens of cyclesreuse, cooperation
L2GPU-wide~hundredsshared cache
HBMGPU-wide~500 cyclesbulk store

L1 and shared often share the same SRAM bank; the split is configurable. Prefer explicit shared tiles when reuse is regular; lean on L1 when access is irregular.

What to do

  1. Size blocks in multiples of 32 — usually 128–512 threads. Tiny blocks waste slots; max-size blocks fight register/shared limits.
  2. Keep warps uniform — same branch path inside a warp when performance matters.
  3. Budget occupancy for memory-bound kernels — enough warps to cover HBM latency; use the occupancy calculator / Nsight, not guesswork.
  4. Coalesce global loads — consecutive addresses per warp; tile through shared when you reuse.
  5. Treat registers as occupancy fuel — spilling or 96+ regs/thread often kills resident warps.
  6. Profile the SM, not the launch — Nsight Compute: warp stall reasons, memory throughput, achieved occupancy.

Generations (short)

The SIMT model stuck; specialized units grew.

GenEraNotable on the SM
Volta2017Tensor Cores, independent thread scheduling
Turing2018RT cores (consumer), concurrent FP32+INT32
Ampere2020128 FP32/SM class, larger shared/L1, sparse TC
Hopper / Ada2022more shared, FP8 / Transformer Engine paths
Blackwell2024denser TC formats (e.g. FP4 class)

You still write for warps, occupancy, and the memory ladder. New silicon adds engines; it does not remove the taxes above.

Further reading

GPU & High-Performance Computing
GPU Pipeline Hazards: Scoreboard Stalls and Warp Divergence

A fine kernel can still profile poorly: warp divergence taxes throughput ~1/N; alone on a long scoreboard wait, SM issue util can hit 0%. Interactive demos.

GPU & High-Performance Computing
GPU Memory Hierarchy & Optimization

Master GPU memory hierarchy from registers to global memory, understand coalescing patterns, bank conflicts, and optimization strategies for maximum performance

GPU & High-Performance Computing
CUDA Contexts: Ownership, Current Stack, Isolation

Deep dive into the CUDA context object: control vs data plane, inventory (memory, modules, streams, events, graphs), push/pop/setCurrent stacks, primary retain/release, flags and limits, isolation, cost, and traps.

GPU & High-Performance Computing
CUDA Context vs Streams vs MPS: Which Layer Fixes What

Decision map for CUDA: a context is per-process GPU state, a stream is an in-order queue inside a context, and MPS shares one context across processes. Pick the layer that matches the problem.

GPU & High-Performance Computing
CUDA Multi-Process Service (MPS): Sharing One GPU Context

Why exclusive CUDA contexts leave SMs idle under multi-process load, how MPS multiplexes clients through a shared context, thread percentage caps, and when to pick exclusive, MPS, or MIG.

GPU & High-Performance Computing
CUDA Streams: Asynchronous Execution and Concurrency

A CUDA stream is an in-order queue of GPU ops. Overlap H→D, kernels, and D→H across streams — and avoid the default-stream trap that serializes everything.

If you found this explanation helpful, consider sharing it with others.

Mastodon