GPU Streaming Multiprocessor (SM)
How a single SM actually runs work: warps of 32, cutaway of cores and memory, divergence tax, latency-hiding occupancy, and coalesced loads — instruments, not a catalog.
7 min readConcept
Explore machine learning concepts related to warp. Clear explanations and practical insights.
How a single SM actually runs work: warps of 32, cutaway of cores and memory, divergence tax, latency-hiding occupancy, and coalesced loads — instruments, not a catalog.