GPU Streaming Multiprocessor (SM)
How a single SM actually runs work: warps of 32, cutaway of cores and memory, divergence tax, latency-hiding occupancy, and coalesced loads — instruments, not a catalog.
16 min readConcept
Explore machine learning concepts related to Hardware Architecture. Clear explanations and practical insights.
How a single SM actually runs work: warps of 32, cutaway of cores and memory, divergence tax, latency-hiding occupancy, and coalesced loads — instruments, not a catalog.