Skip to main content

CUDA Context vs Streams vs MPS: Which Layer Fixes What

Summary
Decision map for CUDA: a context is per-process GPU state, a stream is an in-order queue inside a context, and MPS shares one context across processes. Pick the layer that matches the problem.

Three names, three different bugs

Someone says “we need more contexts.” Nsight is serial. Another process on the same GPU is stuck waiting. A pointer from GPU 0 blows up after cudaSetDevice(1).

Those are not one problem with three names. A context is the per-device container of GPU state. A stream is an in-order queue inside that container. MPS is a userspace front-end that lets many processes share one context so they co-execute instead of time-slicing.

This page is the decision map. Full machinery lives on the sibling pages.

Two rules

1. Streams nest inside contexts
Adding streams never adds contexts. Concurrency within a process is almost always more streams on the primary context — not a second context on the same GPU.

2. Multi-process is not multi-stream
Two processes each get a private primary context by default. Their streams cannot see each other. Sharing means MPS (or MIG), not cudaStreamCreate in both PIDs.

Three layers

Flip the chip for the question each layer answers — device/state, order/overlap, or multi-process share.

The hierarchy

Process owns devices. Each device gets a context (usually one primary under the runtime API). Context owns the stream table. Streams own the operation queues.

What changes across production setups is how many of each layer you create — not the nesting.

Context vs stream, one aspect at a time

Big comparison tables hide the throughline. Pick the aspect you are arguing about.

Which layer for this symptom?

Start from what you see in the profiler or error log.

Mixups people ask in the same breath

Default stream ≠ context. Streams do not own VRAM. Two GPUs mean two contexts first. Create order matters.

Where MPS sits

MPS is not a fourth box inside the context/stream tree. It is a server process that owns the CUDA context and accepts client work so multiple host processes co-execute on one GPU.

Reach for MPS when: small-batch inference replicas, notebook fleets, many short-lived workers paying CUDA init tax.

Skip MPS when: one job already saturates the GPU, each process already pins its own device, or you need hard isolation (look at MIG instead).

Deep dive: CUDA Multi-Process Service.

Traps

Wrong layer is the expensive trap. Dual contexts for “speed,” streams created on the wrong device, and legacy null-stream serialization are the usual sequels.

What to do

  1. In-process overlap → non-default streams, async copies, events — stay on one primary context per device.
  2. Second GPU → cudaSetDevice, then allocate and create streams on that device.
  3. Many processes, one GPU → MPS (throughput share) or MIG (partition), not more streams in each PID.
  4. Nsight serial → hunt default-stream / implicit sync first; do not invent a second context.
  5. Invalid pointer → match current context/device to the inventory that allocated the handle.

Deeper dives

  • CUDA contexts — primary vs current stack, isolation, create cost
  • CUDA streams — default-stream trap, pipelines, pitfalls
  • CUDA MPS — shared context, SM fill, thread caps
GPU & High-Performance Computing
CUDA Contexts: Ownership, Current Stack, Isolation

Deep dive into the CUDA context object: control vs data plane, inventory (memory, modules, streams, events, graphs), push/pop/setCurrent stacks, primary retain/release, flags and limits, isolation, cost, and traps.

GPU & High-Performance Computing
CUDA Streams: Asynchronous Execution and Concurrency

A CUDA stream is an in-order queue of GPU ops. Overlap H→D, kernels, and D→H across streams — and avoid the default-stream trap that serializes everything.

GPU & High-Performance Computing
CUDA Multi-Process Service (MPS): Sharing One GPU Context

Why exclusive CUDA contexts leave SMs idle under multi-process load, how MPS multiplexes clients through a shared context, thread percentage caps, and when to pick exclusive, MPS, or MIG.

GPU & High-Performance Computing
NVIDIA Device Files in /dev/

How /dev/nvidia*, nvidiactl, nvidia-uvm, and DRI nodes map major/minor numbers to the driver, what CUDA opens first, and the minimum mount set for containers.

GPU & High-Performance Computing
NVIDIA vs AMD for Deep Learning: CUDA vs ROCm and the Datacenter Accelerators

NVIDIA vs AMD for deep learning compared at both layers: the CUDA vs ROCm software moat, the microarchitecture (warp vs wavefront, SM vs CU, Tensor vs Matrix Cores), and the datacenter accelerators (H100/H200/B200 vs MI300X/MI325X).

GPU & High-Performance Computing
Flynn's Classification: Taxonomy of Computer Architectures

Flynn's Classification explained — SISD, SIMD, MISD, MIMD with interactive architecture explorer, SIMD evolution from MMX to AMX, branch divergence visualization, and workload-architecture throughput comparison.

If you found this explanation helpful, consider sharing it with others.

Mastodon