Three names, three different bugs
Someone says “we need more contexts.” Nsight is serial. Another process on the same GPU is stuck waiting. A pointer from GPU 0 blows up after cudaSetDevice(1).
Those are not one problem with three names. A context is the per-device container of GPU state. A stream is an in-order queue inside that container. MPS is a userspace front-end that lets many processes share one context so they co-execute instead of time-slicing.
This page is the decision map. Full machinery lives on the sibling pages.
Two rules
1. Streams nest inside contexts
Adding streams never adds contexts. Concurrency within a process is almost always more streams on the primary context — not a second context on the same GPU.
2. Multi-process is not multi-stream
Two processes each get a private primary context by default. Their streams cannot see each other. Sharing means MPS (or MIG), not cudaStreamCreate in both PIDs.
Three layers
Flip the chip for the question each layer answers — device/state, order/overlap, or multi-process share.
The hierarchy
Process owns devices. Each device gets a context (usually one primary under the runtime API). Context owns the stream table. Streams own the operation queues.
What changes across production setups is how many of each layer you create — not the nesting.
Context vs stream, one aspect at a time
Big comparison tables hide the throughline. Pick the aspect you are arguing about.
Which layer for this symptom?
Start from what you see in the profiler or error log.
Mixups people ask in the same breath
Default stream ≠ context. Streams do not own VRAM. Two GPUs mean two contexts first. Create order matters.
Where MPS sits
MPS is not a fourth box inside the context/stream tree. It is a server process that owns the CUDA context and accepts client work so multiple host processes co-execute on one GPU.
Reach for MPS when: small-batch inference replicas, notebook fleets, many short-lived workers paying CUDA init tax.
Skip MPS when: one job already saturates the GPU, each process already pins its own device, or you need hard isolation (look at MIG instead).
Deep dive: CUDA Multi-Process Service.
Traps
Wrong layer is the expensive trap. Dual contexts for “speed,” streams created on the wrong device, and legacy null-stream serialization are the usual sequels.
What to do
- In-process overlap → non-default streams, async copies, events — stay on one primary context per device.
- Second GPU →
cudaSetDevice, then allocate and create streams on that device. - Many processes, one GPU → MPS (throughput share) or MIG (partition), not more streams in each PID.
- Nsight serial → hunt default-stream / implicit sync first; do not invent a second context.
- Invalid pointer → match current context/device to the inventory that allocated the handle.
Deeper dives
- CUDA contexts — primary vs current stack, isolation, create cost
- CUDA streams — default-stream trap, pipelines, pitfalls
- CUDA MPS — shared context, SM fill, thread caps
Related concepts
Deep dive into the CUDA context object: control vs data plane, inventory (memory, modules, streams, events, graphs), push/pop/setCurrent stacks, primary retain/release, flags and limits, isolation, cost, and traps.
A CUDA stream is an in-order queue of GPU ops. Overlap H→D, kernels, and D→H across streams — and avoid the default-stream trap that serializes everything.
Why exclusive CUDA contexts leave SMs idle under multi-process load, how MPS multiplexes clients through a shared context, thread percentage caps, and when to pick exclusive, MPS, or MIG.
How /dev/nvidia*, nvidiactl, nvidia-uvm, and DRI nodes map major/minor numbers to the driver, what CUDA opens first, and the minimum mount set for containers.
NVIDIA vs AMD for deep learning compared at both layers: the CUDA vs ROCm software moat, the microarchitecture (warp vs wavefront, SM vs CU, Tensor vs Matrix Cores), and the datacenter accelerators (H100/H200/B200 vs MI300X/MI325X).
Flynn's Classification explained — SISD, SIMD, MISD, MIMD with interactive architecture explorer, SIMD evolution from MMX to AMX, branch divergence visualization, and workload-architecture throughput comparison.
