You got a pointer. Nobody said context.
You call cudaSetDevice(0) and cudaMalloc. A device pointer comes back. Nothing in that call site says context, yet the driver just retained a primary CUDA context for process × GPU 0: host-side tables, device virtual address space, default stream, and a place to hang every later module, event, and graph.
Most production bugs that look like “random CUDA errors” are really wrong current context, unbalanced push/pop, handle used after destroy, or cross-context pointer. This page is that machinery — what the context is, what lives inside it, and how threads bind to it.
Two numbers
1. Current is per CPU thread
Process-wide “we initialized CUDA” is not enough. Each thread has its own current context (often via a stack). A worker that never called cudaSetDevice gets CUDA_ERROR_INVALID_CONTEXT.
2. Handles do not cross contexts
A cudaMalloc pointer, stream, module, event, or graph is inventory of one context. Same physical GPU, different context → same bit pattern can be invalid. Multi-process default = one private context each (time-sliced) unless MPS shares a server context.
What a context is
A CUDA context (CUcontext) is the synchronized pair:
| Side | Lives on | Holds |
|---|---|---|
| Control plane | Host (driver + your process) | API entry, handle tables, current stack, command submission, flags/limits |
| Data plane | Device | VA space, allocations, loaded code, schedulable work for this owner |
One context ↔ one process × one device association (a process can own many contexts across devices; rare code owns several on one device). The opaque handle is not “a pointer into VRAM” — it names the whole bag.
Inside the context object
When docs say “resources associated with the context,” they mean a concrete inventory. Destroy frees almost all of it. Flip shelves:
Memory shelf
- Device allocations (
cuMemAlloc/cudaMalloc, pitched, many managed paths) live in the context’s address space. - CUDA arrays / mipmaps for texture paths.
- External memory / semaphore imports bound into this world.
- Documented exceptions: some virtual-memory and stream-ordered pool paths have lifetimes that are not simply “die with cuCtxDestroy” — read the destroy notes if you use those APIs.
Code shelf
- Modules (
CUmodule) and functions (CUfunction) from cubin/fatbin loads. - Link / library state used during JIT.
- Device-side dynamic parallelism is gated by context limits (sync depth, pending launch count).
Work shelf
- Default stream (every context has one; legacy vs per-thread default stream modes change serialization).
- User streams and events — created under the current context, unusable under another.
- Graphs and graph execs.
cuCtxSynchronizewaits for all outstanding work in the (current or named) context — blunter than a single stream sync.
Policy shelf
Flags at create time, limits (cuCtxSetLimit), cache config, stream priority range, and on newer stacks execution affinity (e.g. SM count under MPS). These are properties of the context object, not of a single kernel launch (kernels can still override some cache prefs).
Current context and the stack
CUDA routes almost every API call to the current context of the calling thread.
| API | What it does |
|---|---|
cudaSetDevice(i) | Runtime: make device i’s primary context current (retain on first use) |
cuCtxGetCurrent | Read current (NULL if floating) |
cuCtxSetCurrent(ctx) | Bind ctx as current; NULL clears current |
cuCtxPushCurrent(ctx) | Push onto this thread’s stack; top becomes current |
cuCtxPopCurrent | Pop top; previous current restored |
cuDevicePrimaryCtxRetain | Bump primary refcount — does not make it current by itself |
Empty stack / NULL current → invalid context. Two threads can both hold the same context current — isolation is by context, not by thread.
Push, pop, set — the film
Libraries that use the Driver API almost always follow: push my context → do work → pop. That restores the caller’s current context. Unbalanced push is how “works until library X returns” bugs ship.
// Canonical library sandwich CUcontext caller; cuCtxGetCurrent(&caller); // optional: know what you restore cuCtxPushCurrent(myCtx); // ... malloc, launch, sync under myCtx ... cuCtxPopCurrent(NULL); // restore previous current // Primary retain pitfall CUcontext p; cuDevicePrimaryCtxRetain(&p, dev); // retained, NOT current cuCtxSetCurrent(p); // now current — or use cudaSetDevice
cuCtxCreate note: creating a context associates it with the calling thread and makes it current (usage count 1). You must eventually cuCtxDestroy. Destroy pops it from the calling thread if it is current there; other threads that still list it as current see CUDA_ERROR_CONTEXT_IS_DESTROYED on later use.
cuCtxSetCurrent vs push: use push/pop when you must restore the caller. Use set when you own the thread (worker pinned to one context). Pass NULL to float (no current).
Framework code (PyTorch, JAX, etc.) almost always stays on primary contexts and cudaSetDevice. You still need every pool worker to set the device before first CUDA call on that thread.
Primary context vs explicit create
| Primary (Runtime) | Explicit (Driver cuCtxCreate) | |
|---|---|---|
| Who creates | Runtime on first use per device | You |
| Retain API | cuDevicePrimaryCtxRetain / Runtime first call | create returns handle |
| Typical apps | Nearly all ML / CUDA Runtime code | Engines, multi-ctx tools, plugins |
| Lifetime | Refcounted until last release / process teardown | cuCtxDestroy frees inventory |
| Mental model | “This process on GPU i” | “This library’s private GPU world” |
Primary flags (cuDevicePrimaryCtxSetFlags) must be set before the primary is initialized. After retain, you still need current binding.
Fork rule: do not fork() after CUDA has initialized. The child inherits host state that is not safe to use. Prefer fork-before-init, or separate processes from the start.
Flags, limits, and sync scope
The context is not only a bag of pointers — it carries scheduling flags, resource limits, and cache preferences. Wrong layer again: SCHED_SPIN will not fix a default-stream serialize; stream priority will not fix a cross-context pointer.
// Limits (current context) size_t heap; cuCtxGetLimit(&heap, CU_LIMIT_MALLOC_HEAP_SIZE); cuCtxSetLimit(CU_LIMIT_MALLOC_HEAP_SIZE, 128ull << 20); // Blunt sync cuCtxSynchronize(); // all work in current context // Prefer when possible: cuStreamSynchronize(stream); // one queue
Isolation wall
Contexts isolate logical GPU worlds on the same silicon (even the same process can hold two contexts on one GPU via the Driver API — rare, but the wall still applies):
- Device pointers are context virtual addresses, not portable physical tokens. Same GPU ≠ same map.
- Streams, events, modules, graphs, and classic malloc inventory are handles bound to the creating context.
- Threads may share one context (each must make it current). Processes default to private contexts and exclusive time-slicing.
- IPC maps exported memory into the importer’s context (new local pointer). MPS co-schedules clients for SM packing — it does not make raw pointers portable across processes (use IPC for that). MIG is a hardware partition, not a software context share.
# Sketch: two processes, two contexts, no IPC — ptr is not shareable # process A d_a = cuda.malloc(n) # process B cannot use d_a — different context / address space # (cudaIpcGetMemHandle / OpenMemHandle if you need a local map)
When many small processes should co-reside kernels, prefer MPS (or consolidate into one process with many streams) — not “create more contexts per process.” Shared buffers across processes still need IPC (or a single process).
Cost: create, warm path, multi-process slice
Context setup is not free. First touch can pull in driver open paths (device files), primary retain, and module load/JIT. Steady-state launches are cheap by comparison.
Exclusive multi-process scheduling pays a different tax: switch between private contexts while small kernels leave SMs empty — the problem MPS targets.
Persistence mode / nvidia-persistenced keeps the device warm between jobs; it does not replace your process’s CUDA context, but it cuts cold-open pain for short tasks.
Numbers on the tape are pedagogical. Measure first cuda call vs nth launch on your driver and SKU.
Context vs stream (and multi-GPU)
Scaling units people confuse:
| Layer | Scales | Job |
|---|---|---|
| Process | OS isolation | Who talks to the driver |
| Device | Physical GPU | Which chip |
| Context | Ownership / VA / inventory | Isolation boundary |
| Stream | Concurrency lanes | Overlap copy/compute inside a context |
Default mental model for one training process on one GPU:
process └─ GPU 0 └─ primary context ├─ flags · limits · cache config ├─ modules / functions ├─ device allocations (VA space) ├─ default stream ├─ stream 0 (H2D) ├─ stream 1 (compute) └─ stream 2 (D2H)
Multi-GPU: one primary context per device, switch with cudaSetDevice (or a thread pinned per GPU). Peer access and NCCL sit on top of those separate contexts — they do not merge them into one address space.
Deeper stream mechanics: CUDA streams. Side-by-side compare: context vs streams.
Lifecycle (practical)
- Open path — process loads
libcuda, opens/dev/nvidiactl+ device nodes + often UVM. - Bind device —
cudaSetDevice/cuCtxCreate/ first Runtime call retains primary. - Optional policy — primary flags before init; limits and cache config after current.
- Accumulate — mallocs, modules, streams, graphs attach to that context.
- Submit — launches and memcopies use current context on the calling thread.
- Library borrow — push foreign context, work, pop (never leave unbalanced).
- Teardown — process exit, or
cuCtxDestroy/ primary release / device reset: inventory goes.
# See processes holding the GPU (contexts behind them) nvidia-smi # Compute mode matters for multi-process sharing policies nvidia-smi | grep -i 'Compute Mode'
Traps
| Symptom | Check first |
|---|---|
INVALID_CONTEXT on worker | Did this thread call cudaSetDevice / set current? |
INVALID_CONTEXT after library return | Unbalanced push/pop or destroy? |
| Weird pointer / illegal address | Buffer allocated under a different context or device? |
Works in main, dies after fork | CUDA init before fork |
| Multi-process low util | Private contexts time-sliced — consider MPS or one process |
CONTEXT_IS_DESTROYED | Destroy while another thread still had it current |
| Retain but still invalid | Primary retain without SetCurrent |
What to do
- Treat context as ownership inventory and stream as ordering — fix isolation bugs before tuning streams.
- Call
cudaSetDeviceon every thread that will issue CUDA work (or push an explicit context). - Prefer one primary context per device per process; avoid dual contexts on one GPU unless a library requires it.
- If you push, pop on every path; libraries should not leave a foreign current.
- Never fork after CUDA init; size context lifetime to the process or a clear library enter/exit.
- For many small multi-process jobs on one GPU, evaluate MPS; keep devices warm with persistenced when jobs are short.
Related concepts
Decision map for CUDA: a context is per-process GPU state, a stream is an in-order queue inside a context, and MPS shares one context across processes. Pick the layer that matches the problem.
A CUDA stream is an in-order queue of GPU ops. Overlap H→D, kernels, and D→H across streams — and avoid the default-stream trap that serializes everything.
Why exclusive CUDA contexts leave SMs idle under multi-process load, how MPS multiplexes clients through a shared context, thread percentage caps, and when to pick exclusive, MPS, or MIG.
How /dev/nvidia*, nvidiactl, nvidia-uvm, and DRI nodes map major/minor numbers to the driver, what CUDA opens first, and the minimum mount set for containers.
NVIDIA vs AMD for deep learning compared at both layers: the CUDA vs ROCm software moat, the microarchitecture (warp vs wavefront, SM vs CU, Tensor vs Matrix Cores), and the datacenter accelerators (H100/H200/B200 vs MI300X/MI325X).
Flynn's Classification explained — SISD, SIMD, MISD, MIMD with interactive architecture explorer, SIMD evolution from MMX to AMX, branch divergence visualization, and workload-architecture throughput comparison.
