Skip to main content

CUDA Contexts: Ownership, Current Stack, Isolation

Summary
Deep dive into the CUDA context object: control vs data plane, inventory (memory, modules, streams, events, graphs), push/pop/setCurrent stacks, primary retain/release, flags and limits, isolation, cost, and traps.

You got a pointer. Nobody said context.

You call cudaSetDevice(0) and cudaMalloc. A device pointer comes back. Nothing in that call site says context, yet the driver just retained a primary CUDA context for process × GPU 0: host-side tables, device virtual address space, default stream, and a place to hang every later module, event, and graph.

Most production bugs that look like “random CUDA errors” are really wrong current context, unbalanced push/pop, handle used after destroy, or cross-context pointer. This page is that machinery — what the context is, what lives inside it, and how threads bind to it.

Two numbers

1. Current is per CPU thread
Process-wide “we initialized CUDA” is not enough. Each thread has its own current context (often via a stack). A worker that never called cudaSetDevice gets CUDA_ERROR_INVALID_CONTEXT.

2. Handles do not cross contexts
A cudaMalloc pointer, stream, module, event, or graph is inventory of one context. Same physical GPU, different context → same bit pattern can be invalid. Multi-process default = one private context each (time-sliced) unless MPS shares a server context.

What a context is

A CUDA context (CUcontext) is the synchronized pair:

SideLives onHolds
Control planeHost (driver + your process)API entry, handle tables, current stack, command submission, flags/limits
Data planeDeviceVA space, allocations, loaded code, schedulable work for this owner

One context ↔ one process × one device association (a process can own many contexts across devices; rare code owns several on one device). The opaque handle is not “a pointer into VRAM” — it names the whole bag.

Inside the context object

When docs say “resources associated with the context,” they mean a concrete inventory. Destroy frees almost all of it. Flip shelves:

Memory shelf

  • Device allocations (cuMemAlloc / cudaMalloc, pitched, many managed paths) live in the context’s address space.
  • CUDA arrays / mipmaps for texture paths.
  • External memory / semaphore imports bound into this world.
  • Documented exceptions: some virtual-memory and stream-ordered pool paths have lifetimes that are not simply “die with cuCtxDestroy” — read the destroy notes if you use those APIs.

Code shelf

  • Modules (CUmodule) and functions (CUfunction) from cubin/fatbin loads.
  • Link / library state used during JIT.
  • Device-side dynamic parallelism is gated by context limits (sync depth, pending launch count).

Work shelf

  • Default stream (every context has one; legacy vs per-thread default stream modes change serialization).
  • User streams and events — created under the current context, unusable under another.
  • Graphs and graph execs.
  • cuCtxSynchronize waits for all outstanding work in the (current or named) context — blunter than a single stream sync.

Policy shelf

Flags at create time, limits (cuCtxSetLimit), cache config, stream priority range, and on newer stacks execution affinity (e.g. SM count under MPS). These are properties of the context object, not of a single kernel launch (kernels can still override some cache prefs).

Current context and the stack

CUDA routes almost every API call to the current context of the calling thread.

APIWhat it does
cudaSetDevice(i)Runtime: make device i’s primary context current (retain on first use)
cuCtxGetCurrentRead current (NULL if floating)
cuCtxSetCurrent(ctx)Bind ctx as current; NULL clears current
cuCtxPushCurrent(ctx)Push onto this thread’s stack; top becomes current
cuCtxPopCurrentPop top; previous current restored
cuDevicePrimaryCtxRetainBump primary refcount — does not make it current by itself

Empty stack / NULL current → invalid context. Two threads can both hold the same context current — isolation is by context, not by thread.

Push, pop, set — the film

Libraries that use the Driver API almost always follow: push my context → do work → pop. That restores the caller’s current context. Unbalanced push is how “works until library X returns” bugs ship.

code
// Canonical library sandwich CUcontext caller; cuCtxGetCurrent(&caller); // optional: know what you restore cuCtxPushCurrent(myCtx); // ... malloc, launch, sync under myCtx ... cuCtxPopCurrent(NULL); // restore previous current // Primary retain pitfall CUcontext p; cuDevicePrimaryCtxRetain(&p, dev); // retained, NOT current cuCtxSetCurrent(p); // now current — or use cudaSetDevice

cuCtxCreate note: creating a context associates it with the calling thread and makes it current (usage count 1). You must eventually cuCtxDestroy. Destroy pops it from the calling thread if it is current there; other threads that still list it as current see CUDA_ERROR_CONTEXT_IS_DESTROYED on later use.

cuCtxSetCurrent vs push: use push/pop when you must restore the caller. Use set when you own the thread (worker pinned to one context). Pass NULL to float (no current).

Framework code (PyTorch, JAX, etc.) almost always stays on primary contexts and cudaSetDevice. You still need every pool worker to set the device before first CUDA call on that thread.

Primary context vs explicit create

Primary (Runtime)Explicit (Driver cuCtxCreate)
Who createsRuntime on first use per deviceYou
Retain APIcuDevicePrimaryCtxRetain / Runtime first callcreate returns handle
Typical appsNearly all ML / CUDA Runtime codeEngines, multi-ctx tools, plugins
LifetimeRefcounted until last release / process teardowncuCtxDestroy frees inventory
Mental model“This process on GPU i”“This library’s private GPU world”

Primary flags (cuDevicePrimaryCtxSetFlags) must be set before the primary is initialized. After retain, you still need current binding.

Fork rule: do not fork() after CUDA has initialized. The child inherits host state that is not safe to use. Prefer fork-before-init, or separate processes from the start.

Flags, limits, and sync scope

The context is not only a bag of pointers — it carries scheduling flags, resource limits, and cache preferences. Wrong layer again: SCHED_SPIN will not fix a default-stream serialize; stream priority will not fix a cross-context pointer.

code
// Limits (current context) size_t heap; cuCtxGetLimit(&heap, CU_LIMIT_MALLOC_HEAP_SIZE); cuCtxSetLimit(CU_LIMIT_MALLOC_HEAP_SIZE, 128ull << 20); // Blunt sync cuCtxSynchronize(); // all work in current context // Prefer when possible: cuStreamSynchronize(stream); // one queue

Isolation wall

Contexts isolate logical GPU worlds on the same silicon (even the same process can hold two contexts on one GPU via the Driver API — rare, but the wall still applies):

  • Device pointers are context virtual addresses, not portable physical tokens. Same GPU ≠ same map.
  • Streams, events, modules, graphs, and classic malloc inventory are handles bound to the creating context.
  • Threads may share one context (each must make it current). Processes default to private contexts and exclusive time-slicing.
  • IPC maps exported memory into the importer’s context (new local pointer). MPS co-schedules clients for SM packing — it does not make raw pointers portable across processes (use IPC for that). MIG is a hardware partition, not a software context share.
code
# Sketch: two processes, two contexts, no IPC — ptr is not shareable # process A d_a = cuda.malloc(n) # process B cannot use d_a — different context / address space # (cudaIpcGetMemHandle / OpenMemHandle if you need a local map)

When many small processes should co-reside kernels, prefer MPS (or consolidate into one process with many streams) — not “create more contexts per process.” Shared buffers across processes still need IPC (or a single process).

Cost: create, warm path, multi-process slice

Context setup is not free. First touch can pull in driver open paths (device files), primary retain, and module load/JIT. Steady-state launches are cheap by comparison.

Exclusive multi-process scheduling pays a different tax: switch between private contexts while small kernels leave SMs empty — the problem MPS targets.

Persistence mode / nvidia-persistenced keeps the device warm between jobs; it does not replace your process’s CUDA context, but it cuts cold-open pain for short tasks.

Numbers on the tape are pedagogical. Measure first cuda call vs nth launch on your driver and SKU.

Context vs stream (and multi-GPU)

Scaling units people confuse:

LayerScalesJob
ProcessOS isolationWho talks to the driver
DevicePhysical GPUWhich chip
ContextOwnership / VA / inventoryIsolation boundary
StreamConcurrency lanesOverlap copy/compute inside a context

Default mental model for one training process on one GPU:

code
process └─ GPU 0 └─ primary context ├─ flags · limits · cache config ├─ modules / functions ├─ device allocations (VA space) ├─ default stream ├─ stream 0 (H2D) ├─ stream 1 (compute) └─ stream 2 (D2H)

Multi-GPU: one primary context per device, switch with cudaSetDevice (or a thread pinned per GPU). Peer access and NCCL sit on top of those separate contexts — they do not merge them into one address space.

Deeper stream mechanics: CUDA streams. Side-by-side compare: context vs streams.

Lifecycle (practical)

  1. Open path — process loads libcuda, opens /dev/nvidiactl + device nodes + often UVM.
  2. Bind device — cudaSetDevice / cuCtxCreate / first Runtime call retains primary.
  3. Optional policy — primary flags before init; limits and cache config after current.
  4. Accumulate — mallocs, modules, streams, graphs attach to that context.
  5. Submit — launches and memcopies use current context on the calling thread.
  6. Library borrow — push foreign context, work, pop (never leave unbalanced).
  7. Teardown — process exit, or cuCtxDestroy / primary release / device reset: inventory goes.
code
# See processes holding the GPU (contexts behind them) nvidia-smi # Compute mode matters for multi-process sharing policies nvidia-smi | grep -i 'Compute Mode'

Traps

SymptomCheck first
INVALID_CONTEXT on workerDid this thread call cudaSetDevice / set current?
INVALID_CONTEXT after library returnUnbalanced push/pop or destroy?
Weird pointer / illegal addressBuffer allocated under a different context or device?
Works in main, dies after forkCUDA init before fork
Multi-process low utilPrivate contexts time-sliced — consider MPS or one process
CONTEXT_IS_DESTROYEDDestroy while another thread still had it current
Retain but still invalidPrimary retain without SetCurrent

What to do

  1. Treat context as ownership inventory and stream as ordering — fix isolation bugs before tuning streams.
  2. Call cudaSetDevice on every thread that will issue CUDA work (or push an explicit context).
  3. Prefer one primary context per device per process; avoid dual contexts on one GPU unless a library requires it.
  4. If you push, pop on every path; libraries should not leave a foreign current.
  5. Never fork after CUDA init; size context lifetime to the process or a clear library enter/exit.
  6. For many small multi-process jobs on one GPU, evaluate MPS; keep devices warm with persistenced when jobs are short.
GPU & High-Performance Computing
CUDA Context vs Streams vs MPS: Which Layer Fixes What

Decision map for CUDA: a context is per-process GPU state, a stream is an in-order queue inside a context, and MPS shares one context across processes. Pick the layer that matches the problem.

GPU & High-Performance Computing
CUDA Streams: Asynchronous Execution and Concurrency

A CUDA stream is an in-order queue of GPU ops. Overlap H→D, kernels, and D→H across streams — and avoid the default-stream trap that serializes everything.

GPU & High-Performance Computing
CUDA Multi-Process Service (MPS): Sharing One GPU Context

Why exclusive CUDA contexts leave SMs idle under multi-process load, how MPS multiplexes clients through a shared context, thread percentage caps, and when to pick exclusive, MPS, or MIG.

GPU & High-Performance Computing
NVIDIA Device Files in /dev/

How /dev/nvidia*, nvidiactl, nvidia-uvm, and DRI nodes map major/minor numbers to the driver, what CUDA opens first, and the minimum mount set for containers.

GPU & High-Performance Computing
NVIDIA vs AMD for Deep Learning: CUDA vs ROCm and the Datacenter Accelerators

NVIDIA vs AMD for deep learning compared at both layers: the CUDA vs ROCm software moat, the microarchitecture (warp vs wavefront, SM vs CU, Tensor vs Matrix Cores), and the datacenter accelerators (H100/H200/B200 vs MI300X/MI325X).

GPU & High-Performance Computing
Flynn's Classification: Taxonomy of Computer Architectures

Flynn's Classification explained — SISD, SIMD, MISD, MIMD with interactive architecture explorer, SIMD evolution from MMX to AMX, branch divergence visualization, and workload-architecture throughput comparison.

If you found this explanation helpful, consider sharing it with others.

Mastodon