Skip to main content

NVIDIA Device Files in /dev/

Summary
How /dev/nvidia*, nvidiactl, nvidia-uvm, and DRI nodes map major/minor numbers to the driver, what CUDA opens first, and the minimum mount set for containers.

The three paths that were not there

The image has CUDA. The host has an A100. nvidia-smi on the node is fine. Inside the pod:

code
CUDA error: CUDA_ERROR_NOT_INITIALIZED

or a blunt openat(.../dev/nvidiactl) = -1 ENOENT. Nobody rewrote your kernel. The character device nodes the driver publishes under /dev/ never made it into the container — or the open order hit a missing node first.

These files are not decoration. They are the only userspace door into nvidia.ko, nvidia-uvm.ko, and friends. Persistence is about keeping those doors warm; this page is about what the doors are.

Two numbers

1. Three nodes for single-GPU CUDA
/dev/nvidia0 · /dev/nvidiactl · /dev/nvidia-uvm. Forget ctl or uvm and modern stacks look “broken” while ls still shows a GPU device.

2. Major picks the driver · minor picks the instance
Character device 195, 0 is not “file size.” It is “NVIDIA major, GPU 0 minor.” Same major family for nvidiactl (195, 255); UVM lives under major 237.

Flip the instruments until both feel mechanical.

Character devices, not disk files

The leading c in crw-rw-rw- means character device: unbuffered, ioctl-heavy, driver-backed. The pair after owner/group is major, minor.

code
$ ls -la /dev/nvidia* crw-rw-rw- 1 root root 195, 0 /dev/nvidia0 crw-rw-rw- 1 root root 195, 1 /dev/nvidia1 crw-rw-rw- 1 root root 195, 255 /dev/nvidiactl crw-rw-rw- 1 root root 195, 254 /dev/nvidia-modeset crw-rw-rw- 1 root root 237, 0 /dev/nvidia-uvm

Major 195 → classic NVIDIA control/compute nodes. Major 237 → UVM. Major 226 → DRM (/dev/dri/*). Chip through the decoder:

Open order: ctl first

Libraries do not start at /dev/nvidia0. They open /dev/nvidiactl to talk to the driver, enumerate GPUs, then open the per-GPU node, then (for modern paths) UVM.

Typical userspace shape (illustrative ioctl names — real IOCTLs are private):

code
int ctl = open("/dev/nvidiactl", O_RDWR); // enum + driver queries on ctl int gpu = open("/dev/nvidia0", O_RDWR); int uvm = open("/dev/nvidia-uvm", O_RDWR); // then cudaSetDevice / contexts / allocs

UVM is on the floor, not the balcony

cudaMallocManaged, page migration, and many framework init paths need /dev/nvidia-uvm. Mounting only nvidia0 is a classic incomplete hand---device list.

/dev/nvidia-uvm-tools (minor 1) is optional — profilers and stats. Do not confuse it with the required UVM node.

Container: mount the set for the workload

Prefer nvidia-container-toolkit (--gpus) or the Kubernetes device plugin over a permanent hand-curated --device list. When you must reason about the list, match the workload:

code
# Prefer toolkit docker run --rm --gpus all nvidia/cuda:12.6.0-base nvidia-smi # Hand mount (debug only) — single GPU CUDA floor docker run --rm \ --device=/dev/nvidia0 \ --device=/dev/nvidiactl \ --device=/dev/nvidia-uvm \ nvidia/cuda:12.6.0-base nvidia-smi

MIG and fabric add /dev/nvidia-caps/*. Display stacks add modeset and /dev/dri/*.

DRI: card vs render

DRM nodes are a parallel Linux interface (major 226), not a rename of nvidia0. Pure CUDA often never opens them. Desktop and headless GL do.

NodeTypical useGroup
/dev/dri/card0Display + modesetvideo
/dev/dri/renderD128Headless render / compute via DRMrender
/dev/nvidia-modesetNVIDIA KMSvaries

Permissions and traps

Default 0666 on some installs hides access bugs until you harden to 0660 + groups. Missing nodes, incomplete mounts, and wrong GPU indices show up as “CUDA is broken.”

code
# See the first failing open strace -e openat python -c "import torch" 2>&1 | grep nvidia # Host nodes absent? sudo modprobe nvidia nvidia_uvm sudo nvidia-modprobe -u # Groups (after hardening MODE=0660) sudo usermod -aG video,render $USER # re-login or newgrp

Udev sketch (adjust to site policy):

code
# /etc/udev/rules.d/70-nvidia.rules KERNEL=="nvidia*", MODE="0660", GROUP="video" KERNEL=="nvidia_uvm*", MODE="0660", GROUP="video"

Manual mknod is last-resort recovery — wrong major/minor or no module behind the node is worse than an empty ls.

What to do

  1. On the host, confirm /dev/nvidiactl, /dev/nvidiaN, /dev/nvidia-uvm exist before debugging the app.
  2. For CUDA containers, mount ctl + gpu + uvm (or use --gpus / the device plugin).
  3. On NOT_INITIALIZED / weird init, strace the first openat of an nvidia path.
  4. Do not assume nvidia0 is the card in the top slot — check PCI mapping.
  5. Pair with persistence on busy nodes so cold init is not paid every short job.
Systems & Architecture
How Docker Works with GPUs: Device Files, Bind Mounts, and Driver Stacks

Understand how containerized processes access GPU hardware through device files, bind mounts, and the NVIDIA container runtime. Learn the kernel driver vs user-space library distinction.

GPU & High-Performance Computing
CUDA Contexts: Ownership, Current Stack, Isolation

Deep dive into the CUDA context object: control vs data plane, inventory (memory, modules, streams, events, graphs), push/pop/setCurrent stacks, primary retain/release, flags and limits, isolation, cost, and traps.

GPU & High-Performance Computing
CUDA Context vs Streams vs MPS: Which Layer Fixes What

Decision map for CUDA: a context is per-process GPU state, a stream is an in-order queue inside a context, and MPS shares one context across processes. Pick the layer that matches the problem.

GPU & High-Performance Computing
CUDA Multi-Process Service (MPS): Sharing One GPU Context

Why exclusive CUDA contexts leave SMs idle under multi-process load, how MPS multiplexes clients through a shared context, thread percentage caps, and when to pick exclusive, MPS, or MIG.

GPU & High-Performance Computing
CUDA Streams: Asynchronous Execution and Concurrency

A CUDA stream is an in-order queue of GPU ops. Overlap H→D, kernels, and D→H across streams — and avoid the default-stream trap that serializes everything.

GPU & High-Performance Computing
NVIDIA vs AMD for Deep Learning: CUDA vs ROCm and the Datacenter Accelerators

NVIDIA vs AMD for deep learning compared at both layers: the CUDA vs ROCm software moat, the microarchitecture (warp vs wavefront, SM vs CU, Tensor vs Matrix Cores), and the datacenter accelerators (H100/H200/B200 vs MI300X/MI325X).

If you found this explanation helpful, consider sharing it with others.

Mastodon