Skip to main content

NVIDIA GPU Operator: Declarative GPU Stacks on Kubernetes

Summary
How the NVIDIA GPU Operator turns drivers, container toolkit, device plugin, GFD, DCGM, and MIG into DaemonSets reconciling from a ClusterPolicy — boot order, request path, sharing modes, and ops traps.

Thirty-two GPU nodes, one tired runbook

You add a GPU node. The old playbook still works if you enjoy pain: SSH, pin a driver, install the container toolkit, edit containerd, restart the runtime, apply a device-plugin DaemonSet, check nvidia-smi, then do it again for the next machine. By node thirty-two the cluster has three driver versions, two nodes where the plugin never came up, and a wiki page last edited by someone who left.

The NVIDIA GPU Operator exists so that story becomes a ClusterPolicy and a set of DaemonSets. Desired state lives in Git. Nodes converge. Failed plugin pods restart without a human SSH.

This page is the machinery — not a Helm flag encyclopedia.

Two numbers

1. Manual toil scales with nodes
Rough model: about five host steps per GPU node on bootstrap (driver, toolkit, runtime, plugin, verification). At 32 nodes that is ~160 manual actions before you count upgrades.

2. Operator toil is mostly constant
One install + pinned values. The controller keeps every GPU node’s stack matching policy. New nodes that join with the right labels/taints get DaemonSets for free.

Flip node count in the instrument; watch the sticky step count jump on manual and stay flat on operator.

What the operator actually is

Kubernetes Operator pattern: a controller extends the API with a custom resource and continuously reconciles the world toward that resource.

For GPU Operator the CRD is typically ClusterPolicy (via the Helm chart). The controller Deployment:

  1. Discovers GPU-capable nodes (and respects labels / feature discovery).
  2. Creates or updates DaemonSets for driver, toolkit, device plugin, GFD, DCGM, optional MIG manager, NFD pieces.
  3. Reacts when pods die, versions change, or policy fields flip.

It does not replace the need for:

  • Working GPU hardware and device nodes once the driver loads
  • Pods that request nvidia.com/gpu (or MIG resource names)
  • Cluster RBAC, storage for logs/metrics, and sane node OS images

Stack: one layer at a time

Everything that makes a node “GPU-ready” is a separate product. The operator ships them as coordinated DaemonSets.

LayerJob in one line
ControllerReconcile ClusterPolicy → DaemonSets
DriverLoad NVIDIA kmod into the host kernel (or defer to host driver)
Container toolkitConfigure runtime so containers get devices + driver libs
Device pluginAdvertise and allocate nvidia.com/gpu (or MIG)
GFDLabel nodes with product / compute / memory
DCGM exporterPrometheus metrics (:9400)
MIG managerPartition A100/H100 when enabled
NFDBroader hardware labels (CPU, PCI, …)

Boot order is real

You cannot advertise GPUs before the driver works. You cannot inject devices before the toolkit has configured the runtime. Phase 3 components (plugin, GFD, DCGM) only make sense after 1 and 2.

What “Ready” looks like in the cluster

code
kubectl get pods -n gpu-operator # driver, toolkit, device-plugin, gfd, dcgm-exporter … Running kubectl get nodes -o json \ | jq '.items[] | {name: .metadata.name, gpu: .status.allocatable["nvidia.com/gpu"]}'

Empty allocatable usually means device plugin never registered — debug driver and toolkit, not the scheduler.

Request path: YAML → silicon

The operator prepares the node. Your chart still has to ask.

  1. Pod sets resources.limits["nvidia.com/gpu"] (or a MIG name).
  2. Scheduler picks a node with free capacity and matching selectors/taints.
  3. Kubelet + device plugin allocate specific device IDs.
  4. Container toolkit / CDI injects /dev/nvidia*, libraries, env.
  5. Workload runs; CUDA sees only what was allocated.
code
apiVersion: v1 kind: Pod metadata: name: gpu-smoke spec: restartPolicy: Never containers: - name: cuda image: nvidia/cuda:12.2.0-base-ubuntu22.04 command: ['nvidia-smi'] resources: limits: nvidia.com/gpu: 1 # If GPU nodes are tainted: # tolerations: # - key: nvidia.com/gpu # operator: Exists # effect: NoSchedule

Without the limit, the pod can still land on a GPU node and CUDA will report no devices — the runtime never injected anything.

Sharing: exclusive, time-slice, MIG

Default mental model: one GPU resource unit = one full device for one pod. Density knobs:

ModeIsolationResource nameTypical use
ExclusiveFull devicenvidia.com/gpuTraining, heavy inference
Time-slicingSoft; shared memoryOversubscribed nvidia.com/gpuMany small trusted jobs
MIGHardware partitionsnvidia.com/mig-…Multi-tenant A100/H100

MPS is a driver-side sharing path (shared CUDA context). Time-slicing and MIG are Kubernetes scheduling surfaces the operator configures. They solve related density problems at different layers.

code
# Sketch: time-slicing ConfigMap consumed by device plugin # (exact schema follows your operator chart version) version: v1 sharing: timeSlicing: resources: - name: nvidia.com/gpu replicas: 4 # advertise 4 slots per physical GPU
code
# Sketch: MIG profile fragment mig-configs: all-1g.5gb: - devices: all mig-enabled: true mig-devices: '1g.5gb': 7

Workloads must request the name you actually advertised. A chart still saying nvidia.com/gpu: 1 will not consume nvidia.com/mig-1g.5gb.

Install shape (pinned)

code
helm repo add nvidia https://helm.ngc.nvidia.com/nvidia helm repo update helm upgrade --install gpu-operator nvidia/gpu-operator \ --namespace gpu-operator \ --create-namespace \ --version <chart-version> \ -f values.yaml
code
# values.yaml (illustrative — check chart docs for your version) driver: enabled: true version: '535.129.03' # pin in production toolkit: enabled: true devicePlugin: enabled: true dcgmExporter: enabled: true mig: strategy: none # or single / mixed

Host-driver mode (when nodes already have a managed driver):

code
driver: enabled: false # toolkit + plugin still run; they expect a working host nvidia-smi

Day-2 ops that matter

Pin versions

Never ship latest for driver or chart in production. Canary a node pool with the new driver, then roll.

Taints and labels

code
kubectl taint nodes gpu-1 nvidia.com/gpu=present:NoSchedule # Workloads need matching tolerations

GFD labels (nvidia.com/gpu.product, compute capability, memory) feed nodeSelector / affinity for SKU placement.

Quotas

code
apiVersion: v1 kind: ResourceQuota metadata: name: gpu-quota namespace: research spec: hard: requests.nvidia.com/gpu: '8' limits.nvidia.com/gpu: '8'

Metrics

DCGM exporter feeds Prometheus. Alert on high temp, ECC, unexpected zero util with pending GPU pods, or plugin restarts.

Persistence mode

Short jobs still pay cold-driver tax unless the host keeps the device warm via nvidia-persistenced. The operator does not replace that host daemon story.

Traps

Quick debug loop

code
# 1. Operator and policy kubectl get pods -n gpu-operator kubectl get clusterpolicies.nvidia.com -A # 2. Driver kubectl logs -n gpu-operator -l app=nvidia-driver-daemonset --tail=100 kubectl exec -n gpu-operator ds/nvidia-driver-daemonset -- nvidia-smi # 3. Allocatable kubectl describe node <gpu-node> | grep -A2 'Allocatable\|nvidia.com' # 4. Plugin kubectl logs -n gpu-operator -l app=nvidia-device-plugin-daemonset --tail=100

What to do

  1. Choose driver ownership — operator container or host package, never both by accident.
  2. Pin chart + driver versions; canary upgrades.
  3. Verify boot phases before blaming the scheduler (nvidia.com/gpu allocatable).
  4. Always request the extended resource in pods; pair taints with tolerations.
  5. Match sharing mode to tenancy: exclusive for big jobs, time-slice for trusted density, MIG for hard partitions, MPS when the problem is CUDA context co-residency rather than K8s advertising.
GPU & High-Performance Computing
CUDA Contexts: Ownership, Current Stack, Isolation

Deep dive into the CUDA context object: control vs data plane, inventory (memory, modules, streams, events, graphs), push/pop/setCurrent stacks, primary retain/release, flags and limits, isolation, cost, and traps.

GPU & High-Performance Computing
CUDA Context vs Streams vs MPS: Which Layer Fixes What

Decision map for CUDA: a context is per-process GPU state, a stream is an in-order queue inside a context, and MPS shares one context across processes. Pick the layer that matches the problem.

GPU & High-Performance Computing
CUDA Multi-Process Service (MPS): Sharing One GPU Context

Why exclusive CUDA contexts leave SMs idle under multi-process load, how MPS multiplexes clients through a shared context, thread percentage caps, and when to pick exclusive, MPS, or MIG.

GPU & High-Performance Computing
CUDA Streams: Asynchronous Execution and Concurrency

A CUDA stream is an in-order queue of GPU ops. Overlap H→D, kernels, and D→H across streams — and avoid the default-stream trap that serializes everything.

GPU & High-Performance Computing
NVIDIA Device Files in /dev/

How /dev/nvidia*, nvidiactl, nvidia-uvm, and DRI nodes map major/minor numbers to the driver, what CUDA opens first, and the minimum mount set for containers.

GPU & High-Performance Computing
NVIDIA vs AMD for Deep Learning: CUDA vs ROCm and the Datacenter Accelerators

NVIDIA vs AMD for deep learning compared at both layers: the CUDA vs ROCm software moat, the microarchitecture (warp vs wavefront, SM vs CU, Tensor vs Matrix Cores), and the datacenter accelerators (H100/H200/B200 vs MI300X/MI325X).

If you found this explanation helpful, consider sharing it with others.

Mastodon