Thirty-two GPU nodes, one tired runbook
You add a GPU node. The old playbook still works if you enjoy pain: SSH, pin a driver, install the container toolkit, edit containerd, restart the runtime, apply a device-plugin DaemonSet, check nvidia-smi, then do it again for the next machine. By node thirty-two the cluster has three driver versions, two nodes where the plugin never came up, and a wiki page last edited by someone who left.
The NVIDIA GPU Operator exists so that story becomes a ClusterPolicy and a set of DaemonSets. Desired state lives in Git. Nodes converge. Failed plugin pods restart without a human SSH.
This page is the machinery — not a Helm flag encyclopedia.
Two numbers
1. Manual toil scales with nodes
Rough model: about five host steps per GPU node on bootstrap (driver, toolkit, runtime, plugin, verification). At 32 nodes that is ~160 manual actions before you count upgrades.
2. Operator toil is mostly constant
One install + pinned values. The controller keeps every GPU node’s stack matching policy. New nodes that join with the right labels/taints get DaemonSets for free.
Flip node count in the instrument; watch the sticky step count jump on manual and stay flat on operator.
What the operator actually is
Kubernetes Operator pattern: a controller extends the API with a custom resource and continuously reconciles the world toward that resource.
For GPU Operator the CRD is typically ClusterPolicy (via the Helm chart). The controller Deployment:
- Discovers GPU-capable nodes (and respects labels / feature discovery).
- Creates or updates DaemonSets for driver, toolkit, device plugin, GFD, DCGM, optional MIG manager, NFD pieces.
- Reacts when pods die, versions change, or policy fields flip.
It does not replace the need for:
- Working GPU hardware and device nodes once the driver loads
- Pods that request
nvidia.com/gpu(or MIG resource names) - Cluster RBAC, storage for logs/metrics, and sane node OS images
Stack: one layer at a time
Everything that makes a node “GPU-ready” is a separate product. The operator ships them as coordinated DaemonSets.
| Layer | Job in one line |
|---|---|
| Controller | Reconcile ClusterPolicy → DaemonSets |
| Driver | Load NVIDIA kmod into the host kernel (or defer to host driver) |
| Container toolkit | Configure runtime so containers get devices + driver libs |
| Device plugin | Advertise and allocate nvidia.com/gpu (or MIG) |
| GFD | Label nodes with product / compute / memory |
| DCGM exporter | Prometheus metrics (:9400) |
| MIG manager | Partition A100/H100 when enabled |
| NFD | Broader hardware labels (CPU, PCI, …) |
Boot order is real
You cannot advertise GPUs before the driver works. You cannot inject devices before the toolkit has configured the runtime. Phase 3 components (plugin, GFD, DCGM) only make sense after 1 and 2.
What “Ready” looks like in the cluster
kubectl get pods -n gpu-operator # driver, toolkit, device-plugin, gfd, dcgm-exporter … Running kubectl get nodes -o json \ | jq '.items[] | {name: .metadata.name, gpu: .status.allocatable["nvidia.com/gpu"]}'
Empty allocatable usually means device plugin never registered — debug driver and toolkit, not the scheduler.
Request path: YAML → silicon
The operator prepares the node. Your chart still has to ask.
- Pod sets
resources.limits["nvidia.com/gpu"](or a MIG name). - Scheduler picks a node with free capacity and matching selectors/taints.
- Kubelet + device plugin allocate specific device IDs.
- Container toolkit / CDI injects
/dev/nvidia*, libraries, env. - Workload runs; CUDA sees only what was allocated.
apiVersion: v1 kind: Pod metadata: name: gpu-smoke spec: restartPolicy: Never containers: - name: cuda image: nvidia/cuda:12.2.0-base-ubuntu22.04 command: ['nvidia-smi'] resources: limits: nvidia.com/gpu: 1 # If GPU nodes are tainted: # tolerations: # - key: nvidia.com/gpu # operator: Exists # effect: NoSchedule
Without the limit, the pod can still land on a GPU node and CUDA will report no devices — the runtime never injected anything.
Sharing: exclusive, time-slice, MIG
Default mental model: one GPU resource unit = one full device for one pod. Density knobs:
| Mode | Isolation | Resource name | Typical use |
|---|---|---|---|
| Exclusive | Full device | nvidia.com/gpu | Training, heavy inference |
| Time-slicing | Soft; shared memory | Oversubscribed nvidia.com/gpu | Many small trusted jobs |
| MIG | Hardware partitions | nvidia.com/mig-… | Multi-tenant A100/H100 |
MPS is a driver-side sharing path (shared CUDA context). Time-slicing and MIG are Kubernetes scheduling surfaces the operator configures. They solve related density problems at different layers.
# Sketch: time-slicing ConfigMap consumed by device plugin # (exact schema follows your operator chart version) version: v1 sharing: timeSlicing: resources: - name: nvidia.com/gpu replicas: 4 # advertise 4 slots per physical GPU
# Sketch: MIG profile fragment mig-configs: all-1g.5gb: - devices: all mig-enabled: true mig-devices: '1g.5gb': 7
Workloads must request the name you actually advertised. A chart still saying nvidia.com/gpu: 1 will not consume nvidia.com/mig-1g.5gb.
Install shape (pinned)
helm repo add nvidia https://helm.ngc.nvidia.com/nvidia helm repo update helm upgrade --install gpu-operator nvidia/gpu-operator \ --namespace gpu-operator \ --create-namespace \ --version <chart-version> \ -f values.yaml
# values.yaml (illustrative — check chart docs for your version) driver: enabled: true version: '535.129.03' # pin in production toolkit: enabled: true devicePlugin: enabled: true dcgmExporter: enabled: true mig: strategy: none # or single / mixed
Host-driver mode (when nodes already have a managed driver):
driver: enabled: false # toolkit + plugin still run; they expect a working host nvidia-smi
Day-2 ops that matter
Pin versions
Never ship latest for driver or chart in production. Canary a node pool with the new driver, then roll.
Taints and labels
kubectl taint nodes gpu-1 nvidia.com/gpu=present:NoSchedule # Workloads need matching tolerations
GFD labels (nvidia.com/gpu.product, compute capability, memory) feed nodeSelector / affinity for SKU placement.
Quotas
apiVersion: v1 kind: ResourceQuota metadata: name: gpu-quota namespace: research spec: hard: requests.nvidia.com/gpu: '8' limits.nvidia.com/gpu: '8'
Metrics
DCGM exporter feeds Prometheus. Alert on high temp, ECC, unexpected zero util with pending GPU pods, or plugin restarts.
Persistence mode
Short jobs still pay cold-driver tax unless the host keeps the device warm via nvidia-persistenced. The operator does not replace that host daemon story.
Traps
Quick debug loop
# 1. Operator and policy kubectl get pods -n gpu-operator kubectl get clusterpolicies.nvidia.com -A # 2. Driver kubectl logs -n gpu-operator -l app=nvidia-driver-daemonset --tail=100 kubectl exec -n gpu-operator ds/nvidia-driver-daemonset -- nvidia-smi # 3. Allocatable kubectl describe node <gpu-node> | grep -A2 'Allocatable\|nvidia.com' # 4. Plugin kubectl logs -n gpu-operator -l app=nvidia-device-plugin-daemonset --tail=100
What to do
- Choose driver ownership — operator container or host package, never both by accident.
- Pin chart + driver versions; canary upgrades.
- Verify boot phases before blaming the scheduler (
nvidia.com/gpuallocatable). - Always request the extended resource in pods; pair taints with tolerations.
- Match sharing mode to tenancy: exclusive for big jobs, time-slice for trusted density, MIG for hard partitions, MPS when the problem is CUDA context co-residency rather than K8s advertising.
Related concepts
Deep dive into the CUDA context object: control vs data plane, inventory (memory, modules, streams, events, graphs), push/pop/setCurrent stacks, primary retain/release, flags and limits, isolation, cost, and traps.
Decision map for CUDA: a context is per-process GPU state, a stream is an in-order queue inside a context, and MPS shares one context across processes. Pick the layer that matches the problem.
Why exclusive CUDA contexts leave SMs idle under multi-process load, how MPS multiplexes clients through a shared context, thread percentage caps, and when to pick exclusive, MPS, or MIG.
A CUDA stream is an in-order queue of GPU ops. Overlap H→D, kernels, and D→H across streams — and avoid the default-stream trap that serializes everything.
How /dev/nvidia*, nvidiactl, nvidia-uvm, and DRI nodes map major/minor numbers to the driver, what CUDA opens first, and the minimum mount set for containers.
NVIDIA vs AMD for deep learning compared at both layers: the CUDA vs ROCm software moat, the microarchitecture (warp vs wavefront, SM vs CU, Tensor vs Matrix Cores), and the datacenter accelerators (H100/H200/B200 vs MI300X/MI325X).
