One tenant threw a party. The whole machine paid.
A batch job spikes. Latency on the API next door climbs. Memory pressure trips the host OOM killer and kills the wrong process. Nothing was “broken” — nothing was metered. Linux control groups exist so a set of processes can have a budget: how much CPU time, how much RAM, how many PIDs, how much I/O.
Namespaces hide the world. Cgroups price it. Containers need both.
Two numbers
1. 50000 100000
cpu.max as quota and period in microseconds. That pair is half of one core: spend 50 ms of CPU every 100 ms, then throttle until the period resets. Docker --cpus=0.5 writes this file.
2. memory.max
Hard ceiling in bytes. Cross it after reclaim fails and the kernel OOM-kills inside the cgroup — the container dies, the neighbor lives. Soft landing lives one file over: memory.high.
The neighbor without a budget
Flip limits off. The louder tenant takes almost every host slice. Flip them on. Each cgroup is capped; the kernel steals time back at the period boundary.
This is multi-tenant failure in miniature. Shared CI runners, model-serving replicas, and “just run it on the same box” all hit the same strip when nothing enforces a budget.
Hide vs meter
Namespaces answer visibility. Cgroups answer spend. Alone, each leaves a hole.
systemd slices are often cgroup-only (resource accounting for a service). Unshare without cgroup writes is isolation without a leash. Containers assemble both before the entrypoint runs.
One process, one home (or three)
v1 mounts a tree per controller. The same PID can sit under /cpu/… in one place and /memory/… in another. v2 is a single tree under /sys/fs/cgroup: enable controllers on a parent, set limits on the leaf that holds cgroup.procs.
Modern distros default to v2. Hybrid setups exist for legacy containers; treat mixed v1/v2 for the same controller as a footgun, not a feature.
The control plane is a filesystem
# leaf cgroup mkdir /sys/fs/cgroup/myapp echo "+cpu +memory +io +pids" > /sys/fs/cgroup/myapp/cgroup.subtree_control mkdir /sys/fs/cgroup/myapp/worker echo 536870912 > /sys/fs/cgroup/myapp/worker/memory.max echo "50000 100000" > /sys/fs/cgroup/myapp/worker/cpu.max echo $$ > /sys/fs/cgroup/myapp/worker/cgroup.procs
Everything else — Docker flags, systemd unit keys, Kubernetes limits — eventually writes files like these.
CPU: spend, then wait
CFS bandwidth is a period tape. You get a quota of microseconds each period. While quota remains, the cgroup runs. When it is spent, the cgroup is forced idle until the next period — even if the host is otherwise free.
| Want | cpu.max (period 100000) |
|---|---|
| 25% of one core | 25000 100000 |
| half a core | 50000 100000 |
| one full core | 100000 100000 |
| two cores | 200000 100000 |
| unlimited | max 100000 |
cpu.weight is the other knob: proportional share under contention when there is no hard quota. Latency-sensitive services often hate tiny periods with tiny quotas — the throttle edges show up as 10–100 ms stalls. Raise the period or the quota; do not only raise priority with nice.
Memory: soft high, hard max
Usage climbs. Cross memory.high and the kernel starts reclaim and throttling — painful, not fatal. Cross memory.max and the cgroup OOM path fires: a victim inside the group gets SIGKILL.
| File | Role |
|---|---|
memory.current | live usage |
memory.high | soft ceiling · reclaim / throttle |
memory.max | hard ceiling · cgroup OOM |
memory.min / memory.low | protection floors under pressure |
memory.swap.max | swap budget (0 = no quiet death into swap) |
memory.events | oom / oom_kill counters |
Page cache and kernel structures count. A “512 MiB heap” container can still OOM with 400 MiB RSS if buffers filled the rest of the budget.
Docker is a writer of those files
No new mechanism — a directory under the runtime’s cgroup parent, then cpu.max, memory.max, pids.max, sometimes io.max.
docker run \ --cpus=0.5 \ --memory=512m \ --memory-swap=512m \ --pids-limit=100 \ nginx
systemd maps the same ideas to unit keys: CPUQuota=50%, MemoryMax=512M, TasksMax=100. systemd-cgls and systemd-cgtop are the live tree and the scoreboard.
Pressure before failure
PSI (Pressure Stall Information) on cpu.pressure, memory.pressure, and io.pressure reports how often some or all tasks in the cgroup stalled waiting for that resource. Watch avg10 climb before you only notice OOM kills.
cat /sys/fs/cgroup/docker/abc123/memory.pressure # some avg10=5.23 avg60=3.15 … # full avg10=2.11 …
What to do
- Prefer cgroups v2 unified hierarchy; enable controllers with
cgroup.subtree_control, set limits on the leaf that owns the tasks. - Set
memory.maxfor hard multi-tenant isolation; addmemory.highwhen you want reclaim before kill. - Translate CPU intent into
cpu.maxquota/period math — do not confuse weight with a hard cap. - Cap PIDs on untrusted workloads (
pids.max) so a fork bomb dies inside the group. - On hosts that already run systemd, tune unit files (
CPUQuota=,MemoryMax=) instead of hand-rolling trees when a service is the unit of isolation. - Read
memory.eventsand PSI under load; fix budgets from evidence, not from “it felt slow.”
When something smaller is enough
- One batch job should yield under contention →
nicemay be enough. - Per-user FD or nproc caps →
limits.conf/ulimit. - CPU affinity without bandwidth →
taskset, not a full cgroup tree. - Short one-off process → setup cost can exceed the run; skip the hierarchy.
Related concepts
Discover how containers work by combining namespaces, cgroups, and OverlayFS. Build a mental model of Docker internals through interactive visualizations.
Understand how containerized processes access GPU hardware through device files, bind mounts, and the NVIDIA container runtime. Learn the kernel driver vs user-space library distinction.
Master Linux namespaces — the kernel mechanism that makes containers possible. Learn how mount, PID, network, and user namespaces create isolated environments, with interactive demos.
How /dev/nvidia*, nvidiactl, nvidia-uvm, and DRI nodes map major/minor numbers to the driver, what CUDA opens first, and the minimum mount set for containers.
Visualize the complete Linux boot sequence from BIOS/UEFI to login. Learn how GRUB, kernel, and systemd work together with interactive visualizations.
Learn the Btrfs filesystem with built-in snapshots, RAID, and compression. Explore copy-on-write, subvolumes, and self-healing on Linux.
