Skip to main content

Linux cgroups: Resource Limits for Processes

Summary
Control groups as instruments: noisy neighbors, v1 vs v2 membership, CFS cpu.max periods, memory.high vs memory.max OOM, and how Docker flags become cgroup files.

One tenant threw a party. The whole machine paid.

A batch job spikes. Latency on the API next door climbs. Memory pressure trips the host OOM killer and kills the wrong process. Nothing was “broken” — nothing was metered. Linux control groups exist so a set of processes can have a budget: how much CPU time, how much RAM, how many PIDs, how much I/O.

Namespaces hide the world. Cgroups price it. Containers need both.

Two numbers

1. 50000 100000
cpu.max as quota and period in microseconds. That pair is half of one core: spend 50 ms of CPU every 100 ms, then throttle until the period resets. Docker --cpus=0.5 writes this file.

2. memory.max
Hard ceiling in bytes. Cross it after reclaim fails and the kernel OOM-kills inside the cgroup — the container dies, the neighbor lives. Soft landing lives one file over: memory.high.

The neighbor without a budget

Flip limits off. The louder tenant takes almost every host slice. Flip them on. Each cgroup is capped; the kernel steals time back at the period boundary.

This is multi-tenant failure in miniature. Shared CI runners, model-serving replicas, and “just run it on the same box” all hit the same strip when nothing enforces a budget.

Hide vs meter

Namespaces answer visibility. Cgroups answer spend. Alone, each leaves a hole.

systemd slices are often cgroup-only (resource accounting for a service). Unshare without cgroup writes is isolation without a leash. Containers assemble both before the entrypoint runs.

One process, one home (or three)

v1 mounts a tree per controller. The same PID can sit under /cpu/… in one place and /memory/… in another. v2 is a single tree under /sys/fs/cgroup: enable controllers on a parent, set limits on the leaf that holds cgroup.procs.

Modern distros default to v2. Hybrid setups exist for legacy containers; treat mixed v1/v2 for the same controller as a footgun, not a feature.

The control plane is a filesystem

code
# leaf cgroup mkdir /sys/fs/cgroup/myapp echo "+cpu +memory +io +pids" > /sys/fs/cgroup/myapp/cgroup.subtree_control mkdir /sys/fs/cgroup/myapp/worker echo 536870912 > /sys/fs/cgroup/myapp/worker/memory.max echo "50000 100000" > /sys/fs/cgroup/myapp/worker/cpu.max echo $$ > /sys/fs/cgroup/myapp/worker/cgroup.procs

Everything else — Docker flags, systemd unit keys, Kubernetes limits — eventually writes files like these.

CPU: spend, then wait

CFS bandwidth is a period tape. You get a quota of microseconds each period. While quota remains, the cgroup runs. When it is spent, the cgroup is forced idle until the next period — even if the host is otherwise free.

Wantcpu.max (period 100000)
25% of one core25000 100000
half a core50000 100000
one full core100000 100000
two cores200000 100000
unlimitedmax 100000

cpu.weight is the other knob: proportional share under contention when there is no hard quota. Latency-sensitive services often hate tiny periods with tiny quotas — the throttle edges show up as 10–100 ms stalls. Raise the period or the quota; do not only raise priority with nice.

Memory: soft high, hard max

Usage climbs. Cross memory.high and the kernel starts reclaim and throttling — painful, not fatal. Cross memory.max and the cgroup OOM path fires: a victim inside the group gets SIGKILL.

FileRole
memory.currentlive usage
memory.highsoft ceiling · reclaim / throttle
memory.maxhard ceiling · cgroup OOM
memory.min / memory.lowprotection floors under pressure
memory.swap.maxswap budget (0 = no quiet death into swap)
memory.eventsoom / oom_kill counters

Page cache and kernel structures count. A “512 MiB heap” container can still OOM with 400 MiB RSS if buffers filled the rest of the budget.

Docker is a writer of those files

No new mechanism — a directory under the runtime’s cgroup parent, then cpu.max, memory.max, pids.max, sometimes io.max.

code
docker run \ --cpus=0.5 \ --memory=512m \ --memory-swap=512m \ --pids-limit=100 \ nginx

systemd maps the same ideas to unit keys: CPUQuota=50%, MemoryMax=512M, TasksMax=100. systemd-cgls and systemd-cgtop are the live tree and the scoreboard.

Pressure before failure

PSI (Pressure Stall Information) on cpu.pressure, memory.pressure, and io.pressure reports how often some or all tasks in the cgroup stalled waiting for that resource. Watch avg10 climb before you only notice OOM kills.

code
cat /sys/fs/cgroup/docker/abc123/memory.pressure # some avg10=5.23 avg60=3.15 … # full avg10=2.11 …

What to do

  1. Prefer cgroups v2 unified hierarchy; enable controllers with cgroup.subtree_control, set limits on the leaf that owns the tasks.
  2. Set memory.max for hard multi-tenant isolation; add memory.high when you want reclaim before kill.
  3. Translate CPU intent into cpu.max quota/period math — do not confuse weight with a hard cap.
  4. Cap PIDs on untrusted workloads (pids.max) so a fork bomb dies inside the group.
  5. On hosts that already run systemd, tune unit files (CPUQuota=, MemoryMax=) instead of hand-rolling trees when a service is the unit of isolation.
  6. Read memory.events and PSI under load; fix budgets from evidence, not from “it felt slow.”

When something smaller is enough

  • One batch job should yield under contention → nice may be enough.
  • Per-user FD or nproc caps → limits.conf / ulimit.
  • CPU affinity without bandwidth → taskset, not a full cgroup tree.
  • Short one-off process → setup cost can exceed the run; skip the hierarchy.

If you found this explanation helpful, consider sharing it with others.

Mastodon