Slurm Resource Management and Job Priority
How Slurm decides which jobs run first — priority factors, fair-share scheduling, backfill, and monitoring commands (squeue, sinfo, sacct).
8 min readConcept
Explore machine learning concepts related to Resource Management. Clear explanations and practical insights.
How Slurm decides which jobs run first — priority factors, fair-share scheduling, backfill, and monitoring commands (squeue, sinfo, sacct).
How Slurm tracks resource consumption through account hierarchies, TRES billing, and resource limits — sacctmgr, sreport, and the association model explained.
Control groups as instruments: noisy neighbors, v1 vs v2 membership, CFS cpu.max periods, memory.high vs memory.max OOM, and how Docker flags become cgroup files.