Why THP exists
Every memory access needs a virtual → physical translation. Hits in the TLB cost a cycle or two. Misses walk multi-level page tables in RAM — expensive, and common when the working set is large.
With 4 KB pages, even a 1,536-entry second-level TLB only covers ~6 MB. A 10 GB database or model arena almost always misses. Transparent Huge Pages promote regions to 2 MB pages so one TLB entry maps 512× more memory, and the page walk is one level shorter. The kernel does this without rewriting the app — with real costs in fragmentation, latency, and memory bloat.
The page-table walk
On x86-64, a miss walks PGD → PUD → PMD → PTE. With a huge page the walk stops at the PMD (PSE bit set); the remaining bits are the offset inside the 2 MB page.
The 512× coverage jump
This is the core number. Same TLB silicon; larger pages; dramatically more address space cached.
Side-by-side walk depth:
How the kernel creates huge pages
Synchronous allocation
On a 2 MB-aligned fault, the kernel tries a contiguous 2 MB physical region. Success is cheap (~0.1–0.5 ms extra). Failure either falls back to 4 KB pages or compacts memory — compaction can stall the faulting thread for tens of milliseconds.
Asynchronous: khugepaged
khugepaged scans in the background and collapses 512 consecutive 4 KB pages into one 2 MB page when it can. No app stall, but promotion can lag seconds to minutes.
Scan rate and pages-per-pass are tunable: conservative (e.g. every 10 s) for quiet boxes; aggressive for higher coverage at background CPU cost.
Fragmentation
Free memory can be plentiful while contiguous 2 MB runs are gone. That is the main reason THP “does nothing” on long-uptime hosts.
| Mode | Behavior | Latency risk | Typical use |
|---|---|---|---|
| always | Compact on the fault path | Severe (10–100 ms) | Benchmarks only |
| defer | Background compact (kcompactd) | Low for the app | General servers |
| defer+madvise | Background + sync for madvise regions | Controlled | DB + explicit huge regions |
| madvise | Only where the app opts in | Low elsewhere | Production default |
| never | No compact for THP | None | Real-time / hard latency |
always looks attractive until p99 query latency spikes. Prefer madvise (or defer+madvise) and call madvise(MADV_HUGEPAGE) on large, long-lived arenas.
Where it helps
Best fit: large, dense sequential working sets — buffer pools, training tensors, columnar scans. Typical published-style gains: databases ~15–35% throughput when TLB was the bottleneck; ML training often single-digit to mid-teens percent. Always measure.
| Workload | Direction | Suggested mode |
|---|---|---|
| PostgreSQL / Mongo / Redis (large pool) | Often strong win | madvise |
| Spark / ClickHouse style scans | Strong win | madvise |
| PyTorch / large tensors | Moderate win | defer or madvise |
| Prefork / fork-heavy servers | Often hurt | never |
| Sparse random touch | Often hurt | never |
The cost: memory bloat
A small allocation can still be backed by a full 2 MB page. Internal fragmentation balloons RSS for many-tiny-object heaps.
Redis, Node, and some GC region layouts are frequent casualties of enabled=always. Bloat also worsens reclaim and swap (whole 2 MB chunks).
When THP hurts
- Fork + CoW — writing a huge page after
forkcopies 2 MB, not 4 KB. - Sparse access — fault brings 2 MB you never read.
- Churny allocators — khugepaged never stabilizes a region; scan cost with no payoff.
Configuration starting points
| Workload | enabled | defrag | khugepaged |
|---|---|---|---|
| Large buffer-pool DB | madvise | madvise | Aggressive |
| Fork-based web | never | never | Off |
| ML training | always or defer | defer | Moderate |
| Redis / Memcached | madvise | madvise | Moderate |
| Hard real-time | never | never | Off |
# Inspect cat /sys/kernel/mm/transparent_hugepage/enabled cat /sys/kernel/mm/transparent_hugepage/defrag # Safer production defaults (example) echo madvise | sudo tee /sys/kernel/mm/transparent_hugepage/enabled echo madvise | sudo tee /sys/kernel/mm/transparent_hugepage/defrag
Apps opt in with madvise(addr, len, MADV_HUGEPAGE) and opt out with MADV_NOHUGEPAGE on small-object heaps.
Related concepts
Master virtual memory and TLB address translation with interactive demos. Learn page tables, page faults, and memory management optimization.
Master sequential vs strided memory access patterns. Learn how cache efficiency and hardware prefetching affect application performance.
Structure of Arrays vs Array of Structures as instruments: cache-line fill, SIMD gather vs contiguous load, GPU coalescing, and AoSoA hybrids — when layout is a 10× decision.
Deep dive into CPU cache lines — interactive cache simulator with configurable associativity and replacement policies, false sharing MESI protocol visualization, access pattern benchmarks, and optimization techniques.
Explore CPU pipeline stages, instruction-level parallelism, pipeline hazards, and branch prediction through interactive visualizations.
Master pipeline hazards through interactive visualizations of data dependencies, control hazards, structural conflicts, and advanced detection mechanisms.
