Skip to main content

Transparent Huge Pages (THP): Reducing TLB Pressure

Summary
Learn how Transparent Huge Pages (THP) reduces TLB misses by promoting 4KB to 2MB pages. Understand performance benefits and memory bloat tradeoffs.

Why THP exists

Every memory access needs a virtual → physical translation. Hits in the TLB cost a cycle or two. Misses walk multi-level page tables in RAM — expensive, and common when the working set is large.

With 4 KB pages, even a 1,536-entry second-level TLB only covers ~6 MB. A 10 GB database or model arena almost always misses. Transparent Huge Pages promote regions to 2 MB pages so one TLB entry maps 512× more memory, and the page walk is one level shorter. The kernel does this without rewriting the app — with real costs in fragmentation, latency, and memory bloat.

The page-table walk

On x86-64, a miss walks PGD → PUD → PMD → PTE. With a huge page the walk stops at the PMD (PSE bit set); the remaining bits are the offset inside the 2 MB page.

The 512× coverage jump

This is the core number. Same TLB silicon; larger pages; dramatically more address space cached.

Side-by-side walk depth:

How the kernel creates huge pages

Synchronous allocation

On a 2 MB-aligned fault, the kernel tries a contiguous 2 MB physical region. Success is cheap (~0.1–0.5 ms extra). Failure either falls back to 4 KB pages or compacts memory — compaction can stall the faulting thread for tens of milliseconds.

Asynchronous: khugepaged

khugepaged scans in the background and collapses 512 consecutive 4 KB pages into one 2 MB page when it can. No app stall, but promotion can lag seconds to minutes.

Scan rate and pages-per-pass are tunable: conservative (e.g. every 10 s) for quiet boxes; aggressive for higher coverage at background CPU cost.

Fragmentation

Free memory can be plentiful while contiguous 2 MB runs are gone. That is the main reason THP “does nothing” on long-uptime hosts.

ModeBehaviorLatency riskTypical use
alwaysCompact on the fault pathSevere (10–100 ms)Benchmarks only
deferBackground compact (kcompactd)Low for the appGeneral servers
defer+madviseBackground + sync for madvise regionsControlledDB + explicit huge regions
madviseOnly where the app opts inLow elsewhereProduction default
neverNo compact for THPNoneReal-time / hard latency

always looks attractive until p99 query latency spikes. Prefer madvise (or defer+madvise) and call madvise(MADV_HUGEPAGE) on large, long-lived arenas.

Where it helps

Best fit: large, dense sequential working sets — buffer pools, training tensors, columnar scans. Typical published-style gains: databases ~15–35% throughput when TLB was the bottleneck; ML training often single-digit to mid-teens percent. Always measure.

WorkloadDirectionSuggested mode
PostgreSQL / Mongo / Redis (large pool)Often strong winmadvise
Spark / ClickHouse style scansStrong winmadvise
PyTorch / large tensorsModerate windefer or madvise
Prefork / fork-heavy serversOften hurtnever
Sparse random touchOften hurtnever

The cost: memory bloat

A small allocation can still be backed by a full 2 MB page. Internal fragmentation balloons RSS for many-tiny-object heaps.

Redis, Node, and some GC region layouts are frequent casualties of enabled=always. Bloat also worsens reclaim and swap (whole 2 MB chunks).

When THP hurts

  • Fork + CoW — writing a huge page after fork copies 2 MB, not 4 KB.
  • Sparse access — fault brings 2 MB you never read.
  • Churny allocators — khugepaged never stabilizes a region; scan cost with no payoff.

Configuration starting points

Workloadenableddefragkhugepaged
Large buffer-pool DBmadvisemadviseAggressive
Fork-based webneverneverOff
ML trainingalways or deferdeferModerate
Redis / MemcachedmadvisemadviseModerate
Hard real-timeneverneverOff
code
# Inspect cat /sys/kernel/mm/transparent_hugepage/enabled cat /sys/kernel/mm/transparent_hugepage/defrag # Safer production defaults (example) echo madvise | sudo tee /sys/kernel/mm/transparent_hugepage/enabled echo madvise | sudo tee /sys/kernel/mm/transparent_hugepage/defrag

Apps opt in with madvise(addr, len, MADV_HUGEPAGE) and opt out with MADV_NOHUGEPAGE on small-object heaps.

If you found this explanation helpful, consider sharing it with others.

Mastodon