Skip to main content

Pinned Memory and DMA Transfers in PyTorch

Summary
Why pin_memory=True matters: pageable paths pay two host copies and block the CPU; pinned memory enables one DMA hop and real overlap with GPU compute.

The step is short. The bus isn’t.

You profile a training step and the GPU is not the villain. Compute is fine. Between batches the device goes quiet while host memory shuffles into VRAM. The flag everyone recommends is pin_memory=True. It looks like free magic until you see what it actually changes: not a slightly faster memcpy — a different path and the chance to overlap.

Two numbers

1. Path · pageable pays 2 host copies
CUDA will not DMA from ordinary pageable RAM. The runtime copies into a temporary pinned staging buffer, then DMA moves that buffer to the GPU. That is two host-side copies before the batch is device-side. Pinned memory is DMA-legal: one copy, host → device.

2. Overlap · the win is density, not µs
With pinned memory and non_blocking=True, H2D can run while the GPU steps and workers prepare the next batch. Without it, the transfer path holds the CPU and the pipeline goes hollow.

Flip the instruments until both claims feel obvious.

Path: two copies vs one

Pageable is the default. Staging is invisible until you watch the path. Pinned skips the middle box.

Overlap: why the flag exists

Raw H2D bandwidth matters less than whether load, transfer, and compute can stack. Pageable transfers keep the CPU in the way. Pinned + async lets three lanes light up at once.

code
DataLoader( dataset, batch_size=64, num_workers=4, pin_memory=True, persistent_workers=True, prefetch_factor=2, ) images = batch["images"].to("cuda", non_blocking=True)

Dictionary: same bytes, different contract

Pageable pages can leave RAM. Pinned pages cannot. That is the whole OS-side trade.

When pinning hurts

Pinned memory cannot be swapped. Footprint roughly:

code
pinned ≈ num_workers × prefetch_factor × batch_bytes

On a machine with headroom, locking a few hundred MB is free performance. On a 16 GB box already full of model + optimizer + Python, those locked pages push other processes into swap — worse than the staging copy you were trying to avoid.

code
# if Swap used climbs during training, shrink the pin footprint watch -n 1 'free -h | grep Swap' # 1) lower prefetch_factor or num_workers # 2) smaller batch # 3) pin_memory=False last

On multi-socket hosts, pinned allocations follow the NUMA node of the allocating thread. If the GPU lives on the other socket, DMA crosses the interconnect. Bind with numactl when that shows up in profiles — DataLoader will not fix topology for you.

What to do

  1. Enable pin_memory=True for GPU training unless RAM is already critical.
  2. Always pair with .to(device, non_blocking=True) — pinning without async leaves half the win on the table.
  3. Keep workers + prefetch so the pin thread has batches ready; tune those before disabling pin.
  4. Watch swap if the box is small; locked pages are not free.
  5. Confirm in the profiler — look for cudaMemcpyAsync from host pinned, not a double host path.
code
train_loader = DataLoader( train_dataset, batch_size=64, num_workers=4, pin_memory=True, persistent_workers=True, prefetch_factor=2, ) for batch in train_loader: images = batch['images'].to('cuda', non_blocking=True) labels = batch['labels'].to('cuda', non_blocking=True) outputs = model(images) loss = criterion(outputs, labels) loss.backward() optimizer.step()

Further Reading

If you found this explanation helpful, consider sharing it with others.

Mastodon