The step is short. The bus isn’t.
You profile a training step and the GPU is not the villain. Compute is fine. Between batches the device goes quiet while host memory shuffles into VRAM. The flag everyone recommends is pin_memory=True. It looks like free magic until you see what it actually changes: not a slightly faster memcpy — a different path and the chance to overlap.
Two numbers
1. Path · pageable pays 2 host copies
CUDA will not DMA from ordinary pageable RAM. The runtime copies into a temporary pinned staging buffer, then DMA moves that buffer to the GPU. That is two host-side copies before the batch is device-side. Pinned memory is DMA-legal: one copy, host → device.
2. Overlap · the win is density, not µs
With pinned memory and non_blocking=True, H2D can run while the GPU steps and workers prepare the next batch. Without it, the transfer path holds the CPU and the pipeline goes hollow.
Flip the instruments until both claims feel obvious.
Path: two copies vs one
Pageable is the default. Staging is invisible until you watch the path. Pinned skips the middle box.
Overlap: why the flag exists
Raw H2D bandwidth matters less than whether load, transfer, and compute can stack. Pageable transfers keep the CPU in the way. Pinned + async lets three lanes light up at once.
DataLoader( dataset, batch_size=64, num_workers=4, pin_memory=True, persistent_workers=True, prefetch_factor=2, ) images = batch["images"].to("cuda", non_blocking=True)
Dictionary: same bytes, different contract
Pageable pages can leave RAM. Pinned pages cannot. That is the whole OS-side trade.
When pinning hurts
Pinned memory cannot be swapped. Footprint roughly:
pinned ≈ num_workers × prefetch_factor × batch_bytes
On a machine with headroom, locking a few hundred MB is free performance. On a 16 GB box already full of model + optimizer + Python, those locked pages push other processes into swap — worse than the staging copy you were trying to avoid.
# if Swap used climbs during training, shrink the pin footprint watch -n 1 'free -h | grep Swap' # 1) lower prefetch_factor or num_workers # 2) smaller batch # 3) pin_memory=False last
On multi-socket hosts, pinned allocations follow the NUMA node of the allocating thread. If the GPU lives on the other socket, DMA crosses the interconnect. Bind with numactl when that shows up in profiles — DataLoader will not fix topology for you.
What to do
- Enable
pin_memory=Truefor GPU training unless RAM is already critical. - Always pair with
.to(device, non_blocking=True)— pinning without async leaves half the win on the table. - Keep workers + prefetch so the pin thread has batches ready; tune those before disabling pin.
- Watch swap if the box is small; locked pages are not free.
- Confirm in the profiler — look for
cudaMemcpyAsyncfrom host pinned, not a double host path.
train_loader = DataLoader( train_dataset, batch_size=64, num_workers=4, pin_memory=True, persistent_workers=True, prefetch_factor=2, ) for batch in train_loader: images = batch['images'].to('cuda', non_blocking=True) labels = batch['labels'].to('cuda', non_blocking=True) outputs = model(images) loss = criterion(outputs, labels) loss.backward() optimizer.step()
Further Reading
- PyTorch DataLoader —
pin_memoryand multi-process loading - CUDA C Programming Guide: Page-Locked Host Memory —
cudaHostAllocand DMA rules - PyTorch Performance Tuning Guide — data pipeline knobs
Related concepts
A fine kernel can still profile poorly: warp divergence taxes throughput ~1/N; alone on a long scoreboard wait, SM issue util can hit 0%. Interactive demos.
Master GPU memory hierarchy from registers to global memory, understand coalescing patterns, bank conflicts, and optimization strategies for maximum performance
Structure of Arrays vs Array of Structures as instruments: cache-line fill, SIMD gather vs contiguous load, GPU coalescing, and AoSoA hybrids — when layout is a 10× decision.
PyTorch DataLoader deep dive — Dataset, Sampler, Workers, Collate internals, num_workers throughput profiling, memory analysis, serialization costs, production patterns (LMDB, WebDataset), and bottleneck diagnosis.
Why DataLoader num_workers matters: processes hide load latency behind GPU work, how to find the sweet spot, and the memory/GIL pitfalls that come with the pool.
Deep dive into the CUDA context object: control vs data plane, inventory (memory, modules, streams, events, graphs), push/pop/setCurrent stacks, primary retain/release, flags and limits, isolation, cost, and traps.
