Collective Communication for 100k+ GPUs
Min Si, Pavan Balaji, +37
NCCLX is Meta's communication stack for Llama 4: host-driven, zero-copy collectives that start a 96K-GPU job 11x faster and cut decode time by up to 83%.
Explore machine learning papers and reviews related to LLM Inference. Find insights, analysis, and implementation details.
Min Si, Pavan Balaji, +37
NCCLX is Meta's communication stack for Llama 4: host-driven, zero-copy collectives that start a 96K-GPU job 11x faster and cut decode time by up to 83%.
Siyuan Shen, Anton Korzh, +11
How the Every µs Matters paper brings small GPU AllReduce operations close to the hardware speed-of-light bound: it measures a 1.404 µs floor on GB200, removes memory barriers with LL, sentinel and LL128 atomic signalling, and cuts vLLM inter-token latency by up to 13%.
Kan Zhu, Yufei Gao, +14
How NanoFlow raises LLM serving throughput by overlapping compute, memory and network work inside a single GPU: it shows serving is compute-bound, splits each batch into nano-batches, and shares SMs between concurrent kernels, reaching 1.91x the throughput of TensorRT-LLM.
Woosuk Kwon, Zhuohan Li, +7
How PagedAttention (the memory manager behind vLLM) applies OS-style virtual-memory paging to the KV cache — fixed-size blocks, a block table, and copy-on-write prefix sharing — to eliminate fragmentation and dramatically raise LLM serving throughput.