arXiv 2025
Collective Communication for 100k+ GPUs
Min Si, Pavan Balaji, +37
NCCLX is Meta's communication stack for Llama 4: host-driven, zero-copy collectives that start a 96K-GPU job 11x faster and cut decode time by up to 83%.
Explore machine learning papers and reviews related to Collective Communication. Find insights, analysis, and implementation details.
Min Si, Pavan Balaji, +37
NCCLX is Meta's communication stack for Llama 4: host-driven, zero-copy collectives that start a 96K-GPU job 11x faster and cut decode time by up to 83%.
Siyuan Shen, Anton Korzh, +11
How the Every µs Matters paper brings small GPU AllReduce operations close to the hardware speed-of-light bound: it measures a 1.404 µs floor on GB200, removes memory barriers with LL, sentinel and LL128 atomic signalling, and cuts vLLM inter-token latency by up to 13%.