Skip to main content

Collective Communication

Explore machine learning papers and reviews related to Collective Communication. Find insights, analysis, and implementation details.

  • Tagged with
  • 2 papers
Back to all papers

Papers Related to Collective Communication

arXiv 2026

Every µs Matters: Achieving Near Speed-of-Light Latency in GPU Collectives

Siyuan Shen, Anton Korzh, +11

How the Every µs Matters paper brings small GPU AllReduce operations close to the hardware speed-of-light bound: it measures a 1.404 µs floor on GB200, removes memory barriers with LL, sentinel and LL128 atomic signalling, and cuts vLLM inter-token latency by up to 13%.

Mastodon