Skip to main content

GPU Optimization

Explore machine learning papers and reviews related to GPU Optimization. Find insights, analysis, and implementation details.

  • Tagged with
  • 7 papers
Back to all papers

Papers Related to GPU Optimization

arXiv 2019

Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Mohammad Shoeybi, Mostofa Patwary, +4

Megatron-LM splits each transformer layer across GPUs with two all-reduces forward and two backward, and trains an 8.3B-parameter GPT-2 on 512 V100s.

arXiv 2026

Every µs Matters: Achieving Near Speed-of-Light Latency in GPU Collectives

Siyuan Shen, Anton Korzh, +11

How the Every µs Matters paper brings small GPU AllReduce operations close to the hardware speed-of-light bound: it measures a 1.404 µs floor on GB200, removes memory barriers with LL, sentinel and LL128 atomic signalling, and cuts vLLM inter-token latency by up to 13%.

OSDI 2025

NanoFlow: Towards Optimal Large Language Model Serving Throughput

Kan Zhu, Yufei Gao, +14

How NanoFlow raises LLM serving throughput by overlapping compute, memory and network work inside a single GPU: it shows serving is compute-bound, splits each batch into nano-batches, and shares SMs between concurrent kernels, reaching 1.91x the throughput of TensorRT-LLM.

Mastodon