HiveD: Sharing a GPU Cluster for Deep Learning with Guarantees
Hanyu Zhao, Zhenhua Han, +9
HiveD reserves groups of GPUs that sit together, not GPU counts, so a tenant never queues longer in a shared cluster than in a private one of the same size.
Explore machine learning papers and reviews related to GPU Optimization. Find insights, analysis, and implementation details.
Hanyu Zhao, Zhenhua Han, +9
HiveD reserves groups of GPUs that sit together, not GPU counts, so a tenant never queues longer in a shared cluster than in a private one of the same size.
Mohammad Shoeybi, Mostofa Patwary, +4
Megatron-LM splits each transformer layer across GPUs with two all-reduces forward and two backward, and trains an 8.3B-parameter GPT-2 on 512 V100s.
Min Si, Pavan Balaji, +37
NCCLX is Meta's communication stack for Llama 4: host-driven, zero-copy collectives that start a 96K-GPU job 11x faster and cut decode time by up to 83%.
Yunge Li, Lanyu Xu
Reordering image tokens along a Hilbert curve makes 2D local attention block-sparse, for about 4x faster window attention and 18x faster slide attention.
Siyuan Shen, Anton Korzh, +11
How the Every µs Matters paper brings small GPU AllReduce operations close to the hardware speed-of-light bound: it measures a 1.404 µs floor on GB200, removes memory barriers with LL, sentinel and LL128 atomic signalling, and cuts vLLM inter-token latency by up to 13%.
Kan Zhu, Yufei Gao, +14
How NanoFlow raises LLM serving throughput by overlapping compute, memory and network work inside a single GPU: it shows serving is compute-bound, splits each batch into nano-batches, and shares SMs between concurrent kernels, reaching 1.91x the throughput of TensorRT-LLM.
Tri Dao, Daniel Y. Fu, +3
How FlashAttention computes exact attention with linear memory by tiling Q, K, and V into SRAM-resident blocks and fusing the softmax, avoiding the quadratic HBM cost of materializing the full attention matrix.