HiveD: Sharing a GPU Cluster for Deep Learning with Guarantees
Hanyu Zhao, Zhenhua Han, +9
HiveD reserves groups of GPUs that sit together, not GPU counts, so a tenant never queues longer in a shared cluster than in a private one of the same size.
Reviews of machine learning research papers, newest first.
Hanyu Zhao, Zhenhua Han, +9
HiveD reserves groups of GPUs that sit together, not GPU counts, so a tenant never queues longer in a shared cluster than in a private one of the same size.
Mohammad Shoeybi, Mostofa Patwary, +4
Megatron-LM splits each transformer layer across GPUs with two all-reduces forward and two backward, and trains an 8.3B-parameter GPT-2 on 512 V100s.
Min Si, Pavan Balaji, +37
NCCLX is Meta's communication stack for Llama 4: host-driven, zero-copy collectives that start a 96K-GPU job 11x faster and cut decode time by up to 83%.
Yunge Li, Lanyu Xu
Reordering image tokens along a Hilbert curve makes 2D local attention block-sparse, for about 4x faster window attention and 18x faster slide attention.
Leyla Naz Candogan, Arshia Afzal, +2
VIOLIN multiplies ViT attention by decay masks from eight space-filling curves: 1,296 extra parameters on DeiT-B and up to 8.7 points on VTAB-1K spatial tasks.
Siyuan Shen, Anton Korzh, +11
How the Every µs Matters paper brings small GPU AllReduce operations close to the hardware speed-of-light bound: it measures a 1.404 µs floor on GB200, removes memory barriers with LL, sentinel and LL128 atomic signalling, and cuts vLLM inter-token latency by up to 13%.
Kan Zhu, Yufei Gao, +14
How NanoFlow raises LLM serving throughput by overlapping compute, memory and network work inside a single GPU: it shows serving is compute-bound, splits each batch into nano-batches, and shares SMs between concurrent kernels, reaching 1.91x the throughput of TensorRT-LLM.
Tri Dao, Daniel Y. Fu, +3
How FlashAttention computes exact attention with linear memory by tiling Q, K, and V into SRAM-resident blocks and fusing the softmax, avoiding the quadratic HBM cost of materializing the full attention matrix.
Edward J. Hu, Yelong Shen, +6
How LoRA adapts a frozen large language model by learning a low-rank update ΔW = (α/r)·BA to each weight matrix, training under 1% of the parameters of full fine-tuning and adding zero inference latency once merged.