HiveD: Sharing a GPU Cluster for Deep Learning with Guarantees
Hanyu Zhao, Zhenhua Han, +9
HiveD reserves groups of GPUs that sit together, not GPU counts, so a tenant never queues longer in a shared cluster than in a private one of the same size.
Explore machine learning papers and reviews related to Deep Learning. Find insights, analysis, and implementation details.
Hanyu Zhao, Zhenhua Han, +9
HiveD reserves groups of GPUs that sit together, not GPU counts, so a tenant never queues longer in a shared cluster than in a private one of the same size.
Mohammad Shoeybi, Mostofa Patwary, +4
Megatron-LM splits each transformer layer across GPUs with two all-reduces forward and two backward, and trains an 8.3B-parameter GPT-2 on 512 V100s.
Yunge Li, Lanyu Xu
Reordering image tokens along a Hilbert curve makes 2D local attention block-sparse, for about 4x faster window attention and 18x faster slide attention.
Leyla Naz Candogan, Arshia Afzal, +2
VIOLIN multiplies ViT attention by decay masks from eight space-filling curves: 1,296 extra parameters on DeiT-B and up to 8.7 points on VTAB-1K spatial tasks.
Kan Zhu, Yufei Gao, +14
How NanoFlow raises LLM serving throughput by overlapping compute, memory and network work inside a single GPU: it shows serving is compute-bound, splits each batch into nano-batches, and shares SMs between concurrent kernels, reaching 1.91x the throughput of TensorRT-LLM.
Tri Dao, Daniel Y. Fu, +3
How FlashAttention computes exact attention with linear memory by tiling Q, K, and V into SRAM-resident blocks and fusing the softmax, avoiding the quadratic HBM cost of materializing the full attention matrix.
Edward J. Hu, Yelong Shen, +6
How LoRA adapts a frozen large language model by learning a low-rank update ΔW = (α/r)·BA to each weight matrix, training under 1% of the parameters of full fine-tuning and adding zero inference latency once merged.
Woosuk Kwon, Zhuohan Li, +7
How PagedAttention (the memory manager behind vLLM) applies OS-style virtual-memory paging to the KV cache — fixed-size blocks, a block table, and copy-on-write prefix sharing — to eliminate fragmentation and dramatically raise LLM serving throughput.
William Fedus, Barret Zoph, +1
How the Switch Transformer scales to trillions of parameters by routing each token to a single expert (top-1 gating), decoupling model capacity from per-token compute while managing expert load with a capacity factor.
Alex Krizhevsky, Ilya Sutskever, +1
How AlexNet won ImageNet 2012 by a landslide and launched the deep-learning era: a deep convolutional network trained on two GPUs with ReLU activations, dropout regularization, and aggressive data augmentation.
Chao Jia, Yinfei Yang, +8
How ALIGN scales vision-language pre-training to 1.8 billion noisy image-alt-text pairs, showing that the sheer scale of a raw, uncurated dataset can outweigh careful data cleaning for a simple dual-encoder contrastive model.
Jacob Devlin, Ming-Wei Chang, +2
How BERT pre-trains a deep bidirectional Transformer with masked language modeling and next-sentence prediction, then fine-tunes the same model to state-of-the-art results across a wide range of NLP tasks.
Junnan Li, Dongxu Li, +2
How BLIP unifies vision-language understanding and generation in one model, and bootstraps noisy web data with CapFilt — a captioner that synthesizes captions and a filter that removes noisy image-text pairs.
Jiahui Yu, Zirui Wang, +4
How CoCa trains a single image-text foundation model with both a contrastive loss and a generative captioning loss by splitting the text decoder into unimodal and multimodal halves, computed in one forward pass.
Jean-Baptiste Alayrac, Jeff Donahue, +25
How Flamingo bridges a frozen vision encoder and a frozen language model with a Perceiver Resampler and gated cross-attention layers, enabling few-shot learning from interleaved image-text sequences.
Bin Xiao, Haiping Wu, +7
How Florence-2 turns detection, segmentation, grounding, captioning, and OCR into one sequence-to-sequence problem driven by task prompts, expressing spatial outputs as location tokens in the text vocabulary.
Ian J. Goodfellow, Jean Pouget-Abadie, +6
How GANs frame generative modeling as a two-player minimax game: a generator turns noise into samples while a discriminator learns to tell real from fake, driving the generator toward the true data distribution.
Peng Wang, Shuai Bai, +17
How Qwen2-VL perceives images and video at any resolution with naive dynamic resolution (variable visual tokens) and M-RoPE, a multimodal rotary position embedding that decomposes position into temporal, height, and width components.
Nikhila Ravi, Valentin Gabeur, +16
How SAM 2 extends promptable segmentation from images to video with a streaming memory bank and memory attention, propagating masks across frames and recovering objects after occlusion.
Xiaohua Zhai, Basil Mustafa, +2
How SigLIP replaces CLIP’s softmax contrastive loss with a pairwise sigmoid loss, removing the need for a global all-gather normalization and enabling strong language-image pre-training at small batch sizes.
Diederik P. Kingma, Max Welling
How the variational autoencoder turns an autoencoder into a generative model: an encoder predicts a Gaussian latent distribution, the reparameterization trick makes sampling differentiable, and the ELBO balances reconstruction against a KL prior.
Jonathan Ho, Ajay Jain, +1
How diffusion models learn to generate images by reversing a gradual noising process — the foundation of Stable Diffusion, DALL-E, and modern image generation.
Haotian Liu, Chunyuan Li, +2
LLaVA paper: align LLMs with visual information through instruction tuning on image-text pairs, enabling multimodal understanding and reasoning.
Yanghao Li, Hanzi Mao, +2
Investigating the effectiveness of plain Vision Transformers as backbones for object detection and proposing modifications to improve their performance.
Joseph Redmon, Santosh Divvala, +2
Introducing YOLO, a unified, real-time object detection system that frames object detection as a single regression problem.
Mingxing Tan, Quoc V. Le
EfficientNet achieves state-of-the-art image classification accuracy with improved efficiency through a novel compound scaling method for CNNs.
Shaoqing Ren, Kaiming He, +2
Faster R-CNN explained: how Region Proposal Networks (RPN) enable near real-time object detection with shared convolutional features.
Alexander Kirillov, Eric Mintun, +3
SAM is a promptable segmentation model that can segment any object in an image using points, boxes, or text prompts with zero-shot generalization.
Junnan Li, Dongxu Li, +2
BLIP-2 leverages frozen image encoders and LLMs for efficient vision-language pre-training, achieving state-of-the-art multimodal performance.
Nicolas Carion, Francisco Massa, +4
Introducing DETR, a novel end-to-end object detection framework that leverages Transformers to directly predict a set of object bounding boxes.
Alexey Dosovitskiy, Lucas Beyer, +10
Vision Transformer (ViT) explained: how splitting images into 16x16 patches enables pure transformer architecture for state-of-the-art image recognition.
Ze Liu, Yutong Lin, +6
Swin Transformer: hierarchical Vision Transformer using shifted windows for efficient image classification, object detection, and segmentation.
Alec Radford, Jong Wook Kim, +10
CLIP explained: contrastive learning on 400M image-text pairs enables zero-shot image classification and powerful vision-language understanding.
Horace He
Deep learning performance optimization from first principles. Learn to identify compute-bound, memory-bound, and overhead bottlenecks with fusion techniques.
Ashish Vaswani, Noam Shazeer, +6
Deep dive into the Transformer architecture that revolutionized NLP. Understand self-attention, multi-head attention, and positional encoding.
Andrei Ivanov, Nikoli Dryden, +3
Analysis of transformer performance bottlenecks caused by data movement. Learn optimization strategies for memory-bound operations on GPUs.
Kaiming He, Xiangyu Zhang, +2
ResNet analysis: how skip connections and residual learning solved the degradation problem, enabling training of 100+ layer neural networks.