Skip to main content

Deep Learning

Explore machine learning papers and reviews related to Deep Learning. Find insights, analysis, and implementation details.

  • Tagged with
  • 37 papers
Back to all papers

Papers Related to Deep Learning

arXiv 2019

Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Mohammad Shoeybi, Mostofa Patwary, +4

Megatron-LM splits each transformer layer across GPUs with two all-reduces forward and two backward, and trains an 8.3B-parameter GPT-2 on 512 V100s.

OSDI 2025

NanoFlow: Towards Optimal Large Language Model Serving Throughput

Kan Zhu, Yufei Gao, +14

How NanoFlow raises LLM serving throughput by overlapping compute, memory and network work inside a single GPU: it shows serving is compute-bound, splits each batch into nano-batches, and shares SMs between concurrent kernels, reaching 1.91x the throughput of TensorRT-LLM.

SOSP 2023

Efficient Memory Management for Large Language Model Serving with PagedAttention

Woosuk Kwon, Zhuohan Li, +7

How PagedAttention (the memory manager behind vLLM) applies OS-style virtual-memory paging to the KV cache — fixed-size blocks, a block table, and copy-on-write prefix sharing — to eliminate fragmentation and dramatically raise LLM serving throughput.

JMLR 2022

Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity

William Fedus, Barret Zoph, +1

How the Switch Transformer scales to trillions of parameters by routing each token to a single expert (top-1 gating), decoupling model capacity from per-token compute while managing expert load with a capacity factor.

arXiv 2024

Qwen2-VL: Vision-Language Perception at Any Resolution

Peng Wang, Shuai Bai, +17

How Qwen2-VL perceives images and video at any resolution with naive dynamic resolution (variable visual tokens) and M-RoPE, a multimodal rotary position embedding that decomposes position into temporal, height, and width components.

Mastodon