Skip to main content

Memory Management

Explore machine learning papers and reviews related to Memory Management. Find insights, analysis, and implementation details.

  • Tagged with
  • 1 paper
Back to all papers

Papers Related to Memory Management

SOSP 2023

Efficient Memory Management for Large Language Model Serving with PagedAttention

Woosuk Kwon, Zhuohan Li, +7

How PagedAttention (the memory manager behind vLLM) applies OS-style virtual-memory paging to the KV cache — fixed-size blocks, a block table, and copy-on-write prefix sharing — to eliminate fragmentation and dramatically raise LLM serving throughput.

Mastodon