Skip to main content

KV Cache

Explore machine learning papers and reviews related to KV Cache. Find insights, analysis, and implementation details.

  • Tagged with
  • 1 paper
Back to all papers

Papers Related to KV Cache

SOSP 2023

Efficient Memory Management for Large Language Model Serving with PagedAttention

Woosuk Kwon, Zhuohan Li, +7

How PagedAttention (the memory manager behind vLLM) applies OS-style virtual-memory paging to the KV cache — fixed-size blocks, a block table, and copy-on-write prefix sharing — to eliminate fragmentation and dramatically raise LLM serving throughput.

Mastodon