SOSP 2023
Efficient Memory Management for Large Language Model Serving with PagedAttention
Woosuk Kwon, Zhuohan Li, +7
How PagedAttention (the memory manager behind vLLM) applies OS-style virtual-memory paging to the KV cache — fixed-size blocks, a block table, and copy-on-write prefix sharing — to eliminate fragmentation and dramatically raise LLM serving throughput.
- KV Cache
- Transformers
- LLM Inference
- +2 more tags
