NeurIPS 2022
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Tri Dao, Daniel Y. Fu, +3
How FlashAttention computes exact attention with linear memory by tiling Q, K, and V into SRAM-resident blocks and fusing the softmax, avoiding the quadratic HBM cost of materializing the full attention matrix.
- Attention Mechanism
- Transformers
- GPU Optimization
- +2 more tags
