Flash Attention: IO-Aware Exact Attention
Interactive Flash Attention visualization - the IO-aware algorithm achieving memory-efficient exact attention through tiling and kernel fusion.
Related
8 min read · advanced
Clear explanations of core machine learning concepts, from foundational ideas to advanced techniques. Understand attention mechanisms, transformers, skip connections, and more.
Interactive Flash Attention visualization - the IO-aware algorithm achieving memory-efficient exact attention through tiling and kernel fusion.
8 min read · advanced
Deep dive into how different prompt components influence model behavior across transformer layers, from surface patterns to abstract reasoning.
6 min read
NVIDIA vs AMD for deep learning compared at both layers: the CUDA vs ROCm software moat, the microarchitecture (warp vs wavefront, SM vs CU, Tensor vs Matrix Cores), and the datacenter accelerators (H100/H200/B200 vs MI300X/MI325X).
10 min read · intermediate
How C++ object files are linked into executables. Learn symbol resolution, static vs dynamic linking, and linker optimization.
5 min read
XFS internals end-to-end: allocation groups for lock-free parallel metadata, B+ trees instead of bitmaps, extent-based allocation that scales to terabytes, and delayed allocation that turns scattered writes into contiguous extents.
10 min read
Interactive KV cache visualization - how key-value caching in LLM transformers enables fast text generation without quadratic recomputation.
6 min read · intermediate