Mixture of Experts (MoE)
Understanding sparse mixture of experts models - architecture, routing mechanisms, load balancing, and efficient scaling strategies for large language models
Related
12 min read · advanced
Clear explanations of core machine learning concepts, from foundational ideas to advanced techniques. Understand attention mechanisms, transformers, skip connections, and more.
Understanding sparse mixture of experts models - architecture, routing mechanisms, load balancing, and efficient scaling strategies for large language models
12 min read · advanced
Understand internal covariate shift: why layer input distributions change during training, how it slows convergence, and how batch norm fixes it.
8 min read
A deep dive into NCCL internals: communicators and channels, how it picks ring/tree/NVLS algorithms and LL/LL128/Simple protocols, reading NCCL_DEBUG logs, and tuning and debugging distributed training.
12 min read · intermediate
C++ preprocessor visualized: macros, header guards, conditional compilation, and #include directives explained interactively.
5 min read
Structure of Arrays vs Array of Structures as instruments: cache-line fill, SIMD gather vs contiguous load, GPU coalescing, and AoSoA hybrids — when layout is a 10× decision.
12 min read · intermediate
A small draft model proposes tokens, the target model checks them all in one pass, and a rejection rule keeps the output distribution exactly the target's.
10 min read · intermediate