Skip to main content

Sitemap

A visual representation of the site structure to help you navigate through the content.

Site Structure

Main landing page with introduction and recent articles

About/about

Learn more about me, my background, and expertise

Speaking/speaking

My talks, presentations, and speaking engagements

GPU Programming 101 in Python/talks/gpu-programming-101-in-python

BangPypers December Meetup 2025

ArrPy: Array You Fast Enough?/talks/arrpy-array-you-fast-enough

PyCon India 2025

Build with AI - Bangpypers

Rolling with Python: A Deep Dive into Python Wheels/talks/rolling-with-python-intro-to-python-wheels

NammaMUG × BangPypers Meetup

PyCon India 2024, Bangalore / PyCon Japan 2024, Tokyo

Devfest 2023

Introduction to Machine Learning/talks/introduction-to-machine-learning

KJ Somaiya Techfest 2019

Articles/articles

Collection of articles I've written on various topics

Article content

Article content

CUDA Matrix Multiplication Optimization: From Naive to Near-cuBLAS/articles/cuda-matrix-multiplication-optimization

Article content

Numerical Sensitivity: Why FP16 Breaks NAdam/articles/numerical-sensitivity

Article content

SAM's Multi-Mask Ambiguity: A Visual Deep Dive/articles/sam-multi-mask-ambiguity

Article content

Article content

H.264 Implementation & Applications (Part 3 of 3)/articles/h264-implementation-applications

Article content

H.264 Transform & Quantization (Part 2 of 3)/articles/h264-transform-quantization

Article content

Article content

Fix PyTorch "view size is not compatible" Error/articles/view-size-not-compatible

Article content

Article content

Article content

Quantization Deep Dive: From FP32 to INT4/articles/quantization-deep-dive

Article content

Article content

Python's logging Module Is Already a Singleton/articles/python-production-logging

Article content

Article content

Article content

Article content

Article content

Article content

Article content

Article content

Papers/papers

Research papers and publications

Paper content

Hilbert-Guided Sparse Local Attention/papers/hilbert-local-attention

Paper content

Paper content

Paper content

Paper content

Paper content

Visual Instruction Tuning/papers/visual-instruction-tuning

Paper content

Plain ViT Backbones for Object Detection/papers/vit-object-detection

Paper content

Paper content

Paper content

Paper content

ViT: An Image is Worth 16x16 Words/papers/image-worth-16x16

Paper content

A Survey of Techniques for Optimizing Transformer Inference/papers/optimizing-transformer-inference

Paper content

Paper content

Paper content

Paper content

Attention Is All You Need/papers/attention-is-all-you-need

Paper content

Paper content

Deep Residual Learning for Image Recognition/papers/deep-residual-learning

Paper content

Concepts/concepts

Interactive explanations of machine learning concepts

Transformers & LLMs/concepts/category/transformers

Attention mechanisms, large language models, and multimodal architectures: the building blocks of modern AI.

CLS Token in Vision Transformers/concepts/transformers/cls-token

Trace how a learned CLS row joins image patches, gathers evidence through self-attention, and becomes the image-level classification readout.

Hierarchical Attention in Vision Transformers/concepts/transformers/hierarchical-attention

Trace how local windows, shifted cross-window exchange, and patch merging turn one high-resolution token grid into a multi-scale vision hierarchy.

Multi-Head Attention/concepts/transformers/multihead-attention

How multi-head attention runs scaled dot-product attention in parallel across several representation subspaces to build context-aware token embeddings.

Positional Embeddings in Vision Transformers/concepts/transformers/positional-embeddings-vit

Explore how positional embeddings enable Vision Transformers (ViT) to process sequential data by encoding relative positions.

Self-Attention in Vision Transformers/concepts/transformers/self-attention-vit

Follow one image patch through Q/K/V projection, scaled scores, row-wise softmax, value mixing, and the residual update inside a Vision Transformer.

ALiBi: Attention with Linear Biases/concepts/transformers/alibi

Learn ALiBi, the position encoding method that adds linear biases to attention scores for exceptional length extrapolation in transformers.

How Flash Attention, Multi-Head Attention (MHA), Grouped-Query Attention (GQA), and Multi-Query Attention (MQA) compare — algorithm vs architecture, KV-cache memory, quality trade-offs, and how to choose for production transformer inference.

Attention Sinks: Stable Streaming LLMs/concepts/transformers/attention-sinks

Learn about attention sinks, where LLMs concentrate attention on initial tokens, and how preserving them enables streaming inference.

Cross-Attention: Bridging Different Modalities/concepts/transformers/cross-attention

Understand cross-attention, the mechanism that enables transformers to align and fuse information from different sources, sequences, or modalities.

Grouped-Query Attention: Head Sharing and KV Cache/concepts/transformers/grouped-query-attention

Trace how grouped-query attention keeps independent query heads while sharing fewer key/value heads, projections, and compact KV-cache rows during LLM decoding.

Linear Attention Approximations/concepts/transformers/linear-attention-approximations

Explore linear complexity attention mechanisms including Performer, Linformer, and other efficient transformers that scale to very long sequences.

Masked and Causal Attention/concepts/transformers/masked-attention

Trace causal attention from shifted next-token labels through the lower-triangular mask, pre-softmax blocking, parallel training, and incremental decoding.

Multi-Query Attention (MQA)/concepts/transformers/multi-query-attention

Learn Multi-Query Attention (MQA), the optimization that shares keys and values across attention heads for massive memory savings.

Rotary Position Embeddings (RoPE)/concepts/transformers/rotary-position-embeddings

Learn Rotary Position Embeddings (RoPE), the elegant position encoding using rotation matrices, powering LLaMA, Mistral, and modern LLMs.

Scaled Dot-Product Attention/concepts/transformers/scaled-dot-product

Master scaled dot-product attention, the fundamental transformer building block. Learn why scaling is crucial for stable training.

Sliding Window Attention/concepts/transformers/sliding-window-attention

Sliding Window Attention for long sequences: local context windows enable O(n) complexity, used in Mistral and Longformer models.

Sparse Attention Patterns/concepts/transformers/sparse-attention-patterns

Explore sparse attention mechanisms that reduce quadratic complexity to linear or sub-quadratic, enabling efficient processing of long sequences.

The Vision-Language Alignment Problem/concepts/transformers/alignment-problem

How vision-language models align visual and text representations using contrastive learning, cross-modal attention, and CLIP-style training.

Context Windows: The Memory Limits of LLMs/concepts/transformers/context-windows

Interactive visualization of LLM context windows - sliding windows, expanding contexts, and attention patterns that define model memory limits.

Flash Attention: IO-Aware Exact Attention/concepts/transformers/flash-attention

Interactive Flash Attention visualization - the IO-aware algorithm achieving memory-efficient exact attention through tiling and kernel fusion.

KV Cache: The Secret to Fast LLM Inference/concepts/transformers/kv-cache

Interactive KV cache visualization - how key-value caching in LLM transformers enables fast text generation without quadratic recomputation.

The Modality Gap in Multimodal AI/concepts/transformers/modality-gap

The modality gap in CLIP and vision-language models: why image and text embeddings occupy separate regions despite contrastive training.

Multimodal Scaling Laws/concepts/transformers/scaling-laws

Discover how multimodal vision-language models like CLIP, ALIGN, and LLaVA scale with data, parameters, and compute following Chinchilla-style power laws.

Tokenization: Converting Text to Numbers/concepts/transformers/tokenization

Interactive exploration of tokenization methods in LLMs - BPE, SentencePiece, and WordPiece. Understand how text becomes tokens that models can process.

Vision-Language Adapters: Efficient Fine-tuning/concepts/transformers/vision-language-adapters

Master LoRA, bottleneck adapters, and prefix tuning for parameter-efficient fine-tuning of vision-language models like LLaVA with minimal compute and memory.

Mixture of Experts (MoE)/concepts/transformers/mixture-of-experts

Understanding sparse mixture of experts models - architecture, routing mechanisms, load balancing, and efficient scaling strategies for large language models

Speculative Decoding: Draft, Then Verify/concepts/transformers/speculative-decoding

A small draft model proposes tokens, the target model checks them all in one pass, and a rejection rule keeps the output distribution exactly the target's.

Vision-Language Alignment: CLIP, BLIP-2, LLaVA, Flamingo/concepts/transformers/vision-language-alignment

CLIP, BLIP-2, LLaVA and Flamingo align images with text by different routes. Compare their objectives, bridges, data, and which parts train or stay frozen.

Deep Learning/concepts/category/deep-learning

Neural network fundamentals: normalization, convolutions, losses, graph networks, and training dynamics.

Calinski-Harabasz Index: The Variance Ratio Criterion/concepts/deep-learning/calinski-harabasz

How the Calinski-Harabasz index evaluates clustering quality by measuring the ratio of between-cluster to within-cluster variance — fast, intuitive, and ideal for k-selection with convex clusters.

Silhouette Score: Per-Point Clustering Evaluation/concepts/deep-learning/silhouette-score

How the silhouette score measures clustering quality for every individual point — comparing intra-cluster cohesion to nearest-cluster separation, with per-point diagnostics that work for arbitrary cluster shapes.

Davies-Bouldin Index: Worst-Case Cluster Similarity/concepts/deep-learning/davies-bouldin

How the Davies-Bouldin index evaluates clustering quality by finding each cluster's most similar neighbor — a pessimistic, worst-case metric that catches overlapping cluster pairs.

Representation Collapse in Self-Supervised Learning/concepts/deep-learning/collapse-risk

Understanding complete, dimensional, and cluster collapse — the failure modes that every self-supervised method must prevent. Learn why collapse happens and how contrastive, asymmetric, regularization, and masking approaches solve it.

Batch Norm vs Layer Norm: When to Use Which/concepts/deep-learning/batch-vs-layer-norm

BatchNorm normalizes over the batch and spatial axes; LayerNorm normalizes over the channel and spatial axes for each sample. The choice changes whether your model trains stably with batch=1, depends on batch composition at inference, and behaves consistently across train and eval.

Convolution Operation: The Foundation of CNNs/concepts/deep-learning/convolution-operation

Interactive guide to convolution in CNNs: visualize sliding windows, kernels, stride, padding, and feature detection with step-by-step demos.

Cross-Entropy Loss for Classification/concepts/deep-learning/cross-entropy-loss

Understand cross-entropy loss for classification: interactive demos of binary and multi-class CE, the -log(p) curve, softmax gradients, and focal loss.

Dilated Convolutions: Expanding Receptive Fields Efficiently/concepts/deep-learning/dilated-convolutions

Understand dilated (atrous) convolutions: how dilation rates expand receptive fields exponentially without extra parameters and how to avoid gridding artifacts.

Receptive Field in CNNs/concepts/deep-learning/receptive-field

Understand receptive fields in CNNs: how convolutional layers expand their field of view and the gap between theoretical and effective receptive fields.

VAE Latent Space: Understanding Variational Autoencoders/concepts/deep-learning/vae-latent-space

Explore VAE latent space in deep learning. Learn variational autoencoder encoding, decoding, interpolation, and the reparameterization trick.

Contrastive Loss for Representation Learning/concepts/deep-learning/contrastive-loss

Understand contrastive loss for representation learning: interactive demos of InfoNCE, triplet loss, and embedding space clustering with temperature tuning.

Dropout Regularization/concepts/deep-learning/dropout

Understand dropout regularization: how randomly silencing neurons prevents overfitting, the inverted dropout trick, and when to use each dropout variant.

Focal Loss: Focusing on Hard Examples/concepts/deep-learning/focal-loss

Learn focal loss for deep learning: down-weight easy examples, focus on hard ones. Interactive demos of gamma, alpha balancing, and RetinaNet.

He/Kaiming Initialization/concepts/deep-learning/he-initialization

Learn He (Kaiming) initialization for ReLU networks: why ReLU needs special weight initialization, variance flow, and dead neurons explained.

KL Divergence in Machine Learning/concepts/deep-learning/kl-divergence

Learn KL divergence for machine learning: measure distribution differences in VAEs, knowledge distillation, and variational inference.

MSE and MAE Loss Functions/concepts/deep-learning/mse-mae

Interactive guide to MSE vs MAE for regression: explore outlier sensitivity, gradient behavior, and Huber loss with visualizations.

Xavier/Glorot Initialization/concepts/deep-learning/xavier-initialization

Learn Xavier (Glorot) initialization: how it balances forward signals and backward gradients to enable stable deep network training with tanh and sigmoid.

Adaptive Tiling: Efficient Visual Token Generation/concepts/deep-learning/adaptive-tiling

Learn adaptive tiling in vision transformers: dynamically partition images based on visual complexity to reduce token counts while preserving detail.

Emergent Abilities in Large Language Models/concepts/deep-learning/emergent-abilities

Explore emergent abilities in large language models: sudden capabilities at scale thresholds, phase transitions, and the mirage debate.

Prompt Engineering for LLMs/concepts/deep-learning/prompt-engineering

Master prompt engineering for large language models: from basic composition to Chain-of-Thought, few-shot, and advanced techniques.

Prompt Influence Flow Through Transformer Layers/concepts/deep-learning/prompt-influence-flow

Deep dive into how different prompt components influence model behavior across transformer layers, from surface patterns to abstract reasoning.

Neural Scaling Laws Explained/concepts/deep-learning/scaling-laws

Explore neural scaling laws in deep learning: power law relationships between model size, data, and compute that predict AI performance.

Visual Complexity Analysis: Smart Image Processing/concepts/deep-learning/visual-complexity-analysis

Learn visual complexity analysis in deep learning - how neural networks measure entropy, edges, and saliency for adaptive image processing.

Gradient Flow in Deep Networks/concepts/deep-learning/gradient-flow

Learn how gradients propagate through deep neural networks during backpropagation. Understand vanishing and exploding gradient problems.

NAdam: Nesterov-Accelerated Adam/concepts/deep-learning/nadam

Understand the NAdam optimizer that fuses Adam adaptive learning rates with Nesterov look-ahead momentum for faster, smoother convergence in deep learning.

Layer Normalization for Transformers/concepts/deep-learning/layer-normalization

Learn layer normalization for transformers and sequence models: how normalizing across features enables batch-independent training.

Internal Covariate Shift/concepts/deep-learning/internal-covariate-shift

Understand internal covariate shift: why layer input distributions change during training, how it slows convergence, and how batch norm fixes it.

Batch Normalization in Deep Learning/concepts/deep-learning/batch-normalization

Learn batch normalization in deep learning: how normalizing layer inputs accelerates training, improves gradient flow, and acts as regularization.

Skip Connections in Neural Networks/concepts/deep-learning/skip-connections

Learn how skip connections and residual learning enable training of very deep neural networks. Understand the ResNet revolution with interactive visualizations.

Graph Attention Networks (GAT)/concepts/deep-learning/graph-attention-networks

Adaptive attention-based aggregation for graph neural networks - multi-head attention, learned weights, and interpretable graph learning

Graph Centrality & Metrics/concepts/deep-learning/graph-centrality

Understanding node importance through centrality measures, shortest paths, hop distances, clustering coefficients, and fundamental graph metrics

Graph Convolutional Networks (GCN)/concepts/deep-learning/graph-convolutional-networks

Learn Graph Convolutional Networks (GCN) with spectral theory, message passing, and node classification for geometric deep learning.

Graph Embeddings and Node2Vec/concepts/deep-learning/graph-embeddings

Learning low-dimensional vector representations of graphs through random walks, DeepWalk, Node2Vec, and skip-gram models

Graph Pooling Methods/concepts/deep-learning/graph-pooling

Hierarchical graph coarsening techniques - TopK, SAGPool, DiffPool, and readout operations for graph-level representations

Computer Vision/concepts/category/computer-vision

Object detection, feature pyramids, and visual recognition techniques.

Feature Pyramid Networks (FPN) Explained/concepts/computer-vision/feature-pyramid-networks

Learn how Feature Pyramid Networks build multi-scale feature representations through top-down pathways and lateral connections for robust object detection.

ASFF: Adaptive Spatial Feature Fusion/concepts/computer-vision/asff

Learning where to fuse multi-scale features with per-pixel, per-level fusion weights. ASFF challenges FPN's uniform fusion assumption.

RoI Pooling, RoI Align & Deformable RoI Pooling/concepts/computer-vision/roi-pooling

Understanding region-based feature extraction for object detection, from quantized pooling to sub-pixel alignment and adaptive sampling

Anchor-Based vs Anchor-Free Object Detection/concepts/computer-vision/anchor-based-vs-anchor-free

Compare anchor-based vs anchor-free object detection: Faster R-CNN and RetinaNet anchors vs FCOS and CenterNet point-based methods.

Understanding how neural architecture search discovers optimal feature pyramid architectures that outperform hand-designed alternatives

DETR Explained: Object Detection with Transformers/concepts/computer-vision/modern-object-detection

Understanding end-to-end object detection with transformers, from DETR's object queries to bipartite matching and attention-based localization

NMS & Soft-NMS: Removing Duplicate Detections/concepts/computer-vision/nms-soft-nms

Understanding Non-Maximum Suppression algorithms for object detection post-processing, from greedy NMS to soft variants

Visual Complexity Analysis for Token Allocation/concepts/computer-vision/visual-complexity-analysis

Learn how visual complexity analysis optimizes vision transformer token allocation using edge detection, FFT, and entropy metrics.

Embeddings & Retrieval/concepts/category/embeddings

Dense and sparse embeddings, quantization, and vector search for semantic retrieval.

Dense Embeddings/concepts/embeddings/dense-embeddings

How dense embeddings turn meaning into geometry: word2vec, GloVe, and contextual models, vector arithmetic, cosine similarity, and where the field is heading.

How sparse retrieval (BM25/TF-IDF), dense retrieval (BERT-style embeddings), and hybrid systems that combine both compare on recall, semantic understanding, computational cost, and operational complexity for modern search.

BM25 Algorithm for Text Retrieval/concepts/embeddings/bm25-algorithm

Master the BM25 algorithm, the probabilistic ranking function powering Elasticsearch and Lucene for keyword-based document retrieval and search systems.

Pooling Strategies/concepts/embeddings/pooling-strategies

How a transformer’s per-token outputs become one embedding: CLS, mean, max, last-token, and attention pooling — what each does and when to use it.

Contrastive Learning/concepts/embeddings/contrastive-learning

Master contrastive learning for vector embeddings: how InfoNCE loss and self-supervised techniques train models to create high-quality semantic representations.

Matryoshka Embeddings/concepts/embeddings/matryoshka-embeddings

Matryoshka embeddings: nested representations enabling dimension reduction by simple truncation without model retraining for flexible retrieval.

Domain Adaptation for Embeddings/concepts/embeddings/domain-adaptation

Domain adaptation for embeddings: transfer learning to fine-tune retrieval models across domains while preventing catastrophic forgetting.

Cross-Lingual Alignment/concepts/embeddings/cross-lingual-alignment

Learn cross-lingual embedding alignment techniques like VecMap and MUSE for multilingual vector retrieval and zero-shot language transfer in search systems.

Cross-Encoder vs Bi-Encoder/concepts/embeddings/cross-encoder-vs-bi-encoder

Understand the fundamental differences between independent and joint encoding architectures for neural retrieval systems.

Multi-Vector Late Interaction/concepts/embeddings/multi-vector-late-interaction

Explore ColBERT and other multi-vector retrieval models that use fine-grained token-level matching for superior search quality.

Hybrid Retrieval Systems/concepts/embeddings/hybrid-retrieval-systems

Build hybrid retrieval systems combining BM25 sparse search with dense vector embeddings using reciprocal rank fusion for superior semantic search performance.

Quantization Effects Simulator/concepts/embeddings/quantization-effects

Embedding quantization simulator: explore memory-accuracy trade-offs from float32 to int8 and binary representations for retrieval.

Vector Quantization Techniques/concepts/embeddings/vector-quantization

Master vector compression techniques from scalar to product quantization. Learn how to reduce memory usage by 10-100× while preserving search quality.

Binary Embeddings for Fast Search/concepts/embeddings/binary-embeddings

Learn how binary embeddings use 1-bit quantization for ultra-compact vector representations, enabling billion-scale similarity search with 32x memory reduction.

Vector Index Structures/concepts/embeddings/index-structures

Explore the fundamental data structures powering vector databases: trees, graphs, hash tables, and hybrid approaches for efficient similarity search.

How HNSW, IVF-PQ, and LSH compare for approximate nearest neighbor (ANN) search — recall, latency, memory, build cost, and update characteristics — with Annoy, ScaNN, and DiskANN included for completeness.

LSH: Locality Sensitive Hashing/concepts/embeddings/lsh-search

Explore how LSH uses probabilistic hash functions to find similar vectors in sub-linear time, perfect for streaming and high-dimensional data.

Learn how IVF-PQ combines clustering and compression to enable billion-scale vector search with minimal memory footprint.

HNSW: Hierarchical Navigable Small World/concepts/embeddings/hnsw-search

How HNSW navigates a layered proximity graph to find nearest neighbors in logarithmic time — the default in-memory index of modern vector databases.

GPU & High-Performance Computing/concepts/category/gpu-computing

CUDA, tensor cores, multi-GPU communication, and cluster-scale workload orchestration.

Slurm Fundamentals: Job Scheduling on HPC Clusters/concepts/gpu-computing/slurm-fundamentals

Complete guide to Slurm — architecture, core commands, job lifecycle, job scripts, array jobs, dependencies, monitoring with squeue/sacct, and troubleshooting failed jobs on HPC clusters.

Tensor Cores: Mixed Precision & Matrix Acceleration/concepts/gpu-computing/tensor-cores

How NVIDIA Tensor Cores accelerate GEMM: warp-level MMA, precision dials (TF32/FP16/BF16/INT8/FP8), shape alignment cliffs, AMP vs model.half(), and when the speedup actually shows up.

High Bandwidth Memory (HBM)/concepts/gpu-computing/hbm-memory

How HBM feeds AI GPUs: the memory wall, width-over-speed packaging, TSVs and interposers, the on-package hierarchy, and the roofline that decides if your kernel is bandwidth-bound.

GPU Memory Hierarchy & Optimization/concepts/gpu-computing/memory-hierarchy

Master GPU memory hierarchy from registers to global memory, understand coalescing patterns, bank conflicts, and optimization strategies for maximum performance

Multi-GPU Communication: NVLink, PCIe, and NCCL/concepts/gpu-computing/multi-gpu-communication

How GPUs talk: the bandwidth cliff from HBM to Ethernet, NVLink 5 and GB200 NVL72 topologies, ring AllReduce step by step, and choosing between NCCL, Gloo, and MPI.

GPU Streaming Multiprocessor (SM)/concepts/gpu-computing/shared-multiprocessor

How a single SM actually runs work: warps of 32, cutaway of cores and memory, divergence tax, latency-hiding occupancy, and coalesced loads — instruments, not a catalog.

Slurm GPU Allocation for Distributed Training/concepts/gpu-computing/slurm-gpu-allocation

Complete guide to GPU allocation on Slurm — --gres flags, CUDA_VISIBLE_DEVICES remapping, GPU topology and NVLink binding, MIG partitioning, production job scripts, and debugging common GPU errors.

CUDA Contexts: Ownership, Current Stack, Isolation/concepts/gpu-computing/cuda-context

Deep dive into the CUDA context object: control vs data plane, inventory (memory, modules, streams, events, graphs), push/pop/setCurrent stacks, primary retain/release, flags and limits, isolation, cost, and traps.

CUDA Streams: Asynchronous Execution and Concurrency/concepts/gpu-computing/cuda-streams

A CUDA stream is an in-order queue of GPU ops. Overlap H→D, kernels, and D→H across streams — and avoid the default-stream trap that serializes everything.

Slurm Resource Management and Job Priority/concepts/gpu-computing/slurm-resource-management

How Slurm decides which jobs run first — priority factors, fair-share scheduling, backfill, and monitoring commands (squeue, sinfo, sacct).

NVIDIA Unified Virtual Memory/concepts/gpu-computing/unified-memory

NVIDIA Unified Virtual Memory (UVM): on-demand page migration, memory oversubscription, and simplified CPU-GPU memory management.

CUDA Context vs Streams vs MPS: Which Layer Fixes What/concepts/gpu-computing/cuda-context-vs-streams

Decision map for CUDA: a context is per-process GPU state, a stream is an in-order queue inside a context, and MPS shares one context across processes. Pick the layer that matches the problem.

Flynn's Classification: Taxonomy of Computer Architectures/concepts/gpu-computing/flynns-classification

Flynn's Classification explained — SISD, SIMD, MISD, MIMD with interactive architecture explorer, SIMD evolution from MMX to AMX, branch divergence visualization, and workload-architecture throughput comparison.

MPI Fundamentals: Message Passing for Distributed Computing/concepts/gpu-computing/mpi-fundamentals

Complete MPI guide — point-to-point and collective communication with real C and mpi4py code, deadlock simulation, performance benchmarking, communicator splitting, and debugging on HPC clusters.

OpenMP: Shared-Memory Parallel Programming/concepts/gpu-computing/openmp

OpenMP parallel programming: fork-join model, scheduling, data races, false sharing, NUMA thread affinity, and GPU offloading.

Page Migration & Fault Handling/concepts/gpu-computing/page-migration

CUDA page migration and fault handling between CPU and GPU memory. Learn TLB management, DMA transfers, and memory optimization.

Slurm Accounting and Resource Tracking/concepts/gpu-computing/slurm-accounting

How Slurm tracks resource consumption through account hierarchies, TRES billing, and resource limits — sacctmgr, sreport, and the association model explained.

Why exclusive CUDA contexts leave SMs idle under multi-process load, how MPS multiplexes clients through a shared context, thread percentage caps, and when to pick exclusive, MPS, or MIG.

HPC Performance Optimization: Scaling, Profiling, and Tuning/concepts/gpu-computing/hpc-performance-optimization

Amdahl and Gustafson ceilings, strong vs weak scaling, roofline bounds, and hiding all-reduce behind compute — the levers that decide whether more GPUs actually buy science.

Slurm Backfill Scheduling: How Small Jobs Fill the Gaps/concepts/gpu-computing/slurm-backfill

How sched/backfill works — the algorithm that lets small jobs run in gaps while large jobs wait, why accurate time limits matter, and the key tuning parameters (bf_interval, bf_window, bf_max_job_test).

NVIDIA vs AMD for deep learning compared at both layers: the CUDA vs ROCm software moat, the microarchitecture (warp vs wavefront, SM vs CU, Tensor vs Matrix Cores), and the datacenter accelerators (H100/H200/B200 vs MI300X/MI325X).

NVIDIA Device Files in /dev//concepts/gpu-computing/nvidia-device-files

How /dev/nvidia*, nvidiactl, nvidia-uvm, and DRI nodes map major/minor numbers to the driver, what CUDA opens first, and the minimum mount set for containers.

NVIDIA GPU Operator: Declarative GPU Stacks on Kubernetes/concepts/gpu-computing/kubernetes-operator

How the NVIDIA GPU Operator turns drivers, container toolkit, device plugin, GFD, DCGM, and MIG into DaemonSets reconciling from a ClusterPolicy — boot order, request path, sharing modes, and ops traps.

NVIDIA Persistence Daemon: Killing GPU Cold Starts/concepts/gpu-computing/nvidia-persistence-daemon

Why the first nvidia-smi or CUDA open costs seconds, what that cold path rebuilds, how nvidia-persistenced holds device FDs on the host, and when persistence pays off for pods, CI, and batch jobs.

Distributed Parallelism in Deep Learning/concepts/gpu-computing/distributed-parallelism

GPU distributed parallelism: Data Parallel (DDP), Tensor Parallel, Pipeline Parallel, and ZeRO optimization for training large AI models.

GPU Pipeline Hazards: Scoreboard Stalls and Warp Divergence/concepts/gpu-computing/gpu-pipeline-hazards

A fine kernel can still profile poorly: warp divergence taxes throughput ~1/N; alone on a long scoreboard wait, SM issue util can hit 0%. Interactive demos.

NCCL: How NVIDIA Collective Communication Works/concepts/gpu-computing/nccl-communication

A deep dive into NCCL internals: communicators and channels, how it picks ring/tree/NVLS algorithms and LL/LL128/Simple protocols, reading NCCL_DEBUG logs, and tuning and debugging distributed training.

Data vs Tensor vs Pipeline Parallelism: When to Use Each/concepts/gpu-computing/data-vs-tensor-vs-pipeline-parallelism

Data, tensor and pipeline parallelism split the batch, weight matrices or layers across GPUs. Compare memory, traffic and batch limits, then combine them.

Systems & Architecture/concepts/category/systems

CPU pipelines, memory hierarchies, the Linux kernel, filesystems, and networking.

Deep dive into CPU cache lines — interactive cache simulator with configurable associativity and replacement policies, false sharing MESI protocol visualization, access pattern benchmarks, and optimization techniques.

Explore the inner workings of RAM through beautiful animations and interactive visualizations. Understand memory cells, addressing, and the memory hierarchy.

initramfs: The Initial RAM Filesystem Explained/concepts/systems/initramfs-boot-process

Learn how initramfs enables Linux boot by loading essential drivers before the root filesystem mounts. Explore early userspace initialization.

Linux Kernel Architecture: How Your OS Actually Works/concepts/systems/kernel-architecture

Linux kernel architecture explained. Learn syscalls, protection rings, user vs kernel space, and what happens when you run a command.

Filesystems: The Digital DNA of Data Storage/concepts/systems/filesystems-overview

Explore Linux filesystems through interactive visuals. Learn VFS, compare ext4 vs Btrfs vs ZFS, and understand file operations.

Master virtual memory and TLB address translation with interactive demos. Learn page tables, page faults, and memory management optimization.

Filesystem Journaling: Write-Ahead Logging/concepts/systems/filesystem-journaling

Learn how filesystem journaling prevents data loss during crashes. Explore write-ahead logging and recovery in ext4 and XFS.

Understand Linux inodes - the metadata structures behind every file. Learn about hard links, soft links, and inode limits.

Memory Access Patterns: Sequential vs Strided/concepts/systems/memory-access-patterns

Master sequential vs strided memory access patterns. Learn how cache efficiency and hardware prefetching affect application performance.

Understand Copy-on-Write (CoW) in Btrfs and ZFS. Learn how CoW enables instant snapshots, atomic writes, and data integrity.

FUSE: Filesystem in Userspace Explained/concepts/systems/fuse-filesystem

Learn FUSE (Filesystem in Userspace) for building custom filesystems. Understand how NTFS-3G, SSHFS, and cloud storage work.

ext4: The Linux Workhorse Filesystem/concepts/systems/ext4-filesystem

Explore ext4, the default Linux filesystem with journaling, extents, and proven reliability. Learn how ext4 protects your data.

CPU Pipeline Architecture/concepts/systems/cpu-pipeline-detailed

Deep dive into CPU pipeline architecture covering 5-stage RISC pipelines, data hazards, control hazards, superscalar execution, and out-of-order processing.

Explore CPU pipeline stages, instruction-level parallelism, pipeline hazards, and branch prediction through interactive visualizations.

Master pipeline hazards through interactive visualizations of data dependencies, control hazards, structural conflicts, and advanced detection mechanisms.

Master Linux mount options like noatime and async for performance tuning and security hardening. Interactive guide to fstab configuration.

NTFS Filesystem: The Master File Table/concepts/systems/ntfs-filesystem

NTFS internals from the Master File Table outward: 1 KB attribute records, resident vs non-resident $DATA, run lists, alternate data streams, the $LogFile journal, and why dual-boot Linux distros prefer ntfs3 over ntfs-3g.

Btrfs: Modern Copy-on-Write Filesystem/concepts/systems/btrfs-filesystem

Learn the Btrfs filesystem with built-in snapshots, RAID, and compression. Explore copy-on-write, subvolumes, and self-healing on Linux.

Filesystem Data Integrity: Detecting Silent Corruption/concepts/systems/filesystem-integrity

Understand how modern filesystems use Merkle-tree checksums and mirrored pools to detect, repair, and proactively scrub silent data corruption that ext4 and XFS miss entirely.

ZFS: The Ultimate Filesystem/concepts/systems/zfs-filesystem

Master ZFS filesystem with pooled storage, RAID-Z, snapshots, and checksums. Learn enterprise-grade data integrity on Linux.

XFS: High-Performance Parallel Filesystem/concepts/systems/xfs-filesystem

XFS internals end-to-end: allocation groups for lock-free parallel metadata, B+ trees instead of bitmaps, extent-based allocation that scales to terabytes, and delayed allocation that turns scattered writes into contiguous extents.

FAT32 & exFAT: Universal Filesystems/concepts/systems/fat-filesystems

Learn FAT32 and exFAT filesystems for cross-platform USB drives and SD cards. Understand file size limits and compatibility.

Memory Controllers: The Brain Behind RAM Management/concepts/systems/memory-controllers

Learn how memory controllers manage CPU-RAM data flow. Interactive demos of channels, ranks, banks, and command scheduling for optimal bandwidth.

Memory Interleaving: Parallel Memory Access/concepts/systems/memory-interleaving

Discover how memory interleaving distributes addresses across banks for parallel access. Boost memory bandwidth in DDR5 and GPU systems.

NUMA Architecture: Non-Uniform Memory Access/concepts/systems/numa-architecture

Explore NUMA architecture and memory locality in multi-socket systems. Understand local vs remote memory access latency and optimization strategies.

RAID: Redundant Arrays for Speed and Safety/concepts/systems/raid-storage

RAID storage visualized: RAID 0, 1, 5, 6, and 10 levels explained. Learn how they work, when to use them, and disk failure recovery.

SoA vs AoS: Data Layout Optimization/concepts/systems/soa-vs-aos

Structure of Arrays vs Array of Structures as instruments: cache-line fill, SIMD gather vs contiguous load, GPU coalescing, and AoSoA hybrids — when layout is a 10× decision.

Linux Process Management: Fork, Exec, and Beyond/concepts/systems/process-management

Master Linux process management through interactive visualizations. Understand process lifecycle, fork/exec operations, zombies, orphans, and CPU scheduling.

Explore Linux memory management through interactive visualizations. Understand virtual memory, page tables, TLB, swapping, and memory allocation.

Transparent Huge Pages (THP): Reducing TLB Pressure/concepts/systems/transparent-huge-pages

Learn how Transparent Huge Pages (THP) reduces TLB misses by promoting 4KB to 2MB pages. Understand performance benefits and memory bloat tradeoffs.

Linux System Calls: The User-Kernel Interface/concepts/systems/system-calls

Linux system calls visualized: how user programs communicate with the kernel, protection rings, context switching, and syscall performance.

Client-Server Communication: Polling vs WebSockets/concepts/systems/client-server-communication

Learn client-server communication patterns including short polling, long polling, and WebSockets. Compare HTTP protocols for real-time web applications.

Long Polling: The Patient Connection/concepts/systems/long-polling

Learn HTTP long polling - a server-side technique that holds connections open until data arrives. Achieve near real-time updates with standard protocols.

Master the Linux networking stack through interactive visualizations. Understand TCP/IP layers, sockets, iptables, routing, and network namespaces.

Short Polling: The Impatient Client Pattern/concepts/systems/short-polling

Learn short polling in networking - a simple HTTP pattern for periodic data fetching. See why 70-90% of requests waste bandwidth and when to use alternatives.

Understanding TCP/IP Protocol Stack/concepts/systems/tcp-ip

Explore the TCP/IP protocol stack, packet encapsulation, and how data travels through network layers from application to physical transmission.

Master WebSocket protocol for real-time bidirectional communication over TCP. Learn handshakes, frames, and building low-latency web applications.

Linux Boot Process: From Power-On to Login/concepts/systems/boot-process

Visualize the complete Linux boot sequence from BIOS/UEFI to login. Learn how GRUB, kernel, and systemd work together with interactive visualizations.

Linux Init Systems: From SysV to systemd/concepts/systems/init-systems

Compare Linux init systems through interactive visualizations. Understand the evolution from SysV Init to systemd, service management, and boot orchestration.

Master Linux kernel modules through interactive visualizations. Learn how to load, unload, develop, and debug kernel modules that extend Linux functionality.

Master Linux namespaces — the kernel mechanism that makes containers possible. Learn how mount, PID, network, and user namespaces create isolated environments, with interactive demos.

Compare Wayland vs X11 display servers on Linux. Learn about architecture, performance, security, and modern graphics stack.

Control groups as instruments: noisy neighbors, v1 vs v2 membership, CFS cpu.max periods, memory.high vs memory.max OOM, and how Docker flags become cgroup files.

Discover how containers work by combining namespaces, cgroups, and OverlayFS. Build a mental model of Docker internals through interactive visualizations.

Understand how containerized processes access GPU hardware through device files, bind mounts, and the NVIDIA container runtime. Learn the kernel driver vs user-space library distinction.

Learn nvidia-modeset for display configuration on Linux. Understand kernel mode-setting, DRM integration, and GPU drivers.

Language & Framework Internals/concepts/category/language-internals

How CPython, C++, and PyTorch work under the hood: bytecode, memory, linking, and concurrency.

Python Bytecode Compilation/concepts/language-internals/bytecode-compilation

Explore CPython bytecode compilation from source to .pyc files. Learn the dis module, PVM stack operations, and Python 3.11+ adaptive specialization.

PyTorch DataLoader Pipeline/concepts/language-internals/dataloader-pipeline

PyTorch DataLoader deep dive — Dataset, Sampler, Workers, Collate internals, num_workers throughput profiling, memory analysis, serialization costs, production patterns (LMDB, WebDataset), and bottleneck diagnosis.

C++ Stack vs Heap Memory: A Complete Guide/concepts/language-internals/stack-heap

Deep dive into C++ memory allocation — stack frame internals, heap allocator mechanics, fragmentation, performance benchmarks, custom allocators, RAII, and debugging with AddressSanitizer and Valgrind.

Python Memory Management/concepts/language-internals/memory-management

A practical mental model for CPython memory management: names and references, object headers, PyMalloc arenas, reference counting, reuse paths, and memory profiling.

Python Global Interpreter Lock (GIL)/concepts/language-internals/global-interpreter-lock

Learn the CPython Global Interpreter Lock (GIL) from first principles: why it exists, how threads take turns, why I/O still works well, and when to use multiprocessing, asyncio, or native extensions.

Pinned Memory and DMA Transfers in PyTorch/concepts/language-internals/pin-memory

Why pin_memory=True matters: pageable paths pay two host copies and block the CPU; pinned memory enables one DMA hop and real overlap with GPU compute.

C++ Symbol Resolution: How the Linker Connects Your Code/concepts/language-internals/symbol-resolution

Complete guide to C++ symbol resolution — how linkers match references to definitions, name mangling, strong vs weak symbols, ODR, template instantiation, linking order, and debugging undefined reference errors.

C++ Program Loading: From ELF to Running Process/concepts/language-internals/loading

How C++ programs are loaded — ELF segments, the _start to main() chain, dynamic linking with PLT/GOT, ASLR, real readelf/strace/proc maps output, and startup debugging.

Understanding num_workers/concepts/language-internals/num-workers

Why DataLoader num_workers matters: processes hide load latency behind GPU work, how to find the sweet spot, and the memory/GIL pitfalls that come with the pool.

Python Object Model Internals/concepts/language-internals/object-model

Learn how CPython implements PyObject, type objects, and the unified object model. Explore reference counting, memory layout, and Python internals.

Python Garbage Collection/concepts/language-internals/garbage-collection

Understand CPython garbage collection: reference counting, generational GC for circular references, weak references, and gc module tuning strategies.

Python Optimization Techniques/concepts/language-internals/python-optimization

Learn a profiler-first Python optimization workflow: measure bottlenecks, choose the right lever, and verify performance changes.

Python __slots__ Optimization/concepts/language-internals/slots-optimization

Learn when Python __slots__ reduces memory, how slot storage differs from __dict__, and the caveats for dataclasses and inheritance.

Python Green Threads vs OS Threads/concepts/language-internals/green-threads-vs-os-threads

Complete guide to Python concurrency — OS threads, green threads (asyncio), the GIL, event loop internals, Python 3.13 free-threading, and production patterns.

Python asyncio Event Loop/concepts/language-internals/asyncio-event-loop

Deep dive into Python's asyncio library, understanding event loops, coroutines, tasks, and async/await patterns with interactive visualizations.

Python Shared Memory/concepts/language-internals/shared-memory

Master Python multiprocessing.shared_memory for zero-copy IPC. Learn synchronization, NumPy integration, and race condition prevention patterns.

Thread Safety: Concurrent Programming Fundamentals/concepts/language-internals/thread-safety

Complete C++ thread safety guide — race conditions with step-through simulation, mutexes, atomics, condition variables, deadlock detection, memory ordering, and Thread Sanitizer walkthrough.

C++ AST & Parsing Explained/concepts/language-internals/ast-parsing

Explore how C++ code is parsed into an Abstract Syntax Tree (AST). Learn lexical analysis, tokenization, and syntax parsing for systems programming.

C++ Compilation Overview/concepts/language-internals/compilation

Understand the complete C++ compilation pipeline from source code to object files. Learn preprocessing, parsing, code generation, and optimization stages.

C++ Dynamic Linking at Runtime/concepts/language-internals/dynamic-linking

Deep dive into dynamic linking — GOT/PLT lazy resolution, shared library creation, SONAME versioning, RPATH/RUNPATH, dlopen plugin systems, LD_PRELOAD, and debugging with LD_DEBUG.

C++ Linking Overview/concepts/language-internals/linking

How C++ object files are linked into executables. Learn symbol resolution, static vs dynamic linking, and linker optimization.

Memory Management & RAII in C++/concepts/language-internals/memory-raii

Learn Resource Acquisition Is Initialization (RAII) - the cornerstone of C++ memory management. Understand automatic resource cleanup and exception safety.

Modern C++ Features (C++11 and Beyond)/concepts/language-internals/modern-cpp-features

Explore modern C++ features including auto, lambdas, ranges, and coroutines. Learn how C++11/14/17/20 transformed the language.

Object-Oriented Programming in C++/concepts/language-internals/oop-inheritance

Master C++ OOP concepts including inheritance, polymorphism, virtual functions, and modern object-oriented design principles with interactive examples.

C++ Compiler Optimization/concepts/language-internals/optimization

C++ compiler optimization lab notebook — compare optimization levels, inspect compiler rewrites, diagnose auto-vectorization, and build production verification commands.

Pointers & References in C++/concepts/language-internals/pointers-references

Master C++ pointers and references through interactive visualizations. Learn memory addressing, dereferencing, smart pointers, and avoid common pitfalls.

C++ Preprocessor Directives/concepts/language-internals/preprocessor

C++ preprocessor visualized: macros, header guards, conditional compilation, and #include directives explained interactively.

Smart Pointers in Modern C++/concepts/language-internals/smart-pointers

Master C++11 smart pointers through interactive examples. Learn unique_ptr, shared_ptr, and weak_ptr with reference counting visualizations.

Templates & STL in C++/concepts/language-internals/templates-stl

Master C++ templates and the Standard Template Library. Learn generic programming, template metaprogramming, and STL containers and algorithms.

DataParallel vs DistributedDataParallel/concepts/language-internals/data-parallel

Compare PyTorch DataParallel vs DistributedDataParallel for multi-GPU training. Learn GIL limitations, NCCL AllReduce, and DDP best practices.

C++ Virtual Tables & Inheritance/concepts/language-internals/virtual-tables-inheritance

C++ virtual tables (vtables) explained. Learn virtual dispatch, single/multiple inheritance, RTTI, and object memory layout visually.

Comparisons/comparisons

Side-by-side deep dives on technical trade-offs

Expertise/expertise

Technical areas I cover, each linked to its reference page

Uses/uses

Tools, software, and hardware I use

Resume/resume

My professional experience and qualifications

Bookmarks/bookmarks

A curated collection of articles and resources I find valuable

Consulting/consulting

Services and consulting offerings

Sitemap/sitemap

Visual representation of the site structure

Privacy/privacy

How this site handles your data, cookies, and analytics

Mastodon