gpu-kernels-explained

/about

About this site

GPU Kernels Explained shows how a GPU executes the maths behind modern models, one mechanism per chapter, each around an animation. It is the fourth of a family of companion sites: the Transformer Decoder Explainer, LLM Inference Explained, LLM Architectures Explained, Numerics Explained, Systolic Arrays Explained and Inference Trade-offs Explained. Each chapter links the matching slides of the NVIDIA GPU and CUDA series.

The execution model

reference/gpu_model.py is an illustrative, parameterised model of a GPU: SM count, clock, FP32 and tensor peaks, warps, registers, shared memory with 32 banks, L2 and HBM bandwidths. For a kernel configuration it computes the bytes moved at each memory level, arithmetic intensity and the roofline position, occupancy, bank conflicts, coalescing (32-byte sectors per warp request), the active mask of every instruction a diverging warp issues, and a load / compute / store timeline per tile; and, for the later chapters, a GEMM four ways, the steps of four reductions, online softmax and FlashAttention's memory traffic, split-K and quantised layers. It is not cycle-accurate.

How it is checked

The animations

Every animation has play and pause, step back and forward, a scrub bar, speeds from 0.25× to 4× and reset; with the animation focused, Space plays or pauses and the arrow keys step. Each step has a one-line caption, also announced to screen readers. With reduce motion set in your system, nothing plays by itself. Animations pause when scrolled out of view. Colours come from Okabe and Ito's colour-blind-safe palette, one colour per memory level (registers orange, shared memory green, L2 sky blue, HBM purple), the same in light and dark mode; a stall is always a hatched pattern as well as a warning colour.

Source

The code, the model and the tests are on GitHub (MIT licence). The design system is copied from the companion sites; the README records where each piece came from.