/about
About this site
GPU Kernels Explained shows how a GPU executes the maths behind modern models, one mechanism per chapter, each around an animation. It is the fourth of a family of companion sites: the Transformer Decoder Explainer, LLM Inference Explained, LLM Architectures Explained, Numerics Explained, Systolic Arrays Explained and Inference Trade-offs Explained. Each chapter links the matching slides of the NVIDIA GPU and CUDA series.
The execution model
reference/gpu_model.py is an illustrative, parameterised model of a GPU: SM count, clock, FP32 and tensor peaks, warps, registers, shared memory with 32 banks, L2 and HBM bandwidths. For a kernel configuration it computes the bytes moved at each memory level, arithmetic intensity and the roofline position, occupancy, bank conflicts, coalescing (32-byte sectors per warp request), the active mask of every instruction a diverging warp issues, and a load / compute / store timeline per tile; and, for the later chapters, a GEMM four ways, the steps of four reductions, online softmax and FlashAttention's memory traffic, split-K and quantised layers. It is not cycle-accurate.
- Every preset figure has a source and a status (see the GPU presets).
- Kernel time is a hierarchical roofline: each memory level and the ALUs work in parallel at peak, and the slowest sets the time. Latency, instruction issue and cache behaviour beyond a simple fits-in-L2 rule are not modelled.
- Occupancy follows the CUDA Toolkit's
cuda_occupancy.hfor compute capability 8.0 and 9.0, with the shared-memory carveout at its maximum. - Divergence follows the classic SIMT order (if-path, else-path, reconverge); since Volta the scheduler may order the paths differently, at the same cost.
How it is checked
- Closed forms. tests/python/test_gpu_model.py checks the reference against its sources' worked examples: the CUDA Programming Guide's 12.5% coalescing worst case, two-way and 32-way bank conflicts and the padding fix, the 75% and 50% occupancy examples, and the GEMM traffic formulas.
- Exact parity. The TypeScript port, src/lib/gpu/model.ts, repeats every function with the same operations in the same order; scripts/make_fixtures.py writes the reference's results over grids of every parameter, and the unit tests require the port to reproduce all of them exactly, with no tolerance (the online softmax, which calls
exp, to a relative 10−14, because JavaScript's and the C library'sexpcan differ in the last bit). CI fails if the fixtures are out of date. - Animations from the model. Every animation draws a sequence of states the model computes; a frame is a pure function of one state. The end-to-end tests set chosen frames of every animation and require the caption to match the caption built from the Python reference's state for that frame.
- Numbers in the prose are printed from the model when the page is built, not typed.
The animations
Every animation has play and pause, step back and forward, a scrub bar, speeds from 0.25× to 4× and reset; with the animation focused, Space plays or pauses and the arrow keys step. Each step has a one-line caption, also announced to screen readers. With reduce motion set in your system, nothing plays by itself. Animations pause when scrolled out of view. Colours come from Okabe and Ito's colour-blind-safe palette, one colour per memory level (registers orange, shared memory green, L2 sky blue, HBM purple), the same in light and dark mode; a stall is always a hatched pattern as well as a warning colour.
Source
The code, the model and the tests are on GitHub (MIT licence). The design system is copied from the companion sites; the README records where each piece came from.