GPU Kernels Explained
How a GPU actually executes the maths.
A GPU can do arithmetic far faster than it can fetch the numbers to do it on. An A100's FP32 units need 12.5 flops of work on every byte from its HBM just to keep busy. Every fast kernel is a way of closing that gap: reuse data close to the ALUs, fetch it in whole transactions, avoid serialising on shared memory, keep enough warps in flight.
Each chapter here is built around an animation, and every frame is computed by a small GPU execution model that is tested against a Python reference. The hardware figures come from NVIDIA's whitepapers and datasheets, the CUDA Programming Guide and published microbenchmarks.
01
The memory hierarchy
Registers, shared memory, L2 and HBM: how much each holds, how fast it moves data, and which one sets a kernel's pace.
02
The roofline
Arithmetic intensity decides whether a kernel waits on memory or on the ALUs; the ridge point is where they meet.
03
Warps, SIMT and divergence
32 threads share one instruction stream; a branch they disagree on runs both ways, with lanes masked off.
04
Coalescing
How a warp's 32 addresses become 32-byte memory transactions, and what stride and alignment cost.
05
Shared-memory bank conflicts
32 banks, one word each per cycle: why a column read serialises 32 ways, and how one word of padding fixes it.
06
Occupancy
Registers, shared memory and warp slots decide how many blocks share an SM, and so how much latency it can hide.
07
GEMM, step by step
Naive, tiled in shared memory, register-blocked, tensor cores: each step reuses data closer to the ALUs and raises the arithmetic intensity.
08
Reductions and warp shuffles
Summing a block's values as a tree in shared memory, three ways, then register to register with warp shuffles.
09
Softmax and FlashAttention
Online softmax keeps a running maximum and rescales; FlashAttention uses it to never write the score matrix to HBM.
10
Split-K, streams and overlap
Double buffering hides copies behind compute; split-K makes enough blocks to fill the GPU when the output is small.
11
Quantised kernels
Low-precision weights cut the bytes a decode step reads; the kernel dequantises them in registers, just before the multiply.
Every number on the GPUs the model uses, with its source: the GPU presets.
Part of a family of companion sites: the Transformer Decoder Explainer (one forward pass), LLM Inference Explained (serving it), LLM Architectures Explained (how the models differ), Numerics Explained (the numbers themselves) and Systolic Arrays Explained (the matrix hardware of TPUs). This site is the layer underneath: the kernels. How it was built, and how to check it: about.