gpu-kernels-explained

GPU Kernels Explained

How a GPU actually executes the maths.

A GPU can do arithmetic far faster than it can fetch the numbers to do it on. An A100's FP32 units need 12.5 flops of work on every byte from its HBM just to keep busy. Every fast kernel is a way of closing that gap: reuse data close to the ALUs, fetch it in whole transactions, avoid serialising on shared memory, keep enough warps in flight.

Each chapter here is built around an animation, and every frame is computed by a small GPU execution model that is tested against a Python reference. The hardware figures come from NVIDIA's whitepapers and datasheets, the CUDA Programming Guide and published microbenchmarks.

Registers117 TB/sShared memory19.49 TB/sL2 cache7.219 TB/sHBM1.555 TB/s
The A100's memory levels, bandwidth to scale. HBM is the sliver at the bottom.

Every number on the GPUs the model uses, with its source: the GPU presets.

Part of a family of companion sites: the Transformer Decoder Explainer (one forward pass), LLM Inference Explained (serving it), LLM Architectures Explained (how the models differ), Numerics Explained (the numbers themselves) and Systolic Arrays Explained (the matrix hardware of TPUs). This site is the layer underneath: the kernels. How it was built, and how to check it: about.