/learn
How a GPU runs a kernel
One chapter per mechanism, each opening with an animation. Every frame is computed by this site's GPU execution model (checked against its Python reference), and every hardware figure comes from a published source. Toggle layers (Concept / Maths / Code) inside any chapter to choose how deep to go. The model architectures these kernels serve are on LLM Architectures Explained; for the slides behind each chapter, see the NVIDIA GPU and CUDA series.
Registers, shared memory, L2 and HBM: how much each holds, how fast it moves data, and which one sets a kernel's pace.
Arithmetic intensity decides whether a kernel waits on memory or on the ALUs; the ridge point is where they meet.
32 threads share one instruction stream; a branch they disagree on runs both ways, with lanes masked off.
How a warp's 32 addresses become 32-byte memory transactions, and what stride and alignment cost.
32 banks, one word each per cycle: why a column read serialises 32 ways, and how one word of padding fixes it.
Registers, shared memory and warp slots decide how many blocks share an SM, and so how much latency it can hide.
Naive, tiled in shared memory, register-blocked, tensor cores: each step reuses data closer to the ALUs and raises the arithmetic intensity.
Summing a block's values as a tree in shared memory, three ways, then register to register with warp shuffles.
Online softmax keeps a running maximum and rescales; FlashAttention uses it to never write the score matrix to HBM.
Double buffering hides copies behind compute; split-K makes enough blocks to fill the GPU when the output is small.
Low-precision weights cut the bytes a decode step reads; the kernel dequantises them in registers, just before the multiply.