Occupancy
Registers, shared memory and warp slots decide how many blocks share an SM, and so how much latency it can hide.
Loading the animation…
Concept
An SM keeps many warps resident at once and switches between them every cycle: while one waits hundreds of cycles for memory (chapter 1's latencies), the others compute. Occupancy is "the ratio of the number of active warps to the maximum number of active warps supported by the SM" (CUDA Programming Guide §2.3.7). On the A100 and H100 the maximum is 64 warps (2,048 threads) per SM.
How many warps fit depends on what each block asks for. The scheduler places a block on an SM only if the SM still has room for all of it: its warp slots, its registers (the SM has 65,536), its shared memory (up to 164 KB on the A100, 228 KB on the H100) and one of the 32 block slots. The animation places blocks one at a time until the next one does not fit; the resource that ran out is the limiter.
Try the guide's two examples on the A100 (shared memory at 0):
- 768 threads per block: two blocks use 48 of 64 warp slots, and a third would need 72. Occupancy 75%.
- 32 threads per block: each block is one warp, so the 32-block limit stops the SM at 32 warps. Occupancy 50%.
Then raise the registers: at 64 registers per thread, 256-thread blocks stop at 4 blocks (32 warps, 50%). Add 48 KB of shared memory and only 3 fit on the A100 (37.5%), while the H100's larger shared memory still takes 4.
High occupancy is a means, not the goal. A kernel that keeps enough independent work in each thread (instruction-level parallelism, such as the 8×8 register tiles of chapter 1's fastest GEMM) can hide latency with fewer warps, and the registers that buys are often worth more than the occupancy they cost.
Maths
A block of threads is warps. With registers per thread, each warp is allocated registers, from one of 4 sub-partitions of registers. With bytes of shared memory per block, the block is allocated bytes. Then the SM holds
blocks, and the occupancy is .
Worked example (A100, 256 threads, 64 registers, 48 KB): ; , so warps per sub-partition, blocks; , so blocks. , occupancy , limited by shared memory.
Code
The register limit, cut from src/lib/gpu/model.ts (a port of the same rule in cuda_occupancy.h):
const regsPerWarp = roundUp(regs * p.warp_size, p.reg_alloc_unit);
const regsPerBlock = regsPerWarp * warpsPerBlock;
const regsAssumed = regsPerWarp * roundUp(warpsPerBlock, p.sub_partitions);