gpu-kernels-explained
← /learn · 06

Occupancy

Registers, shared memory and warp slots decide how many blocks share an SM, and so how much latency it can hide.

Loading the animation…

Concept

An SM keeps many warps resident at once and switches between them every cycle: while one waits hundreds of cycles for memory (chapter 1's latencies), the others compute. Occupancy is "the ratio of the number of active warps to the maximum number of active warps supported by the SM" (CUDA Programming Guide §2.3.7). On the A100 and H100 the maximum is 64 warps (2,048 threads) per SM.

How many warps fit depends on what each block asks for. The scheduler places a block on an SM only if the SM still has room for all of it: its warp slots, its registers (the SM has 65,536), its shared memory (up to 164 KB on the A100, 228 KB on the H100) and one of the 32 block slots. The animation places blocks one at a time until the next one does not fit; the resource that ran out is the limiter.

Try the guide's two examples on the A100 (shared memory at 0):

  • 768 threads per block: two blocks use 48 of 64 warp slots, and a third would need 72. Occupancy 75%.
  • 32 threads per block: each block is one warp, so the 32-block limit stops the SM at 32 warps. Occupancy 50%.

Then raise the registers: at 64 registers per thread, 256-thread blocks stop at 4 blocks (32 warps, 50%). Add 48 KB of shared memory and only 3 fit on the A100 (37.5%), while the H100's larger shared memory still takes 4.

High occupancy is a means, not the goal. A kernel that keeps enough independent work in each thread (instruction-level parallelism, such as the 8×8 register tiles of chapter 1's fastest GEMM) can hide latency with fewer warps, and the registers that buys are often worth more than the occupancy they cost.

Maths

A block of TT threads is w=⌈T/32⌉w = \lceil T / 32 \rceil warps. With rr registers per thread, each warp is allocated rw=256⌈32r/256⌉r_w = 256 \lceil 32r / 256 \rceil registers, from one of 4 sub-partitions of 65536/4=1638465536 / 4 = 16384 registers. With ss bytes of shared memory per block, the block is allocated sb=128⌈(s+1024)/128⌉s_b = 128 \lceil (s + 1024)/128 \rceil bytes. Then the SM holds

B=min⁡(⌊64w⌋, ⌊4⌊16384/rw⌋w⌋, ⌊SSMsb⌋, 32)B = \min\left(\left\lfloor \frac{64}{w} \right\rfloor,\ \left\lfloor \frac{4 \lfloor 16384 / r_w \rfloor}{w} \right\rfloor,\ \left\lfloor \frac{S_{\text{SM}}}{s_b} \right\rfloor,\ 32\right)

blocks, and the occupancy is B w/64B\,w / 64.

Worked example (A100, 256 threads, 64 registers, 48 KB): w=8w = 8; rw=2048r_w = 2048, so ⌊16384/2048⌋=8\lfloor 16384/2048 \rfloor = 8 warps per sub-partition, 4×8/8=44 \times 8 / 8 = 4 blocks; sb=49×1024s_b = 49 \times 1024, so ⌊164/49⌋=3\lfloor 164/49 \rfloor = 3 blocks. B=min⁡(8,4,3,32)=3B = \min(8, 4, 3, 32) = 3, occupancy 3×8/64=37.5%3 \times 8 / 64 = 37.5\%, limited by shared memory.

Code

The register limit, cut from src/lib/gpu/model.ts (a port of the same rule in cuda_occupancy.h):

const regsPerWarp = roundUp(regs * p.warp_size, p.reg_alloc_unit);
const regsPerBlock = regsPerWarp * warpsPerBlock;
const regsAssumed = regsPerWarp * roundUp(warpsPerBlock, p.sub_partitions);