gpu-kernels-explained

/gpus

The GPU presets

The model is parameterised: any GPU is a set of the figures below. Two presets ship, each figure with a status. spec: NVIDIA states it. rule: the CUDA documentation or toolkit states it. measured: a published microbenchmark. derived: computed from other figures, with the formula. approximation: a stand-in, named as such. The presets are approximations of the real parts and the model is not cycle-accurate; the figures are what the animations compute from.

A100 SXM4 40 GB

Ampere (GA100), compute capability 8.0

Registers117 TB/sShared memory19.49 TB/sL2 cache7.219 TB/sHBM1.555 TB/s
FigureValueStatusSource
Streaming multiprocessors (SMs)108speca100-wp: Table 4, p. 36
Boost clock1,410 MHzspeca100-wp: Table 4 (GPU Boost Clock), p. 36
FP32 cores per SM64speca100-wp: Table 4, p. 36
Peak FP32 (CUDA cores)19.5 TFLOP/sspeca100-wp: Table 4 (Peak FP32 TFLOPS, non-Tensor), p. 36
Peak BF16 tensor (dense)312 TFLOP/sspeca100-wp: Table 4 (Peak BF16 Tensor TFLOPS with FP32 accumulate, dense), p. 36
HBM bandwidth1.555 TB/sspeca100-wp: Table 4 (Memory Bandwidth), p. 37
HBM capacity40 GBspeca100-wp: Table 4 (Memory Size 40 GB), p. 36
L2 cache40 MBspeca100-wp: Table 4 (L2 Cache Size 40960 KB), p. 37
L2 bandwidth per clock5120 B/clkspeca100-wp: p. 35: 'The A100 L2 read bandwidth is 5120 Bytes/clk'
Shared memory per SM (max)164 KBspeca100-wp: Table 5 (configurable up to 164 KB), p. 43; CUDA guide Table 32
Shared memory per block (max)163 KBrulecuda-cc: Table 32, compute capability 8.0 (max per thread block, opt-in)
32-bit registers per SM65,536speca100-wp: Table 5 (Max 32-bit Registers / SM), p. 43
Registers per block (max)65,536speca100-wp: Table 5 (Max Registers / Block), p. 43
Registers per thread (max)255speca100-wp: Table 5 (Max Registers / Thread), p. 43
Warps per SM (max)64speca100-wp: Table 5 (Max Warps / SM), p. 43
Blocks per SM (max)32speca100-wp: Table 5 (Max Thread Blocks / SM), p. 43
Threads per block (max)1024speca100-wp: Table 5 (Max Thread Block Size), p. 43
Shared-memory latency29 cyclesmeasuredluo2024: Table IV, A100 PCIe (cycles)
L2 latency261.5 cyclesmeasuredluo2024: Table IV, A100 PCIe (cycles)
HBM (global) latency466.3 cyclesmeasuredluo2024: Table IV, A100 PCIe (cycles)
mma_shapemma.sync m16n8k16 (BF16, one warp)ruleptx: #warp-level-matrix-shape
peak_int8_tensor624000000000000speca100-wp: Table 4 (Peak INT8 Tensor TOPS, dense), p. 36

Derived

FigureValueFormula
Register bandwidth117 TB/s6 bytes per flop × peak FP32 (three 4-byte operands per 2-flop FFMA)
Shared-memory bandwidth19.49 TB/sSMs × 32 banks × 4 bytes × clock
L2 bandwidth7.219 TB/sL2 bytes per clock × clock
FP32 ridge point12.5 flop/bytepeak FP32 / HBM bandwidth
BF16 tensor ridge point201 flop/bytepeak tensor / HBM bandwidth

H100 SXM5 80 GB

Hopper (GH100), compute capability 9.0

Registers402 TB/sShared memory33.5 TB/sL2 cache8.867 TB/sHBM3.35 TB/s
FigureValueStatusSource
Streaming multiprocessors (SMs)132spech100-wp: Table 3 (H100 SXM5), p. 39
Boost clock1,982.7 MHzderivedh100-page: peak FP32 / (2 x FP32 cores): 67e12 / (2 x 132 x 128) Hz; the whitepaper lists the clock as not finalised
FP32 cores per SM128spech100-wp: Table 3 (FP32 Cores / SM), p. 39
Peak FP32 (CUDA cores)67 TFLOP/sspech100-page: FP32 67 teraFLOPS
Peak BF16 tensor (dense)990 TFLOP/sderivedh100-page: BFLOAT16 Tensor Core 1,979 teraFLOPS is 'with sparsity'; dense is half
HBM bandwidth3.35 TB/sspech100-page: GPU Memory Bandwidth 3.35TB/s
HBM capacity80 GBspech100-page: GPU Memory 80GB
L2 cache50 MBspech100-wp: Table 3 (L2 Cache Size 50 MB), p. 40
L2 bandwidth per clock4472.3 B/clkapproximationluo2024: Table V: L2 throughput measured on H800 PCIe (same GH100 die), FP32 loads; NVIDIA does not publish H100's
Shared memory per SM (max)228 KBspech100-wp: Table 4 (configurable up to 228 KB), p. 41; CUDA guide Table 32
Shared memory per block (max)227 KBrulecuda-cc: Table 32, compute capability 9.0 (max per thread block, opt-in)
32-bit registers per SM65,536spech100-wp: Table 4 (Max 32-bit Registers / SM), p. 41
Registers per block (max)65,536spech100-wp: Table 4 (Max Registers / Thread Block), p. 41
Registers per thread (max)255spech100-wp: Table 4 (Max Registers / Thread), p. 41
Warps per SM (max)64spech100-wp: Table 4 (Max Warps / SM), p. 41
Blocks per SM (max)32spech100-wp: Table 4 (Max Thread Blocks / SM), p. 41
Threads per block (max)1024spech100-wp: Table 4 (Max Thread Block Size), p. 41
Shared-memory latency29 cyclesmeasuredluo2024: Table IV, H800 PCIe (cycles)
L2 latency263 cyclesmeasuredluo2024: Table IV, H800 PCIe (cycles)
HBM (global) latency478.8 cyclesmeasuredluo2024: Table IV, H800 PCIe (cycles)
mma_shapewgmma.mma_async m64nNk16 (BF16, one warpgroup, N up to 256)ruleptx: #asynchronous-warpgroup-level-matrix-shape
peak_int8_tensor1979000000000000derivedh100-page: INT8 Tensor Core 3,958 TOPS is 'with sparsity'; dense is half

Derived

FigureValueFormula
Register bandwidth402 TB/s6 bytes per flop × peak FP32 (three 4-byte operands per 2-flop FFMA)
Shared-memory bandwidth33.5 TB/sSMs × 32 banks × 4 bytes × clock
L2 bandwidth8.867 TB/sL2 bytes per clock × clock
FP32 ridge point20 flop/bytepeak FP32 / HBM bandwidth
BF16 tensor ridge point295 flop/bytepeak tensor / HBM bandwidth

Rules shared by both

FigureValueStatusSource
Threads per warp32rulecuda-kernels: §2.3 (a warp is 32 threads)
Shared-memory banks32rulecuda-kernels: #shared-memory-access-patterns: 32 banks
Bank width4 bytes per clockrulecuda-kernels: #shared-memory-access-patterns: successive 32-bit words, 32 bits per bank per clock
Global-memory transaction (sector)32 bytesrulecuda-kernels: #coalesced-global-memory-access: 32-byte transactions
Register allocation unit (per warp)256 registersrulecuda-occ: cudaOccRegAllocationGranularity: 256 registers per warp allocation
Shared-memory allocation unit128 bytesrulecuda-occ: cudaOccSMemAllocationGranularity: 128 bytes (CC 8.x, 9.x)
Shared memory reserved per block1024 bytesrulecuda-occ: reservedSharedMemPerBlock: 1 KB per block (per-SM minus per-block maximum, Table 32)
SM sub-partitions4rulecuda-occ: cudaOccSubPartitionsPerMultiprocessor: 4

Sources

The H100 whitepaper was published before launch and lists its specifications as preliminary; the H100 figures that changed (FP32 and tensor peaks, HBM bandwidth) are taken from NVIDIA's current product page instead, and its clock is derived from the FP32 peak. NVIDIA does not publish H100's L2 bandwidth, so the model uses the per-clock rate Luo et al. measured on an H800 (the same GH100 die). Luo et al. measured about 2,008 bytes per clock on an A100 against NVIDIA's stated 5,120, so the A100's L2 figure is a peak the model treats as attainable.