/gpus
The GPU presets
The model is parameterised: any GPU is a set of the figures below. Two presets ship, each figure with a status. spec: NVIDIA states it. rule: the CUDA documentation or toolkit states it. measured: a published microbenchmark. derived: computed from other figures, with the formula. approximation: a stand-in, named as such. The presets are approximations of the real parts and the model is not cycle-accurate; the figures are what the animations compute from.
A100 SXM4 40 GB
Ampere (GA100), compute capability 8.0
| Figure | Value | Status | Source |
|---|---|---|---|
| Streaming multiprocessors (SMs) | 108 | spec | a100-wp: Table 4, p. 36 |
| Boost clock | 1,410 MHz | spec | a100-wp: Table 4 (GPU Boost Clock), p. 36 |
| FP32 cores per SM | 64 | spec | a100-wp: Table 4, p. 36 |
| Peak FP32 (CUDA cores) | 19.5 TFLOP/s | spec | a100-wp: Table 4 (Peak FP32 TFLOPS, non-Tensor), p. 36 |
| Peak BF16 tensor (dense) | 312 TFLOP/s | spec | a100-wp: Table 4 (Peak BF16 Tensor TFLOPS with FP32 accumulate, dense), p. 36 |
| HBM bandwidth | 1.555 TB/s | spec | a100-wp: Table 4 (Memory Bandwidth), p. 37 |
| HBM capacity | 40 GB | spec | a100-wp: Table 4 (Memory Size 40 GB), p. 36 |
| L2 cache | 40 MB | spec | a100-wp: Table 4 (L2 Cache Size 40960 KB), p. 37 |
| L2 bandwidth per clock | 5120 B/clk | spec | a100-wp: p. 35: 'The A100 L2 read bandwidth is 5120 Bytes/clk' |
| Shared memory per SM (max) | 164 KB | spec | a100-wp: Table 5 (configurable up to 164 KB), p. 43; CUDA guide Table 32 |
| Shared memory per block (max) | 163 KB | rule | cuda-cc: Table 32, compute capability 8.0 (max per thread block, opt-in) |
| 32-bit registers per SM | 65,536 | spec | a100-wp: Table 5 (Max 32-bit Registers / SM), p. 43 |
| Registers per block (max) | 65,536 | spec | a100-wp: Table 5 (Max Registers / Block), p. 43 |
| Registers per thread (max) | 255 | spec | a100-wp: Table 5 (Max Registers / Thread), p. 43 |
| Warps per SM (max) | 64 | spec | a100-wp: Table 5 (Max Warps / SM), p. 43 |
| Blocks per SM (max) | 32 | spec | a100-wp: Table 5 (Max Thread Blocks / SM), p. 43 |
| Threads per block (max) | 1024 | spec | a100-wp: Table 5 (Max Thread Block Size), p. 43 |
| Shared-memory latency | 29 cycles | measured | luo2024: Table IV, A100 PCIe (cycles) |
| L2 latency | 261.5 cycles | measured | luo2024: Table IV, A100 PCIe (cycles) |
| HBM (global) latency | 466.3 cycles | measured | luo2024: Table IV, A100 PCIe (cycles) |
| mma_shape | mma.sync m16n8k16 (BF16, one warp) | rule | ptx: #warp-level-matrix-shape |
| peak_int8_tensor | 624000000000000 | spec | a100-wp: Table 4 (Peak INT8 Tensor TOPS, dense), p. 36 |
Derived
| Figure | Value | Formula |
|---|---|---|
| Register bandwidth | 117 TB/s | 6 bytes per flop × peak FP32 (three 4-byte operands per 2-flop FFMA) |
| Shared-memory bandwidth | 19.49 TB/s | SMs × 32 banks × 4 bytes × clock |
| L2 bandwidth | 7.219 TB/s | L2 bytes per clock × clock |
| FP32 ridge point | 12.5 flop/byte | peak FP32 / HBM bandwidth |
| BF16 tensor ridge point | 201 flop/byte | peak tensor / HBM bandwidth |
H100 SXM5 80 GB
Hopper (GH100), compute capability 9.0
| Figure | Value | Status | Source |
|---|---|---|---|
| Streaming multiprocessors (SMs) | 132 | spec | h100-wp: Table 3 (H100 SXM5), p. 39 |
| Boost clock | 1,982.7 MHz | derived | h100-page: peak FP32 / (2 x FP32 cores): 67e12 / (2 x 132 x 128) Hz; the whitepaper lists the clock as not finalised |
| FP32 cores per SM | 128 | spec | h100-wp: Table 3 (FP32 Cores / SM), p. 39 |
| Peak FP32 (CUDA cores) | 67 TFLOP/s | spec | h100-page: FP32 67 teraFLOPS |
| Peak BF16 tensor (dense) | 990 TFLOP/s | derived | h100-page: BFLOAT16 Tensor Core 1,979 teraFLOPS is 'with sparsity'; dense is half |
| HBM bandwidth | 3.35 TB/s | spec | h100-page: GPU Memory Bandwidth 3.35TB/s |
| HBM capacity | 80 GB | spec | h100-page: GPU Memory 80GB |
| L2 cache | 50 MB | spec | h100-wp: Table 3 (L2 Cache Size 50 MB), p. 40 |
| L2 bandwidth per clock | 4472.3 B/clk | approximation | luo2024: Table V: L2 throughput measured on H800 PCIe (same GH100 die), FP32 loads; NVIDIA does not publish H100's |
| Shared memory per SM (max) | 228 KB | spec | h100-wp: Table 4 (configurable up to 228 KB), p. 41; CUDA guide Table 32 |
| Shared memory per block (max) | 227 KB | rule | cuda-cc: Table 32, compute capability 9.0 (max per thread block, opt-in) |
| 32-bit registers per SM | 65,536 | spec | h100-wp: Table 4 (Max 32-bit Registers / SM), p. 41 |
| Registers per block (max) | 65,536 | spec | h100-wp: Table 4 (Max Registers / Thread Block), p. 41 |
| Registers per thread (max) | 255 | spec | h100-wp: Table 4 (Max Registers / Thread), p. 41 |
| Warps per SM (max) | 64 | spec | h100-wp: Table 4 (Max Warps / SM), p. 41 |
| Blocks per SM (max) | 32 | spec | h100-wp: Table 4 (Max Thread Blocks / SM), p. 41 |
| Threads per block (max) | 1024 | spec | h100-wp: Table 4 (Max Thread Block Size), p. 41 |
| Shared-memory latency | 29 cycles | measured | luo2024: Table IV, H800 PCIe (cycles) |
| L2 latency | 263 cycles | measured | luo2024: Table IV, H800 PCIe (cycles) |
| HBM (global) latency | 478.8 cycles | measured | luo2024: Table IV, H800 PCIe (cycles) |
| mma_shape | wgmma.mma_async m64nNk16 (BF16, one warpgroup, N up to 256) | rule | ptx: #asynchronous-warpgroup-level-matrix-shape |
| peak_int8_tensor | 1979000000000000 | derived | h100-page: INT8 Tensor Core 3,958 TOPS is 'with sparsity'; dense is half |
Derived
| Figure | Value | Formula |
|---|---|---|
| Register bandwidth | 402 TB/s | 6 bytes per flop × peak FP32 (three 4-byte operands per 2-flop FFMA) |
| Shared-memory bandwidth | 33.5 TB/s | SMs × 32 banks × 4 bytes × clock |
| L2 bandwidth | 8.867 TB/s | L2 bytes per clock × clock |
| FP32 ridge point | 20 flop/byte | peak FP32 / HBM bandwidth |
| BF16 tensor ridge point | 295 flop/byte | peak tensor / HBM bandwidth |
Rules shared by both
| Figure | Value | Status | Source |
|---|---|---|---|
| Threads per warp | 32 | rule | cuda-kernels: §2.3 (a warp is 32 threads) |
| Shared-memory banks | 32 | rule | cuda-kernels: #shared-memory-access-patterns: 32 banks |
| Bank width | 4 bytes per clock | rule | cuda-kernels: #shared-memory-access-patterns: successive 32-bit words, 32 bits per bank per clock |
| Global-memory transaction (sector) | 32 bytes | rule | cuda-kernels: #coalesced-global-memory-access: 32-byte transactions |
| Register allocation unit (per warp) | 256 registers | rule | cuda-occ: cudaOccRegAllocationGranularity: 256 registers per warp allocation |
| Shared-memory allocation unit | 128 bytes | rule | cuda-occ: cudaOccSMemAllocationGranularity: 128 bytes (CC 8.x, 9.x) |
| Shared memory reserved per block | 1024 bytes | rule | cuda-occ: reservedSharedMemPerBlock: 1 KB per block (per-SM minus per-block maximum, Table 32) |
| SM sub-partitions | 4 | rule | cuda-occ: cudaOccSubPartitionsPerMultiprocessor: 4 |
Sources
- a100-wp: NVIDIA A100 Tensor Core GPU Architecture (whitepaper)
- cuda-cc: CUDA Programming Guide, Compute capabilities (Table 32: shared memory)
- cuda-kernels: CUDA Programming Guide, §2.3 Writing CUDA kernels
- cuda-occ: cuda_occupancy.h, the CUDA Toolkit's occupancy calculator header (CUDA 12.5 copy)
- h100-page: NVIDIA H100 product page, Product Specifications (H100 SXM column)
- h100-wp: NVIDIA H100 Tensor Core GPU Architecture (whitepaper, GTC 2022; preliminary specifications)
- luo2024: Luo et al., Benchmarking and Dissecting the Nvidia Hopper GPU Architecture (arXiv 2402.13499)
- ptx: PTX ISA: matrix shapes for mma and wgmma
The H100 whitepaper was published before launch and lists its specifications as preliminary; the H100 figures that changed (FP32 and tensor peaks, HBM bandwidth) are taken from NVIDIA's current product page instead, and its clock is derived from the FP32 peak. NVIDIA does not publish H100's L2 bandwidth, so the model uses the per-clock rate Luo et al. measured on an H800 (the same GH100 die). Luo et al. measured about 2,008 bytes per clock on an A100 against NVIDIA's stated 5,120, so the A100's L2 figure is a peak the model treats as attainable.