Quantised kernels
Low-precision weights cut the bytes a decode step reads; the kernel dequantises them in registers, just before the multiply.
Loading the animation…
Concept
Chapter 2 put batch-1 decoding far down the roofline's slope: each weight is read from HBM for a single multiply-add. When a kernel is memory-bound, the way to make it faster is to read fewer bytes, and the bytes are almost all weights. Store each weight in 8 or 4 bits instead of 16 and the time falls by nearly the same factor, as long as the kernel stays memory-bound.
The arithmetic still happens in BF16 (for weight-only formats such as W4A16): the kernel loads the packed weights, dequantises them in registers, just before the multiply-add, and never writes the BF16 weights anywhere. The animation does it for one 32-bit register holding eight 4-bit weights: shift and mask out each 4-bit integer , subtract the offset 8 to centre it, multiply by the scale of its group (here one scale per 128 weights, stored in 16 bits), and multiply-add with the activation. Production kernels unpack several weights at once with bit-manipulation tricks and feed the results to the tensor cores; the order of operations is the same.
When the whole layer has enough tokens to be compute-bound, the bytes stop mattering and the tensor cores' rate does. A W4A16 kernel still runs at the BF16 tensor peak; only formats whose maths is narrower go faster there: INT8 tensor cores (624 TFLOP/s dense on the A100, twice its BF16 rate) for W8A8, or FP8 on the H100.
Loading the animation…
Concept
The second animation applies one 8192 × 8192 layer (67 million weights) to a growing batch of tokens. At batch 1 on the A100 the model gives 86.3 µs in BF16, 43.2 µs in INT8 and 22.3 µs in INT4: 34.6 MB of INT4 weights and scales instead of 134 MB. The lines stay flat while the batch shares one read of the weights, then bend upwards where the tensor cores take over, and the INT4 line bends first, because it has fewer bytes to hide behind. From there on, the weight-only formats are no faster than BF16, and only W8A8, with its INT8 tensor cores, keeps an advantage.
So quantisation is mainly a decode-time (small-batch) optimisation for weights, and a compute-time one only when the maths itself is narrower. The accuracy side, which formats and scales keep a model's quality, is the subject of Numerics Explained: its outliers chapter shows why W8A8 needs care, and its GPTQ and AWQ and NF4 chapters how 4-bit weights keep their accuracy.
Maths
A weight matrix in -bit integers, with one 16-bit scale per group of weights, takes
bytes: for INT4 with , against in BF16, a ratio of 3.88. For a batch of tokens the layer does flops, so its intensity at batch is about
and it reaches the ridge at : for BF16 on the A100 that is 201, while INT4 gets there at about a quarter of it. Above , whatever the storage format.
Code
The dequantise step, cut from src/lib/gpu/model.ts:
const qi = (word >>> (4 * i)) & 0xf;
const w = (qi - 8) * scale;
The bytes of a quantised layer:
let wbytes = idiv(rows * cols * f.bits, 8);
if (f.group === -1) wbytes += 2 * rows;
else if (f.group > 0) wbytes += 2 * idiv(rows * cols, f.group);
The same unpacking in CUDA, illustrative and not compiled by this site's CI:
uint32_t r = packed[i]; // eight 4-bit weights
#pragma unroll
for (int j = 0; j < 8; ++j) {
int q = (r >> (4 * j)) & 0xF; // unpack
float w = (q - 8) * scale; // dequantise in a register
acc = fmaf(w, x[8 * i + j], acc); // and use it at once
}
One 8192 × 8192 layer (model)
| Format | Weight bytes | Batch 1, A100 | Batch 1, H100 | Batch 1024, A100 |
|---|---|---|---|---|
| BF16 | 134 MB | 86.3 µs | 40.1 µs | 441 µs |
| INT8 weights (W8A16) | 67.1 MB | 43.2 µs | 20 µs | 441 µs |
| INT4 weights, g = 128 (W4A16) | 34.6 MB | 22.3 µs | 10.3 µs | 441 µs |
| W8A8 (INT8 tensor cores) | 67.1 MB | 43.2 µs | 20 µs | 220 µs |