Skip to content
teach

Lesson 15. NF4 and Double Quantisation

Mission link: These are the two ideas QLoRA contributed. Knowing them separates understanding the paper from having read the abstract.
Primary source: Paper: "QLoRA: Efficient Finetuning of Quantized LLMs", Dettmers et al., arXiv:2305.14314
Prerequisites: Lesson 14

Warm-up

  1. ▢ Why blockwise rather than per-tensor quantisation?
Check

An outlier only sets the scale within its own block, containing the damage to 64 or 128 values instead of the whole tensor.

  1. ▢ True bytes per parameter for 4-bit weights with fp32 scales at block size 64?
Check

0.5625, so 4.5 bits, with 0.5 bits of pure bookkeeping overhead.

  1. ▢ What precision is the matmul in, for a 4-bit base?
Check

A higher one, typically bf16. Blocks are dequantised on the fly and the dequantized copies discarded.

Know this

QLoRA contributed three things: NF4, double quantisation, and paged optimizers. The first two are ideas about representation and are the interesting ones. The third is an engineering mitigation, covered in Lesson 16.

NF4, or NormalFloat

With 4 bits you get 16 levels. Where should you put them?

Uniform spacing, as in plain int4, places them at equal intervals across the range. That is optimal if the values are uniformly distributed. Neural network weights are not uniformly distributed. They are approximately zero-centred and normally distributed, so most values cluster near zero and few are near the extremes.

Uniform levels therefore waste resolution: many levels sit out in the tails where almost no weights live, while the dense region near zero gets too few.

NF4's answer: place the 16 levels at the quantiles of a standard normal distribution, so each level covers an equal amount of probability mass rather than an equal amount of range. Levels are dense near zero and sparse in the tails, matching where the weights actually are. The paper describes this as information-theoretically optimal for zero-mean normally distributed data.

For this to work, each block must first be normalised into [−1, 1] by dividing by its absmax, which is exactly what blockwise quantisation was already storing. So NF4 is not extra machinery; it is a better choice of the 16 levels within the scheme from Lesson 14.

int4:  levels evenly spaced        → resolution wasted in empty tails
nf4:   levels at normal quantiles  → resolution where the weights are

The consequence is that NF4 beats int4 at the same bit-width, for free. No extra storage, just a better codebook.

Two honest limits:

  • The gain depends on the weights actually being close to normal. They usually are, per-block, after pretraining. It is an empirical property, not a guarantee.
  • NF4 is a weight format. It is not the right tool for activations, which have different and less well-behaved distributions.

Double quantisation

Return to the overhead from Lesson 14. Per parameter, at block size 64 with fp32 scales:

weights            0.5     bytes
scales   4/64  =   0.0625  bytes
                  ────────
                   0.5625  bytes  (4.5 bits)

Those scales are themselves just numbers in a tensor, and there are a lot of them. So quantise them too.

Double quantisation takes the fp32 block scales, groups them (the paper uses blocks of 256), and stores them at 8 bits with their own second-level scale. The result:

first-level scales at 8 bits:   1/64      = 0.015625 bytes/param
second-level scales:            4/(64·256) ≈ 0.000244 bytes/param
weights:                                     0.5      bytes/param
                                            ──────────
                                             0.5159   bytes/param  (~4.13 bits)

The paper reports the saving as approximately 0.37 bits per parameter on average, which for a 65B model is around 3 GB. That is a real amount of memory recovered from pure bookkeeping.

Three bars of bits per parameter. The four-bit weight segment is identical in all three. Alone it is 4 bits; with fp32 block scales a half-bit segment is added, reaching 4.5; with double quantisation that segment shrinks to about an eighth of a bit, reaching about 4.13.

Note the shape of the idea: the quantisation constants were a fixed tax on the scheme, and double quantisation reduced the tax rather than the payload.

The bars say the same thing without the arithmetic. The grey segment is the same length three times over, because nothing about the weights changed; every difference between 4, 4.5 and 4.13 bits is the small accented piece on the end. A scheme that had reduced the payload instead would have moved the grey. It is a small, clever, entirely mechanical win, and it is why 4-bit QLoRA lands near 4.1 bits per parameter in practice rather than 4.5.

The memory result

Recompute Lesson 6's example with a 4-bit base. A 7B model with a 20M-parameter adapter:

bf16 base 4-bit base
Frozen base weights 14.0 GB ~3.6 GB
Adapter state (16 B/param) 0.32 GB 0.32 GB
Subtotal 14.3 GB 3.9 GB

Against 112 GB for full fine-tuning, the 4-bit adapter run's fixed cost is under 4 GB. Activations (Lesson 7) are now decisively the dominant term, which means gradient checkpointing matters more here than anywhere else in the workspace.

This is the result that made QLoRA significant: it moved fine-tuning of large models from clusters to single devices.

The configuration

quant_config := {
    load_in_4bit: true,
    quant_type: "nf4",            # not "fp4"
    use_double_quant: true,       # the 0.37 bits/param saving
    compute_dtype: "bfloat16",    # storage is 4-bit; compute is not
}

Four settings, each mapping to something in this lesson. bnb_4bit_compute_dtype is the one people leave at its default and should not: it is the precision the dequantised blocks are multiplied in, and leaving it at fp32 discards much of the speed benefit.

As always, confirm these names against the installed version's documentation.

Practice

  1. ▢ Why does NF4 beat int4 at the same bit-width?
Check

It places its 16 levels at the quantiles of a normal distribution rather than at uniform intervals, so resolution is concentrated where weights actually cluster, near zero, instead of being spent on empty tails.

Same storage, better codebook. The precondition is normalising each block to [−1, 1] first, which blockwise quantisation already does.

  1. ▢ What exactly does double quantisation compress, and roughly how much does it save?
Check

The first-level block scales, meaning the quantisation constants rather than the weights. They are stored at 8 bits in groups of 256, with a small second-level scale.

The paper reports about 0.37 bits per parameter on average, roughly 3 GB for a 65B model. It takes 4-bit-with-overhead from about 4.5 bits down to about 4.1.

  1. ▢ Compute total fixed memory for a 13B base in NF4 with double quantisation, plus a 30M adapter.
Check

Base: 13e9 × 0.5159 ≈ 6.7 GB. Adapter: 3e7 × 16 = 0.48 GB. About 7.2 GB fixed.

Compare 26 GB for a bf16 base plus adapter, and 208 GB for full fine-tuning. Activations then sit on top of the 7.2 GB.

  1. ▢ Which QLoRA component saves the most memory?

    • a) NF4 rather than plain uniform int4 quantisation
    • b) Storing the frozen base at 4 bits rather than bf16
    • c) Double quantisation of the first-level block scales
    • d) Paged optimizers spilling state during memory spikes
Check

b) Storing the frozen base at 4 bits rather than bf16, a 4× reduction on the largest fixed term.

NF4 improves quality at identical storage. Double quantisation saves about 0.37 bits per parameter. Paged optimizers prevent crashes at peaks rather than lowering steady-state use.

  1. ▢ Why is NF4 inappropriate for quantising activations?
Check

Its level placement is derived from an assumption that values are zero-mean and approximately normal. Post-pretraining weights fit that reasonably; activations do not, because they carry systematic large-magnitude outlier features, so normal quantiles are the wrong codebook for them.

  1. ▢ You set load_in_4bit=True and leave bnb_4bit_compute_dtype at its default. What have you likely given up?
Check

Speed. Storage is still 4-bit and memory is still saved, but if the compute dtype defaults to fp32 then every dequantised block is multiplied at full precision, forfeiting much of the throughput advantage.

Set it to bf16 deliberately. Storage precision and compute precision are separate decisions.

Real-world reps

  • [ ] Work out on paper the bytes per parameter for 4-bit with and without double quantisation, at block size 64. Confirm the difference is close to 0.37 bits.
  • [ ] Load the same model in bf16 and in NF4, measuring memory both times. Compare against your prediction.
  • [ ] Tomorrow: generate the same prompt from both, greedily, and read the outputs side by side. Form a first impression of the quality cost; Lesson 17 makes it a measurement.

Going further


Not landing? Reread the primary source at the top, since this lesson compresses it and compression is where understanding leaks. Check the glossary for any term that felt slippery.

If the lesson itself is unclear rather than the material, that is a defect: open an issue.

Table of contents