Skip to content
teach

Memory Budget

Byte accounting for adapter fine-tuning. Built for lookup during a run.

Bytes per parameter

Format Bytes Bits
fp32 4 32
bf16 / fp16 2 16
int8 1 8
4-bit, raw 0.5 4
4-bit + fp32 scales, block 64 0.5625 4.5
4-bit + double quant ~0.516 ~4.13

Reflex: bf16 gigabytes ≈ 2 × parameters in billions.

Per-parameter training cost

Category fp32 Mixed precision
Weights 4 2 bf16 + 4 fp32 master
Gradients 4 2
AdamW moment 1 4 4
AdamW moment 2 4 4
Trainable total 16 16
Frozen total 4 2

Mixed precision does not reduce the total. It buys arithmetic throughput.

Optimizer variant Bytes/param
AdamW, fp32 moments 16
AdamW, 8-bit moments 10
SGD with momentum 12

Parameter counts

attention/layer = 2·d² + 2·d·d_kv          # q,o full; k,v narrowed by GQA
mlp/layer       = 3·d·d_ff                  # gate, up, down
model           ≈ n_layers·(attn + mlp) + vocab·d·(1 if tied else 2)

The MLP is typically 75 to 80% of a block.

Adapter parameter count

adapter = Σ over targeted layers:  r · (d_in + d_out)

Dominated by the larger dimension, so GQA narrowing barely shrinks an adapter.

Base Targets Rank Adapter % of base
~1.1B, d=2048 attention 8 2.5M 0.22%
~7B, d=4096 q,v 16 6.8M 0.10%
~7B, d=4096 all-linear 64 ~160M ~2.3%

Activations

activations ≈ n_layers · batch · seq_len · d · ~20 bytes

The constant is a starting point to measure against, not a law. Linear in batch and sequence length; the attention-score quadratic term is removed by a memory-efficient attention kernel.

Config Estimate
24 layers, d=2048, batch 4, seq 2048 ~8 GB
32 layers, d=4096, batch 2, seq 4096 ~21 GB

Gradient checkpointing cuts this by roughly an order of magnitude for ~20 to 40% more step time.

In an adapter run, activations usually dominate. This inverts the full fine-tuning profile.

Two stacked bars on one scale in gigabytes. A full fine-tune of 7B with AdamW is 112 GB fixed plus 21 GB of activations, so activations are 16% of it. A QLoRA run on an NF4 7B base with a 20M adapter is 3.9 GB fixed plus the same 21 GB, so activations are 84% of it.

Both bars take the fixed memory from the table below and the activations from the 32-layer, d=4096, batch 2, seq 4096 row above, with no gradient checkpointing. The activation block is the same size in both, because it depends on the shape of the forward pass rather than on how many parameters are trainable. Only the memory around it collapses, which is why the thing to cut is rarely the adapter.

Worked totals, excluding activations

Setup Fixed memory
Full FT, 1.5B, AdamW 24 GB
Full FT, 7B, AdamW 112 GB
Full FT, 70B, AdamW 1120 GB
bf16 7B base + 20M adapter 14.3 GB
NF4 7B base + 20M adapter ~3.9 GB
NF4 13B base + 30M adapter ~7.2 GB

When a run does not fit

In order of what you give up:

  1. Gradient checkpointing, costs time only
  2. Lower batch size, raise gradient accumulation, statistically equivalent
  3. Shorter sequence length, but check the token-length distribution first; this changes the task
  4. Quantise the base, costs some quality
  5. Lower rank, saves the least and costs capacity

Rank is the instinct and nearly the worst option. Halving rank on a 4M adapter saves ~32 MB.

Effective batch size

effective = per_device_batch × grad_accumulation × num_devices

Always report the effective number.

Table of contents