Memory Budget
Byte accounting for adapter fine-tuning. Built for lookup during a run.
Bytes per parameter
| Format | Bytes | Bits |
|---|---|---|
| fp32 | 4 | 32 |
| bf16 / fp16 | 2 | 16 |
| int8 | 1 | 8 |
| 4-bit, raw | 0.5 | 4 |
| 4-bit + fp32 scales, block 64 | 0.5625 | 4.5 |
| 4-bit + double quant | ~0.516 | ~4.13 |
Reflex: bf16 gigabytes ≈ 2 × parameters in billions.
Per-parameter training cost
| Category | fp32 | Mixed precision |
|---|---|---|
| Weights | 4 | 2 bf16 + 4 fp32 master |
| Gradients | 4 | 2 |
| AdamW moment 1 | 4 | 4 |
| AdamW moment 2 | 4 | 4 |
| Trainable total | 16 | 16 |
| Frozen total | 4 | 2 |
Mixed precision does not reduce the total. It buys arithmetic throughput.
| Optimizer variant | Bytes/param |
|---|---|
| AdamW, fp32 moments | 16 |
| AdamW, 8-bit moments | 10 |
| SGD with momentum | 12 |
Parameter counts
attention/layer = 2·d² + 2·d·d_kv # q,o full; k,v narrowed by GQA
mlp/layer = 3·d·d_ff # gate, up, down
model ≈ n_layers·(attn + mlp) + vocab·d·(1 if tied else 2)
The MLP is typically 75 to 80% of a block.
Adapter parameter count
adapter = Σ over targeted layers: r · (d_in + d_out)
Dominated by the larger dimension, so GQA narrowing barely shrinks an adapter.
| Base | Targets | Rank | Adapter | % of base |
|---|---|---|---|---|
| ~1.1B, d=2048 | attention | 8 | 2.5M | 0.22% |
| ~7B, d=4096 | q,v | 16 | 6.8M | 0.10% |
| ~7B, d=4096 | all-linear | 64 | ~160M | ~2.3% |
Activations
activations ≈ n_layers · batch · seq_len · d · ~20 bytes
The constant is a starting point to measure against, not a law. Linear in batch and sequence length; the attention-score quadratic term is removed by a memory-efficient attention kernel.
| Config | Estimate |
|---|---|
| 24 layers, d=2048, batch 4, seq 2048 | ~8 GB |
| 32 layers, d=4096, batch 2, seq 4096 | ~21 GB |
Gradient checkpointing cuts this by roughly an order of magnitude for ~20 to 40% more step time.
In an adapter run, activations usually dominate. This inverts the full fine-tuning profile.
Both bars take the fixed memory from the table below and the activations from the 32-layer, d=4096, batch 2, seq 4096 row above, with no gradient checkpointing. The activation block is the same size in both, because it depends on the shape of the forward pass rather than on how many parameters are trainable. Only the memory around it collapses, which is why the thing to cut is rarely the adapter.
Worked totals, excluding activations
| Setup | Fixed memory |
|---|---|
| Full FT, 1.5B, AdamW | 24 GB |
| Full FT, 7B, AdamW | 112 GB |
| Full FT, 70B, AdamW | 1120 GB |
| bf16 7B base + 20M adapter | 14.3 GB |
| NF4 7B base + 20M adapter | ~3.9 GB |
| NF4 13B base + 30M adapter | ~7.2 GB |
When a run does not fit
In order of what you give up:
- Gradient checkpointing, costs time only
- Lower batch size, raise gradient accumulation, statistically equivalent
- Shorter sequence length, but check the token-length distribution first; this changes the task
- Quantise the base, costs some quality
- Lower rank, saves the least and costs capacity
Rank is the instinct and nearly the worst option. Halving rank on a 4M adapter saves ~32 MB.
Effective batch size
effective = per_device_batch × grad_accumulation × num_devices
Always report the effective number.