Learning: Adapter Fine-Tuning (LoRA, QLoRA, DoRA)
Be the person who can take a base model and a task, decide whether adapter fine-tuning is the right answer at all, run it, prove it worked with an eval that would catch a regression, and ship it, and who can argue convincingly for prompting or retrieval instead when those would do the job better.
Start here: 0001. What a Base Model Actually Is
Latest lesson: 0027. When Not to Fine-Tune
Success looks like
- Derive the memory cost of full fine-tuning versus LoRA for a given model and account for every byte.
- Choose rank, alpha, target modules and learning rate for a new task and justify each from the task rather than from a blog post's defaults.
- Train LoRA, QLoRA and DoRA adapters on the same task and explain the measured differences, not the advertised ones.
- Build an eval that catches a regression before shipping, instead of reading a loss curve and hoping.
- Serve an adapter through an inference stack and measure what it cost in quality and latency.
- Diagnose a failed run and name the cause: overfitting, catastrophic forgetting, wrong chat template, tokenizer mismatch, wrong target modules.
- Read a new PEFT paper and judge whether its claimed win would survive on your own task.
Constraints
- Assumes no prior machine-learning background. Matrix multiplication is the only mathematics taken for granted; everything above it is taught where it is needed.
- Adapter training on a base model in the 1B to 3B class fits roughly 16 GB of accelerator or unified memory, which puts the whole arc within reach of ordinary single-device hardware.
- Backend support is uneven, not exclusive. CUDA remains the most complete and best-optimised path and is where published numbers reproduce most reliably; other backends work with varying kernel coverage and slower steps. Every concept transfers unchanged; only speed and kernel availability differ.
- Reps are long-latency: a training run takes minutes to hours, so sessions batch rather than fitting an evening, with reading in between.
- The tooling moves faster than any book. Library APIs are read from the installed version's documentation, never recalled.
- Nothing in the arc requires paid infrastructure, though renting a larger accelerator will make some comparisons faster.
Out of scope
- Full fine-tuning at scale, and pretraining.
- RLHF, DPO and preference optimisation: adjacent, and a separate workspace later.
- Inference quantisation and distillation as topics in their own right; touched only where QLoRA needs them.
- Building serving infrastructure, beyond running an adapter through an existing stack.
- Training a model from scratch.
The arc
Seven stages, zero to senior. A stage takes several lessons and the boundaries are soft; what makes a stage done is the capability, not the lesson count.
| Stage | Lessons | Covers | Done when |
|---|---|---|---|
| 1. Ground floor | 0001 to 0004 | What a forward pass computes, where the weight matrices live, tokenizers, chat templates, what training changes | Can point at which tensors an adapter would touch and why |
| 2. The memory argument | 0005 to 0007 | Parameters vs gradients vs optimizer state, AdamW's two moments, activation memory, gradient checkpointing | Can compute why full fine-tuning fails on a given device, in bytes |
| 3. LoRA | 0008 to 0013 | ΔW = BA, rank, alpha and scaling, initialisation, target module choice, running it, reading the run, merging | Trains an adapter and can defend every hyperparameter |
| 4. Quantisation and QLoRA | 0014 to 0017 | int8/int4, NF4, double quantisation, backprop through a frozen quantized base, what degrades | Ran real NF4 and can say what it cost in quality |
| 5. DoRA and the variant landscape | 0018 to 0020 | Magnitude/direction decomposition, why it helps most at low rank, how to assess a new variant's claim | Can predict when DoRA beats LoRA before running it |
| 6. Data and evaluation | 0021 to 0024 | Dataset construction, contamination, held-out design, task metrics vs loss, regression suites | Has an eval that has actually caught a regression |
| 7. Operate and judge | 0025 to 0027 | Serving adapters vs merged, multi-adapter serving, cost and latency, when not to fine-tune | Trusted to decide whether a task should be fine-tuned at all |
Lessons
Work through these in order.
| # | Lesson | Teaches |
|---|---|---|
| 0001 | What a Base Model Actually Is | The model is one next-token function; base is not instruct |
| 0002 | Where the Weights Live | Naming every projection an adapter could attach to |
| 0003 | Tokenizers and Chat Templates | Train on the rendering you will serve |
| 0004 | What Training Actually Changes | The four operations in one training step |
| 0005 | Counting Parameters and Bytes | Turning a model config into gigabytes |
| 0006 | Gradients and Optimizer State | Sixteen bytes per trainable parameter, two per frozen |
| 0007 | Activations, Batch Size and Checkpointing | Why an adapter run still runs out of memory |
| 0008 | The Low-Rank Idea | ΔW = BA, and counting an adapter |
| 0009 | Rank, Alpha and Initialisation | Only α/r matters, and why B starts at zero |
| 0010 | Choosing Target Modules | Attention-only is a 2021 ablation, not a default |
| 0011 | Your First Adapter | A run that proves the pipeline before it proves anything else |
| 0012 | Reading a Training Run | Diagnosing by curve shape, and what loss cannot see |
| 0013 | Merging, Saving and Shipping an Adapter | Merged is exact; an adapter is a diff that needs its base |
| 0014 | Quantisation from First Principles | Blockwise scales, and why one outlier ruins a tensor |
| 0015 | NF4 and Double Quantisation | The two ideas QLoRA actually contributed |
| 0016 | Training Through a Quantized Base | How gradients flow through frozen 4-bit weights |
| 0017 | What QLoRA Actually Costs | Separating base degradation from training degradation |
| 0018 | DoRA: Magnitude and Direction | Renormalisation decouples the two, and that is the method |
| 0019 | When DoRA Wins, and When It Doesn't | Predicting the win before running it |
| 0020 | Judging a New PEFT Variant | Six triage questions, and the baseline to beat |
| 0021 | Building the Dataset | Data beats every hyperparameter in this workspace |
| 0022 | Contamination and Held-Out Design | Splitting by the right key, and sizing for the question |
| 0023 | Metrics That Mean Something | Loss selects checkpoints; task metrics make decisions |
| 0024 | The Regression Suite | Catching the damage every other number hides |
| 0025 | Serving Adapters | Merge for one task, route for many |
| 0026 | Cost, Latency and Throughput | Prefill versus decode, and where the money actually goes |
| 0027 | When Not to Fine-Tune | Fine-tuning is sixth on the list, and why that matters |
Reference
- Glossary: canonical terms for this topic
- Resources: trusted sources and communities
- Memory budget: byte accounting, and what to try when a run will not fit
- LoRA hyperparameters: rank, alpha, targets, learning rate, variants
- Failure modes: symptom to cause, and the silent failures
How this works
Each lesson is short and self-contained. Answer keys are collapsed: recall first, then open them. The real-world reps matter more than the reading, and spacing them out is the point. Anything still unclear at the end of a lesson is worth chasing to its primary source before moving on.