Skip to content
teach

Learning: Adapter Fine-Tuning (LoRA, QLoRA, DoRA)

Be the person who can take a base model and a task, decide whether adapter fine-tuning is the right answer at all, run it, prove it worked with an eval that would catch a regression, and ship it, and who can argue convincingly for prompting or retrieval instead when those would do the job better.

Start here: 0001. What a Base Model Actually Is
Latest lesson: 0027. When Not to Fine-Tune

Success looks like

  • Derive the memory cost of full fine-tuning versus LoRA for a given model and account for every byte.
  • Choose rank, alpha, target modules and learning rate for a new task and justify each from the task rather than from a blog post's defaults.
  • Train LoRA, QLoRA and DoRA adapters on the same task and explain the measured differences, not the advertised ones.
  • Build an eval that catches a regression before shipping, instead of reading a loss curve and hoping.
  • Serve an adapter through an inference stack and measure what it cost in quality and latency.
  • Diagnose a failed run and name the cause: overfitting, catastrophic forgetting, wrong chat template, tokenizer mismatch, wrong target modules.
  • Read a new PEFT paper and judge whether its claimed win would survive on your own task.

Constraints

  • Assumes no prior machine-learning background. Matrix multiplication is the only mathematics taken for granted; everything above it is taught where it is needed.
  • Adapter training on a base model in the 1B to 3B class fits roughly 16 GB of accelerator or unified memory, which puts the whole arc within reach of ordinary single-device hardware.
  • Backend support is uneven, not exclusive. CUDA remains the most complete and best-optimised path and is where published numbers reproduce most reliably; other backends work with varying kernel coverage and slower steps. Every concept transfers unchanged; only speed and kernel availability differ.
  • Reps are long-latency: a training run takes minutes to hours, so sessions batch rather than fitting an evening, with reading in between.
  • The tooling moves faster than any book. Library APIs are read from the installed version's documentation, never recalled.
  • Nothing in the arc requires paid infrastructure, though renting a larger accelerator will make some comparisons faster.

Out of scope

  • Full fine-tuning at scale, and pretraining.
  • RLHF, DPO and preference optimisation: adjacent, and a separate workspace later.
  • Inference quantisation and distillation as topics in their own right; touched only where QLoRA needs them.
  • Building serving infrastructure, beyond running an adapter through an existing stack.
  • Training a model from scratch.

The arc

Seven stages, zero to senior. A stage takes several lessons and the boundaries are soft; what makes a stage done is the capability, not the lesson count.

Stage Lessons Covers Done when
1. Ground floor 0001 to 0004 What a forward pass computes, where the weight matrices live, tokenizers, chat templates, what training changes Can point at which tensors an adapter would touch and why
2. The memory argument 0005 to 0007 Parameters vs gradients vs optimizer state, AdamW's two moments, activation memory, gradient checkpointing Can compute why full fine-tuning fails on a given device, in bytes
3. LoRA 0008 to 0013 ΔW = BA, rank, alpha and scaling, initialisation, target module choice, running it, reading the run, merging Trains an adapter and can defend every hyperparameter
4. Quantisation and QLoRA 0014 to 0017 int8/int4, NF4, double quantisation, backprop through a frozen quantized base, what degrades Ran real NF4 and can say what it cost in quality
5. DoRA and the variant landscape 0018 to 0020 Magnitude/direction decomposition, why it helps most at low rank, how to assess a new variant's claim Can predict when DoRA beats LoRA before running it
6. Data and evaluation 0021 to 0024 Dataset construction, contamination, held-out design, task metrics vs loss, regression suites Has an eval that has actually caught a regression
7. Operate and judge 0025 to 0027 Serving adapters vs merged, multi-adapter serving, cost and latency, when not to fine-tune Trusted to decide whether a task should be fine-tuned at all

Lessons

Work through these in order.

# Lesson Teaches
0001 What a Base Model Actually Is The model is one next-token function; base is not instruct
0002 Where the Weights Live Naming every projection an adapter could attach to
0003 Tokenizers and Chat Templates Train on the rendering you will serve
0004 What Training Actually Changes The four operations in one training step
0005 Counting Parameters and Bytes Turning a model config into gigabytes
0006 Gradients and Optimizer State Sixteen bytes per trainable parameter, two per frozen
0007 Activations, Batch Size and Checkpointing Why an adapter run still runs out of memory
0008 The Low-Rank Idea ΔW = BA, and counting an adapter
0009 Rank, Alpha and Initialisation Only α/r matters, and why B starts at zero
0010 Choosing Target Modules Attention-only is a 2021 ablation, not a default
0011 Your First Adapter A run that proves the pipeline before it proves anything else
0012 Reading a Training Run Diagnosing by curve shape, and what loss cannot see
0013 Merging, Saving and Shipping an Adapter Merged is exact; an adapter is a diff that needs its base
0014 Quantisation from First Principles Blockwise scales, and why one outlier ruins a tensor
0015 NF4 and Double Quantisation The two ideas QLoRA actually contributed
0016 Training Through a Quantized Base How gradients flow through frozen 4-bit weights
0017 What QLoRA Actually Costs Separating base degradation from training degradation
0018 DoRA: Magnitude and Direction Renormalisation decouples the two, and that is the method
0019 When DoRA Wins, and When It Doesn't Predicting the win before running it
0020 Judging a New PEFT Variant Six triage questions, and the baseline to beat
0021 Building the Dataset Data beats every hyperparameter in this workspace
0022 Contamination and Held-Out Design Splitting by the right key, and sizing for the question
0023 Metrics That Mean Something Loss selects checkpoints; task metrics make decisions
0024 The Regression Suite Catching the damage every other number hides
0025 Serving Adapters Merge for one task, route for many
0026 Cost, Latency and Throughput Prefill versus decode, and where the money actually goes
0027 When Not to Fine-Tune Fine-tuning is sixth on the list, and why that matters

Reference

How this works

Each lesson is short and self-contained. Answer keys are collapsed: recall first, then open them. The real-world reps matter more than the reading, and spacing them out is the point. Anything still unclear at the end of a lesson is worth chasing to its primary source before moving on.

Table of contents