Skip to content
teach

Adapter Fine-Tuning Resources

Knowledge

The three methods

Current practice

  • Docs: "LoRA Without Regret", Hugging Face TRL
    The case that all-linear targets with adequate rank close the gap to full fine-tuning. Use for: target-module and rank choice, and as the baseline any new variant must beat. This supersedes the attention-only convention inherited from the original paper.

  • Blog: "Practical Tips for Finetuning LLMs Using LoRA", Sebastian Raschka, Ahead of AI
    Ablations over rank, alpha and target modules, with the experiments shown. Use for: sanity-checking hyperparameter folklore against measurements.

  • Docs: PEFT, Hugging Face
    Reference implementation of all three methods and the configuration surface they expose. Use for: what a hyperparameter is called in practice, and which variants are established enough to be a flag. Version-sensitive, so check against the installed release rather than from memory.

  • Docs: SFTTrainer, Hugging Face TRL
    The standard supervised fine-tuning loop and its dataset formats. Use for: stage 3 onward. Version-sensitive; parameter names have moved between releases.

Foundations

Quantisation

Data and evaluation

Serving

The alternative

Wisdom (Communities)

  • r/LocalLLaMA
    The largest concentration of people fine-tuning on consumer hardware, across every backend. Use for: searching what already worked within a memory budget, which folklore is worth testing, and replication reports on new methods. Read the archive; posting is a rep, not a source.

  • Hugging Face Forums
    Searchable and archived, with library authors present. Use for: PEFT, TRL and transformers behaviour that the docs leave ambiguous.

  • EleutherAI Blog
    Write-ups from a group that reads its own field critically, including negative results. Use for: stage 5 and 7 judgment calls on whether a claimed result generalises. Their Discord is where this gets discussed first, and is deliberately not listed: nothing said there is retrievable later.

  • MLX Discussions, ml-explore/mlx
    Public, searchable, and answered by the maintainers. Use for: reading whether a given operation is supported on Apple Silicon yet, usually already asked.

Gaps

  • No source tracks which torch, PEFT, TRL and quantisation-library versions actually work together on a given Python version and platform. This blocks a first run more often than any concept does, and each environment has to be verified against release notes rather than assumed.
  • Hardware backend coverage for 4-bit operations changes release to release. No document is reliable here; the library source is the only authority, and any tutorial's claim about hardware requirements should be treated as expired.
  • No trusted source yet for evaluation design specific to small-model task adaptation. This is the weakest link in stage 6 and the hardest part of the mission. The general evaluation literature is aimed at benchmarking foundation models, not at proving a narrow adapter helped.
  • Most published LoRA hyperparameter advice targets 7B models and above. Little addresses whether it transfers to the 1B to 3B class, which is where most people actually start.
  • No independent replication surveyed for DoRA at the scale LoRA and QLoRA enjoy. Treat its reported gains as needing local measurement.
  • No source chosen for constrained decoding, which Lesson 27 recommends as the correct answer to schema-validity problems. It is named without a reference behind it.
Table of contents