Failure Modes
Diagnostic sheet. Keep it open during a run.
The four checks, before anything else
# 1. Did the adapter attach?
model.print_trainable_parameters() # compare to your arithmetic
# 2. Is the data what you think?
print(repr(tokenizer.decode(ds[0]["input_ids"])))
print(sum(1 for l in ds[0]["labels"] if l != -100), "of", len(ds[0]["labels"]), "scored")
# 3. Does loss move? # logging_steps=1 over 50 steps
# 4. Can it overfit two examples? # loss → ~0 in 100 steps, or the pipeline is broken
Check 2 catches template mismatch, missing end-of-turn tokens, masking bugs and truncation in one read. It is the highest-value ten seconds in the process.
Symptom → cause
The first eleven rows below name a curve rather than an event, and several of them share a shape: three say flat, two say barely moves. Match the sketch, then read every row that names it, since one shape can have more than one cause. From row twelve the symptom was observed somewhere else, in an output, a load error or a regression suite, and there is no curve to match.
| Symptom | Likely cause | Check |
|---|---|---|
| Loss exactly flat from step 1 | Adapter matched nothing | print_trainable_parameters() |
| Loss exactly flat from step 1 | All labels -100 | Count scored positions |
| Loss exactly flat from step 1 | LR orders of magnitude too low | Should be ~1e-4 |
| Loss barely moves | Target module names wrong for this architecture | named_modules() |
| Loss barely moves | α/r accidentally tiny | Raised r, forgot α |
| Loss plateaus high | Capacity ceiling | More rank or more targets, not more epochs |
| Held-out rises, train falls | Overfitting | Ship the held-out minimum |
| Train → 0 fast, held-out rises at once | Dataset far too small for capacity | More data, less capacity |
| Loss staircases at epoch boundaries | Memorisation | Fewer epochs |
| Loss spikes then recovers | One pathological batch | Find the outlier example |
Loss → NaN | Overflow, or bad data | Prefer bf16; lower LR; add warm-up |
| Held-out suspiciously good | Contamination across the split | Dedup train vs test, read top-similarity pairs |
| Model never stops generating | No end-of-turn token in training data | Decode an example |
| Output ignores requested format | Template mismatch train vs serve | Compare with repr |
| Answers cut off mid-sentence | max_length truncating targets | Check the length distribution |
| Answers to questions not asked | Packing broke loss masking or boundaries | Verify -100 survives packing |
| Worse at unrelated things | Catastrophic forgetting | Regression suite vs stored baseline |
| No longer refuses what base refused | Safety regression | Blocks the ship |
| Merged ≠ unmerged output | Merge bug, or sampling was on | Re-verify greedily |
| Merged model degraded | Merged into a quantized base | Merge into full precision, then quantise |
| Adapter will not load | Wrong base model or revision | An adapter is a diff; it needs its exact base |
| Gradient errors on a 4-bit base | Skipped prepare_model_for_kbit_training | Add it |
| Great eval, bad production | Train/serve distribution mismatch | Train on real inputs, not clean ones |
| Result vanished on rerun | Seed noise, never a real effect | Multiple seeds, paired comparison |
Silent failures: no error, wrong result
The dangerous set. Nothing raises, everything looks fine.
- Template mismatch between training and serving.
- Missing end-of-turn token. Trains fine, never stops.
- Loss masking wrong. Training on prompts unintentionally.
- Contamination. Held-out measures memorisation, reports generalisation.
- Group or temporal leakage. Split by the wrong key.
- Catastrophic forgetting. Every number you watch improves.
- Adapter attached to nothing. Trains, logs a loss, learns nothing.
- Evaluating a different artifact than you ship, for example bf16 eval and 4-bit deploy.
Each has a specific detection step above. None announces itself.
Reducing forgetting
Cheapest first:
- Fewer steps, or an earlier checkpoint
- Lower learning rate
- Lower rank, or fewer target modules
- Mix a few percent general instruction data into training
- Keep the adapter unmerged and route, so the base stays intact and forgetting stops mattering
Comparison hygiene
- One variable at a time; fix seed, data and split
- Multiple seeds, since variance often exceeds the effect
- Paired comparison on identical examples
- Report
nand an interval, never a bare number - Read actual generations, including failures, every time
What to record per run
Base model and revision · adapter config verbatim · dataset and exact split · seed · effective batch size · LR and schedule · steps · chosen checkpoint and the number justifying it · library versions · eval and regression results.