Skip to content
teach

LLM Post-Training Resources

Knowledge

Gaps

  • No primary source yet on preference-data quality: annotator disagreement, labeling instructions, and how preference-data collection choices bias a trained reward model. Needed before stage 3's data-quality material can be fully grounded.
  • Alignment-tax measurement beyond InstructGPT's own report is thin; a more recent, model-agnostic treatment would strengthen stage 9.
  • No production account yet of a full DPO-versus-PPO-versus-GRPO cost comparison from a single team's infrastructure; stage 10's judgment lesson currently has to synthesize this from separate papers rather than one source that already made the comparison.
Table of contents