LLM Pretraining Resources
Knowledge
- Paper: "Scaling Laws for Neural Language Models", Kaplan et al., 2020
Fits the original power-law relationships between loss, compute, parameter count, and data, and argues for scaling parameters faster than data. Use for: stage 3, as the baseline that Chinchilla later revises. - Paper: "Training Compute-Optimal Large Language Models", Hoffmann et al., 2022
The Chinchilla paper. Refits the scaling laws with a wider sweep and finds parameters and tokens should scale together, roughly twenty tokens per parameter, overturning Kaplan's parameter-heavy recommendation. Use for: stage 3's central result, and the token budget most modern training runs actually use. - Paper: "ZeRO: Memory Optimizations Toward Training Trillion Parameter Models", Rajbhandari et al., 2019
Defines the three ZeRO stages, sharding optimizer state, then gradients, then parameters across data-parallel devices. Use for: stage 5's central technique, which FSDP implements. - Documentation: "Fully Sharded Data Parallel", PyTorch
The API reference for PyTorch's ZeRO-stage-3-style sharding, including the sharding strategies and wrapping policies a lesson's lab actually runs. Use for: stage 5's hands-on reps. - Paper: "Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM", Narayanan et al., 2021
Combines tensor parallelism, pipeline parallelism, and data parallelism in one system, with the throughput and memory tradeoffs of each measured against the others. Use for: stage 6's central technique and its worked numbers. - Documentation: "DeepSpeed", Microsoft
A second production implementation of ZeRO alongside 3D parallelism, with configuration examples a lesson can point to as a second concrete system next to FSDP and Megatron-LM. Use for: stages 5 and 6, as the cross-check that a claim is about the technique rather than one library's API. - Paper: "The Pile: An 800GB Dataset of Diverse Text for Language Modeling", Gao et al., 2020
A fully documented pretraining corpus: what sources went in, at what proportions, and why. Use for: stage 1, as a worked example of a data-mixing decision made explicit. - Paper: "Deduplicating Training Data Makes Language Models Better", Lee et al., 2021
Shows near-duplicate text and long repeated substrings cause verbatim memorization, and gives two deduplication techniques (exact substring matching via a suffix array, and near-duplicate matching via MinHash), each measured against a real corpus. Use for: stage 1's deduplication material, with concrete before/after numbers. - Paper: "Language Models are Few-Shot Learners", Brown et al., 2020
The GPT-3 paper. Appendix A documents a concrete quality-filtering pipeline: a classifier trained to distinguish curated text from raw Common Crawl, used to re-sample it toward higher-scoring documents, alongside a separate fuzzy-deduplication pass. Use for: stage 1's quality-filtering material, as the worked example the deduplication paper above does not itself provide. - Paper: "Neural Machine Translation of Rare Words with Subword Units", Sennrich et al., 2015
Introduces byte-pair encoding for subword tokenization, the algorithm most modern tokenizers still build on. Use for: stage 2's core algorithm. - Paper: "SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing", Kudo and Richardson, 2018
Treats tokenization as operating on raw text directly, without a language-specific pre-tokenization step, and covers unigram-language-model tokenization as an alternative to BPE. Use for: stage 2's second algorithm and the vocabulary-size tradeoff. - Paper: "Scaling Language Models: Methods, Analysis & Insights from Training Gopher", Rae et al., 2021
A detailed, honest account of a large training run, including where loss spikes and instability appeared and what the authors did about them. Use for: stage 7's central case study. - Paper: "LLaMA: Open and Efficient Foundation Language Models", Touvron et al., 2023
A published, reproducible training recipe: data mixture, tokenizer, schedule, and hyperparameters, stated plainly enough to check a lesson's claims against. Use for: cross-checking stages 1 through 3 and 7 against a real, complete recipe. - Paper: "OPT: Open Pre-trained Transformer Language Models", Zhang et al., 2022
Released alongside a public logbook of the infrastructure problems the authors actually hit training a 175B model: hardware failures, manual and automatic restarts, loss divergences, and the concrete recovery procedure used for each. Use for: stage 8's central case study, and a second real example of the loss-divergence pattern stage 7 covers. - Paper: "Don't Stop Pretraining: Adapt Language Models to Domains and Tasks", Gururangan et al., 2020
Shows that continuing pretraining on an existing broad-coverage model (domain-adaptive pretraining) consistently improves target-domain performance across four domains and eight tasks, in both high- and low-resource settings, without training a new model from scratch. Use for: stage 10's central case that pretraining from scratch is usually the wrong default.
Gaps
- OPT's paper documents real failure frequency and recovery procedure but does not state a specific checkpoint-interval policy (how often checkpoints were actually saved). Stage 8's checkpoint-frequency tradeoff is taught from general engineering reasoning rather than a stated policy; a primary source with a concrete interval would strengthen it.
- GPT-3's classifier-based filtering (2020) is the only worked quality-filtering example listed so far. Corpora built since 2023 (FineWeb, Dolma) filter far more aggressively and document it in more depth; a more recent primary source would strengthen stage 1.