Learning: Evals
Be able to build an eval that catches a regression a vibe check would miss, and use it to make a go/no-go call on a model change (a fine-tune, a prompt change, a RAG change) that you can defend with a number instead of a feeling.
Latest lesson: 19. Reviewing an Eval Report
Success looks like
- Design a held-out eval set for a given model change that resists contamination and would catch a regression a casual read-through would miss.
- Use that eval to make and defend a go/no-go call on shipping the change, naming the number and why it is trustworthy.
- Compare an LLM-as-judge approach against task-specific metrics for a given case, and choose between them on the merits rather than by default.
- Review someone else's eval report or harness and name specifically what would make its number untrustworthy (an unheld-out set, an uncalibrated judge, a single run with no variance estimate, a leaderboard rank standing in for a task-specific call), rather than saying it feels thin.
Constraints
- No domain restriction: covers evaluating fine-tunes, prompt changes, and RAG changes alike, since the same eval discipline applies to all three.
- No fixed stack or prior-experience assumption.
Out of scope
- Judging a variant, held-out design, metrics and the regression suite as taught inside one fine-tuning arc: see
llm/finetuning(lessons 0020-0024), subordinated to fine-tuning specifically. This workspace owns evaluation as the subject and links back rather than restating.
The arc
Thirteen stages, from a trustworthy eval to a defended go/no-go call that accounts for humans, safety, agents, RAG, production, cost, and the pull of a public leaderboard, finishing with the judgment to review someone else's eval and maintain one honestly as shared infrastructure. A stage takes several lessons and the boundaries are soft; what makes a stage done is the capability, not the lesson count.
| Stage | Lessons | Covers | Done when |
|---|---|---|---|
| 1. Held-out data and contamination | 0001 to 0002 | Train/test split discipline, contamination detection and its sources | Can design a held-out set that resists contamination |
| 2. Task-specific metrics | 0003 to 0004 | Exact match, F1, BLEU/ROUGE, code-execution metrics, choosing per task | Can pick and justify a metric for a stated task |
| 3. LLM-as-judge | 0005 to 0006 | Judge-prompt design, calibration, position and verbosity bias, agreement with humans | Can build a judge prompt and name its failure modes |
| 4. Building the harness | 0007 to 0008 | Eval frameworks, reproducible runs, turning a design into a number | Has a runnable eval harness that produces a trustworthy number |
| 5. The go/no-go call | 0009 to 0010 | Statistical significance vs noise, regression thresholds, communicating the decision | Can defend a ship/no-ship call with a number and say why it's trustworthy |
| 6. Human evaluation as a method | 0011 | Annotation guidelines, sampling for review, Cohen's kappa, and what a human label costs | Can measure inter-rater agreement precisely and diagnose a low value |
| 7. Safety and red-teaming evaluation | 0012 | Red-teaming, jailbreaks and attack success rate, over-refusal and XSTest | Can report both a refusal-rate and an over-refusal metric, and explain why neither alone is a safety eval |
| 8. Evaluating agents and tool use | 0013 | Trajectories, task success vs. step accuracy, pass^k reliability | Can explain why task success beats step matching, and why single-trial success isn't repeated-trial reliability |
| 9. RAG-specific evaluation | 0014 | Faithfulness/groundedness, citation correctness, separating a retrieval failure from a generation failure | Can measure faithfulness and citation correctness independently, and diagnose which side of a RAG pipeline actually failed |
| 10. Online evaluation | 0015 to 0016 | A/B tests, OEC and guardrail metrics, production telemetry, drift, training-serving skew | Can choose a defensible OEC with a guardrail metric, and correctly attribute an offline/online gap to variance, drift, or an engineering error |
| 11. Cost and latency as eval dimensions | 0017 | Tail latency vs. average, per-request/token cost, folding both into the go/no-go call | Can defend a go/no-go call that weighs quality, tail latency, and cost together, not quality alone |
| 12. Public benchmarks | 0018 | Static-benchmark contamination/saturation, arena-style selective disclosure, style bias at leaderboard scale | Can explain why a leaderboard rank fails the OEC test and isn't a substitute for a task-specific go/no-go call |
| 13. Judgment | 0019 | Reviewing eval reports, settling disputed claims from primary sources, maintaining an eval suite others depend on | Trusted to review someone else's eval and defend its number or name what makes it untrustworthy, and to maintain an eval suite as shared infrastructure without breaking others' gates |
Lessons
Work through these in order.
| # | Lesson | Teaches |
|---|---|---|
| 0001 | Held-out Data and Contamination | Why an eval number is only as trustworthy as what the model never saw |
| 0002 | Designing Contamination Resistance | Preventing contamination in a custom eval set from the start, instead of only detecting it after the fact |
| 0003 | Task-Specific Metrics | Exact match, token-level F1, and BLEU/ROUGE, and the failure mode each one has |
| 0004 | Code-Execution Metrics and Choosing a Metric | Functional correctness, the pass@k estimator, and a decision principle for picking a metric per task |
| 0005 | Judge-Prompt Design and Calibration | How to write a judge prompt that grades consistently, and what it means for a judge to be calibrated before trusting it |
| 0006 | Judge Bias and Human Agreement | Position bias and verbosity bias in LLM-as-judge, and how to measure a judge's agreement with human raters |
| 0007 | Eval Frameworks and Building a Harness | The four stages every eval harness has, and choosing between building custom eval-as-code and reusing a standardized benchmark harness |
| 0008 | Reproducible Runs | Why the same eval can produce different numbers on different runs, and what a trustworthy result has to log alongside the score |
| 0009 | Statistical Significance vs Noise | Why a score difference has to be measured against the noise that could produce it by chance, and why paired comparison beats treating two scores as independent |
| 0010 | Defending the Go/No-Go Call | How to set a regression threshold honestly, and everything a complete go/no-go defense has to cite |
| 0011 | Human Evaluation as a Method | Lesson 6 measured a judge against human raters without ever teaching how a human rating is actually produced, and the raw agreement number it leaned on turns out to need its own correction for chance before it means anything |
| 0012 | Safety and Red-Teaming Evaluation | A safety eval that only measures how often a model refuses genuinely unsafe prompts is measuring half a trade-off, since the same tuning that raises that number can just as easily raise how often the model refuses prompts that were never unsafe at all |
| 0013 | Evaluating Agents and Tool Use | An agent task has many valid paths to the same correct outcome, so scoring it against one reference sequence of actions repeats exact match's mistake at trajectory scale, and even a single trial's success rate hides how often the same agent fails the same task on a second try |
| 0014 | RAG-Specific Evaluation | llm/rag drew the line between a retrieval failure and a generation failure but left the generation-side check undefined, and that check turns out to need two separate numbers, not one, since a faithful claim and a correctly cited claim can fail independently of each other |
| 0015 | A/B Tests and Guardrail Metrics | Every stage so far has been offline, but shipping to real users needs its own decision metric, and that metric alone is exactly as dangerous as the single safety numbers earlier stages already warned against trusting in isolation |
| 0016 | Production Telemetry and Drift | An offline number and a live number disagreeing isn't one problem, it's three different ones, and treating an engineering bug as drift, or drift as an engineering bug, sends the fix to the wrong team entirely |
| 0017 | Cost and Latency as Eval Dimensions | A quality win measured in isolation from what it costs to serve is half a go/no-go call, and an average latency number hides exactly the tail that determines whether users actually experience the system as fast |
| 0018 | Public Benchmarks | A leaderboard position is a movable metric optimized by people who aren't you, for users who aren't your users, scored by a process with its own documented gaming vectors, which makes it exactly the kind of metric lesson 15 warned against trusting as a go/no-go signal |
| 0019 | Reviewing an Eval Report | Name specifically what would make an eval report's number untrustworthy (contamination, an uncalibrated judge, a single run with no variance estimate, a leaderboard rank standing in for a task-specific call), rather than saying it feels thin |
Reference
- Glossary: canonical terms for this topic
- Resources: trusted sources
- Held-Out Data and Contamination: the two contamination pathways, the guided-instruction test, and the prevention toolkit for a custom eval set
- Metrics: exact match, F1, BLEU/ROUGE and pass@k side by side, with the pass@k formula and the decision principle for choosing among them
- LLM-as-Judge: the judge-prompt design checklist, calibration and score compression, position and verbosity bias, and the human-vs-human agreement bar
- Eval Harness: the four-stage pipeline, openai/evals vs lm-evaluation-harness, and what a defensible number logs alongside the score
- Go/No-Go: the standard-error formula, paired comparison, setting a threshold honestly, and everything a complete defense cites
How this works
Each lesson is short and self-contained. Answer keys are collapsed: recall first, then open them. The real-world reps matter more than the reading, and spacing them out is the point. Anything still unclear at the end of a lesson is worth chasing to its primary source before moving on.