Skip to content
teach

Learning: Evals

Be able to build an eval that catches a regression a vibe check would miss, and use it to make a go/no-go call on a model change (a fine-tune, a prompt change, a RAG change) that you can defend with a number instead of a feeling.

Latest lesson: 19. Reviewing an Eval Report

Success looks like

  • Design a held-out eval set for a given model change that resists contamination and would catch a regression a casual read-through would miss.
  • Use that eval to make and defend a go/no-go call on shipping the change, naming the number and why it is trustworthy.
  • Compare an LLM-as-judge approach against task-specific metrics for a given case, and choose between them on the merits rather than by default.
  • Review someone else's eval report or harness and name specifically what would make its number untrustworthy (an unheld-out set, an uncalibrated judge, a single run with no variance estimate, a leaderboard rank standing in for a task-specific call), rather than saying it feels thin.

Constraints

  • No domain restriction: covers evaluating fine-tunes, prompt changes, and RAG changes alike, since the same eval discipline applies to all three.
  • No fixed stack or prior-experience assumption.

Out of scope

  • Judging a variant, held-out design, metrics and the regression suite as taught inside one fine-tuning arc: see llm/finetuning (lessons 0020-0024), subordinated to fine-tuning specifically. This workspace owns evaluation as the subject and links back rather than restating.

The arc

Thirteen stages, from a trustworthy eval to a defended go/no-go call that accounts for humans, safety, agents, RAG, production, cost, and the pull of a public leaderboard, finishing with the judgment to review someone else's eval and maintain one honestly as shared infrastructure. A stage takes several lessons and the boundaries are soft; what makes a stage done is the capability, not the lesson count.

Stage Lessons Covers Done when
1. Held-out data and contamination 0001 to 0002 Train/test split discipline, contamination detection and its sources Can design a held-out set that resists contamination
2. Task-specific metrics 0003 to 0004 Exact match, F1, BLEU/ROUGE, code-execution metrics, choosing per task Can pick and justify a metric for a stated task
3. LLM-as-judge 0005 to 0006 Judge-prompt design, calibration, position and verbosity bias, agreement with humans Can build a judge prompt and name its failure modes
4. Building the harness 0007 to 0008 Eval frameworks, reproducible runs, turning a design into a number Has a runnable eval harness that produces a trustworthy number
5. The go/no-go call 0009 to 0010 Statistical significance vs noise, regression thresholds, communicating the decision Can defend a ship/no-ship call with a number and say why it's trustworthy
6. Human evaluation as a method 0011 Annotation guidelines, sampling for review, Cohen's kappa, and what a human label costs Can measure inter-rater agreement precisely and diagnose a low value
7. Safety and red-teaming evaluation 0012 Red-teaming, jailbreaks and attack success rate, over-refusal and XSTest Can report both a refusal-rate and an over-refusal metric, and explain why neither alone is a safety eval
8. Evaluating agents and tool use 0013 Trajectories, task success vs. step accuracy, pass^k reliability Can explain why task success beats step matching, and why single-trial success isn't repeated-trial reliability
9. RAG-specific evaluation 0014 Faithfulness/groundedness, citation correctness, separating a retrieval failure from a generation failure Can measure faithfulness and citation correctness independently, and diagnose which side of a RAG pipeline actually failed
10. Online evaluation 0015 to 0016 A/B tests, OEC and guardrail metrics, production telemetry, drift, training-serving skew Can choose a defensible OEC with a guardrail metric, and correctly attribute an offline/online gap to variance, drift, or an engineering error
11. Cost and latency as eval dimensions 0017 Tail latency vs. average, per-request/token cost, folding both into the go/no-go call Can defend a go/no-go call that weighs quality, tail latency, and cost together, not quality alone
12. Public benchmarks 0018 Static-benchmark contamination/saturation, arena-style selective disclosure, style bias at leaderboard scale Can explain why a leaderboard rank fails the OEC test and isn't a substitute for a task-specific go/no-go call
13. Judgment 0019 Reviewing eval reports, settling disputed claims from primary sources, maintaining an eval suite others depend on Trusted to review someone else's eval and defend its number or name what makes it untrustworthy, and to maintain an eval suite as shared infrastructure without breaking others' gates

Lessons

Work through these in order.

# Lesson Teaches
0001 Held-out Data and Contamination Why an eval number is only as trustworthy as what the model never saw
0002 Designing Contamination Resistance Preventing contamination in a custom eval set from the start, instead of only detecting it after the fact
0003 Task-Specific Metrics Exact match, token-level F1, and BLEU/ROUGE, and the failure mode each one has
0004 Code-Execution Metrics and Choosing a Metric Functional correctness, the pass@k estimator, and a decision principle for picking a metric per task
0005 Judge-Prompt Design and Calibration How to write a judge prompt that grades consistently, and what it means for a judge to be calibrated before trusting it
0006 Judge Bias and Human Agreement Position bias and verbosity bias in LLM-as-judge, and how to measure a judge's agreement with human raters
0007 Eval Frameworks and Building a Harness The four stages every eval harness has, and choosing between building custom eval-as-code and reusing a standardized benchmark harness
0008 Reproducible Runs Why the same eval can produce different numbers on different runs, and what a trustworthy result has to log alongside the score
0009 Statistical Significance vs Noise Why a score difference has to be measured against the noise that could produce it by chance, and why paired comparison beats treating two scores as independent
0010 Defending the Go/No-Go Call How to set a regression threshold honestly, and everything a complete go/no-go defense has to cite
0011 Human Evaluation as a Method Lesson 6 measured a judge against human raters without ever teaching how a human rating is actually produced, and the raw agreement number it leaned on turns out to need its own correction for chance before it means anything
0012 Safety and Red-Teaming Evaluation A safety eval that only measures how often a model refuses genuinely unsafe prompts is measuring half a trade-off, since the same tuning that raises that number can just as easily raise how often the model refuses prompts that were never unsafe at all
0013 Evaluating Agents and Tool Use An agent task has many valid paths to the same correct outcome, so scoring it against one reference sequence of actions repeats exact match's mistake at trajectory scale, and even a single trial's success rate hides how often the same agent fails the same task on a second try
0014 RAG-Specific Evaluation llm/rag drew the line between a retrieval failure and a generation failure but left the generation-side check undefined, and that check turns out to need two separate numbers, not one, since a faithful claim and a correctly cited claim can fail independently of each other
0015 A/B Tests and Guardrail Metrics Every stage so far has been offline, but shipping to real users needs its own decision metric, and that metric alone is exactly as dangerous as the single safety numbers earlier stages already warned against trusting in isolation
0016 Production Telemetry and Drift An offline number and a live number disagreeing isn't one problem, it's three different ones, and treating an engineering bug as drift, or drift as an engineering bug, sends the fix to the wrong team entirely
0017 Cost and Latency as Eval Dimensions A quality win measured in isolation from what it costs to serve is half a go/no-go call, and an average latency number hides exactly the tail that determines whether users actually experience the system as fast
0018 Public Benchmarks A leaderboard position is a movable metric optimized by people who aren't you, for users who aren't your users, scored by a process with its own documented gaming vectors, which makes it exactly the kind of metric lesson 15 warned against trusting as a go/no-go signal
0019 Reviewing an Eval Report Name specifically what would make an eval report's number untrustworthy (contamination, an uncalibrated judge, a single run with no variance estimate, a leaderboard rank standing in for a task-specific call), rather than saying it feels thin

Reference

  • Glossary: canonical terms for this topic
  • Resources: trusted sources
  • Held-Out Data and Contamination: the two contamination pathways, the guided-instruction test, and the prevention toolkit for a custom eval set
  • Metrics: exact match, F1, BLEU/ROUGE and pass@k side by side, with the pass@k formula and the decision principle for choosing among them
  • LLM-as-Judge: the judge-prompt design checklist, calibration and score compression, position and verbosity bias, and the human-vs-human agreement bar
  • Eval Harness: the four-stage pipeline, openai/evals vs lm-evaluation-harness, and what a defensible number logs alongside the score
  • Go/No-Go: the standard-error formula, paired comparison, setting a threshold honestly, and everything a complete defense cites

How this works

Each lesson is short and self-contained. Answer keys are collapsed: recall first, then open them. The real-world reps matter more than the reading, and spacing them out is the point. Anything still unclear at the end of a lesson is worth chasing to its primary source before moving on.

Table of contents