Evals Resources
Knowledge
- Docs: "Define success criteria and build evaluations", Claude Platform Docs
Practitioner walkthrough of turning a vague "did it get better" question into measurable success criteria and a held-out eval set. Use for: designing the eval itself before reaching for a framework. - Repo: openai/evals, OpenAI
Official framework for defining and running an eval as code: prompts, grading logic, and a registry of existing evals to read as worked examples. Use for: how to structure and run a custom eval. - Repo: lm-evaluation-harness, EleutherAI
The de facto standard harness for running a model against standardized benchmarks, with the task configs showing how held-out sets are structured and scored in practice. Use for: running or adapting an existing benchmark rather than building an eval from zero. - Paper: "Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena", Zheng et al., 2023
Introduces LLM-as-judge for open-ended tasks and measures its biases against human preference: position bias, verbosity bias, self-enhancement bias. Use for: deciding whether an LLM judge is trustworthy for a given case, and what to correct for if it is. - Docs: Evaluate, Hugging Face
Library of standard task-specific metrics (BLEU, ROUGE, exact match, F1, and more) with the definition and failure modes of each. Use for: the task-specific-metric side of the metric-vs-LLM-judge comparison. - Paper: "Time Travel in LLMs: Tracing Data Contamination in Large Language Models", Golchin and Surdeanu, 2023
A concrete method for testing whether a benchmark's data leaked into a model's training set, with the guessing-the-rest-of-the-instance technique that catches it. Use for: defending a held-out set's honesty against the specific claim "the model just memorized this". - Paper: "Beyond the Imitation Game: Quantifying and Extrapolating the Capabilities of Language Models" (BIG-bench), Srivastava et al., 2022
Introduces the canary-string convention (a unique marker phrase embedded in benchmark data, asking crawlers to exclude it from training corpora) as a preventive contamination-resistance technique for a benchmark's own release. Use for: designing a custom eval set to resist contamination from the start, rather than detecting it after the fact. -
Paper: "Evaluating Large Language Models Trained on Code" (Codex), Chen et al., 2021
Introduces functional correctness (execute generated code against test cases rather than comparing text) and the unbiased pass@k estimator, with the combinatorial formula that avoids the high variance of directly re-sampling k completions. Use for: evaluating code generation, and for the general principle of checking behavior over text similarity wherever a task is executable. -
Article: "Cohen's kappa", Wikipedia
The chance-corrected inter-rater agreement statistic: its formula, why a kappa of 0 means no better than chance, and its known tendency to underestimate agreement on a rare category. Use for: measuring human-rater agreement precisely instead of eyeballing a raw agreement percentage. - Article: "Inter-rater reliability", Wikipedia
Covers why raw joint-probability agreement is misleading (inflated by chance, worse with fewer categories) and Krippendorff's alpha as the generalization of chance-corrected agreement to any number of raters and any level of measurement. Use for: choosing the right agreement statistic for more than two raters or non-categorical ratings. - Paper: "HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal", Mazeika et al., 2024
Introduces a standardized framework for evaluating automated red-teaming methods against models and defenses, at scale and comparably, where the field previously lacked one. Use for: attack success rate as a metric, and the distinction between an attack-generation method and a rigorous way to evaluate it. - Paper: "XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models", Röttger et al., 2023
A test suite of safe prompts written to resemble unsafe ones, specifically to surface over-refusal. Use for: the precise tension between harmlessness (refuse unsafe prompts) and helpfulness (don't refuse safe ones), and a concrete way to measure a model landing badly on that trade-off. - Paper: "τ-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains", Yao et al., 2024
Evaluates agents by comparing the end-of-conversation database state to an annotated goal state, and introducespass^kfor measuring an agent's reliability across repeated trials of the same task. Use for: why task success beats step-by-step trajectory matching, and the precise, easy-to-confuse distinction betweenpass^kand lesson 4'spass@k. - Paper: "AgentBench: Evaluating LLMs as Agents", Liu et al., 2023
A multi-dimensional benchmark across 8 distinct interactive environments testing an LLM's reasoning and decision-making as an agent, finding a significant gap between top models and smaller ones specifically in agentic settings. Use for: evidence that agentic capability is a distinct thing to measure, not implied by single-turn benchmark performance. - Paper: "Ragas: Automated Evaluation of Retrieval Augmented Generation", Es et al., 2023
A reference-free framework separating a RAG system's retrieval-quality, faithfulness, and generation-quality dimensions into independent metrics. Use for: computing a faithfulness score without needing a human-written gold answer, and for keeping retrieval failure and generation failure diagnosable as separate questions. - Paper: "Enabling Large Language Models to Generate Text with Citations" (ALCE), Gao et al., 2023
Introduces automatic citation-quality metrics correlated with human judgment, and the finding that even top systems lack complete citation support for roughly half their claims on one test set. Use for: citation correctness as a stricter, separate check from faithfulness. - Paper: "Seven Rules of Thumb for Web Site Experimenters", Kohavi et al., 2014
Kohavi and coauthors' practitioner guidance on running online controlled experiments, including the OEC concept (movable and causally connected to the real outcome) and guardrail metrics. Use for: choosing a defensible OEC instead of an easily-moved but uncausal metric, and why a guardrail metric has to be tracked alongside it. - Guide: "Rules of Machine Learning", Zinkevich, Google
Google's internal ML engineering guidance, including Rule #37 (measuring training/serving skew as several distinct comparisons, not one) and Rule #8 (freshness requirements). Use for: precisely which of an offline/online gap's several possible causes, ordinary variance, drift, or an engineering error, actually applies, and why each needs a different fix. - Paper: "The Tail at Scale", Dean and Barroso, 2013
The canonical explanation of why tail latency, not average latency, determines real-world responsiveness at scale, especially once a request fans out to many backend calls. Use for: why a go/no-go call needs a percentile (p95/p99), not just an average, before ruling out a latency regression. - Paper: "The Leaderboard Illusion", Singh et al., 2025
Documents selective disclosure on Chatbot Arena: private pre-release testing of many model variants (27 for Llama 4 alone) and large, asymmetric data access favoring a handful of closed-model providers. Use for: a concrete, current, well-documented example of why a leaderboard rank isn't a neutral, trustworthy measurement on its own. - Article: "Does Style Matter?", LMArena
The arena's own analysis showing rankings shift meaningfully once response length and markdown formatting are controlled for. Use for: confirming lesson 6's verbosity-bias mechanism operates identically at full leaderboard scale, with human raters instead of a single LLM judge.
Gaps
- No dedicated source yet on statistical significance testing for eval score differences (standard error of a proportion, paired significance tests like McNemar's). The mission needs this for stage 5's "is this difference real or noise" question; lesson 9 teaches it from stable, standard statistical method rather than a single cited source, and this gap should close once a good practitioner-level source is found.