Stanford CME295 Transformers & LLMs | Autumn 2025 | Lecture 8 - LLM Evaluation

Stanford CME295 Transformers & LLMs | Autumn 2025 | Lecture 8 - LLM Evaluation

🎙 Afshine Amidi, Shervine Amidi 👥 1.2M 📅 December 2, 2025 ⏱ 109 min 👁 156K 📄 lecture 🧭 2026-08-06
Available in: English (current) Français

Keywords

evaluationLLMmetricsbenchmarksLLM-as-a-judge

Summary

This lecture from Stanford’s CME295 course focuses on evaluating large language models (LLMs). It begins by discussing the challenges of evaluating free-form outputs, including subjectivity and cost of human ratings. The lecture introduces inter-rater agreement metrics like Cohen’s kappa to quantify consistency among human evaluators. It then covers rule-based metrics such as METEOR, BLEU, and ROUGE for comparing model outputs to references. The core of the lecture is the LLM-as-a-judge paradigm, where a language model is used to evaluate outputs, discussing best practices and biases like position, verbosity, and self-enhancement. The lecture also covers factuality evaluation, agent evaluation, and various benchmarks including MMLU for knowledge, AIME and PIQA for reasoning, SWE-bench for coding, HarmBench for safety, and Tau-Bench for agents. The instructors emphasize the importance of evaluation for improving LLM performance and provide practical guidance for implementing evaluation systems.

139 words

Critical Evaluation

The lecture provides a comprehensive overview of LLM evaluation, covering both traditional metrics and modern approaches. The content is well-structured, starting with the motivation for evaluation and progressing through human evaluation, rule-based metrics, and LLM-as-a-judge. The explanations are clear and accessible, with concrete examples that illustrate key concepts. The discussion of inter-rater agreement metrics is particularly valuable, as it highlights the importance of consistency in human evaluation and introduces statistical measures to quantify agreement beyond chance. The lecture also addresses practical considerations, such as the cost and speed of human evaluation, and presents LLM-as-a-judge as a scalable alternative. The coverage of biases in LLM-as-a-judge, including position, verbosity, and self-enhancement, is thorough and provides actionable insights for practitioners. The section on benchmarks is useful, though it could benefit from more depth on each benchmark’s design and limitations. The lecture does not cite specific research papers, but it is based on established methods and is likely informed by current literature. The presentation style is engaging, with the instructors effectively using the whiteboard to explain formulas and concepts. Overall, this is a high-quality educational resource that would benefit students and practitioners seeking to understand LLM evaluation. The only minor weakness is the lack of external references, which could help viewers explore topics in more depth.

212 words

Title / Content Match

The title accurately reflects the content, which focuses on LLM evaluation methods and practices.

Quality & Reliability

8/10

Lecture from Stanford University, presented by adjunct lecturers, covering established evaluation methods with clear explanations and examples. The content is well-structured and aligns with current practices in LLM evaluation, though it does not include original research or external citations beyond course materials.

Chapters

Cited Sources

Concurring Sources

  • LLM-as-a-judge paper — Research paper on using LLMs as judges, aligning with the lecture's discussion of LLM-as-a-judge.
  • MMLU benchmark — The MMLU benchmark paper, which the lecture references for knowledge evaluation.
  • SWE-bench — The SWE-bench paper, which the lecture mentions for coding evaluation.

Contribution & Novelties

The lecture provides a structured overview of LLM evaluation, synthesizing common practices and highlighting key considerations. It offers practical guidance on implementing evaluation systems, including the use of inter-rater agreement metrics and LLM-as-a-judge. The discussion of biases in LLM-as-a-judge is particularly valuable for practitioners.

Pour aller plus loin :

  • LLM-as-a-judge paper — Discusses the use of LLMs as judges for evaluating other LLMs, a core topic of the lecture.
  • MMLU benchmark — The Massive Multitask Language Understanding benchmark, referenced in the lecture for knowledge evaluation.
  • SWE-bench — A benchmark for evaluating LLMs on real-world software engineering tasks, mentioned in the lecture.

101 words

Radar Profile

The radar chart shows a balanced profile with high scores in information quantity, quality, and reliability, and a slightly lower score in technical depth. This indicates a comprehensive and reliable lecture that is accessible to a broad audience, though it may not delve into the most advanced technical details.

Reliability 8/10