[M2L 2025] 1.4 LLM Evaluation - Emine Yilmaz

[M2L 2025] 1.4 LLM Evaluation - Emine Yilmaz

🎙 Emine Yilmaz 👥 3K 📅 November 10, 2025 ⏱ 65 min 👁 109 📄 lecture 🧭 2026-08-15
Available in: English (current) Français

Keywords

evaluationmetricsBLEUROUGELLM-as-a-judge

Summary

The lecture by Emine Yilmaz at the Mediterranean Machine Learning summer school provides a comprehensive introduction to LLM evaluation. It begins by emphasizing the importance of evaluation in model development, highlighting risks like misinformation and bias. The talk categorizes evaluation metrics into quantitative (e.g., accuracy, BLEU, ROUGE, perplexity) and qualitative (e.g., fluency, coherence, factuality, safety) dimensions. It discusses the challenges of reference-based metrics and the shift towards LLM-based judges. The lecture covers human annotation methods, their costs, and the emergence of LLM-as-a-judge, including prompt engineering and the use of RAG and agents. It concludes with research directions on improving evaluation reliability and efficiency.

103 words

Critical Evaluation

Value of the Information & Strength of the Argument

The lecture provides a solid overview of LLM evaluation, systematically covering quantitative and qualitative metrics, and discussing their strengths and limitations. The argumentation is coherent, moving from why evaluation matters to specific metrics and evaluation methods. The speaker supports claims with examples and references to research, though not exhaustive. The value lies in its structured presentation of key concepts, making it a useful primer for researchers and practitioners.

Scientific Rigor, Source Quality, Title Accuracy

The lecture demonstrates scientific rigor by referencing established metrics and research, such as BLEU, ROUGE, and the GPT-4 judge paper. However, specific citations are not always provided, and the talk is more of an overview than a deep dive. The title accurately reflects the content, and the lecture is well-organized. The speaker’s expertise adds credibility, but the lack of detailed source citations limits the ability to verify all claims.

152 words

Title / Content Match

The title accurately reflects the content, which is a focused lecture on LLM evaluation.

Quality & Reliability

8/10

The lecture is delivered by an established researcher in information retrieval and evaluation, providing a structured overview of LLM evaluation methods. It covers both quantitative and qualitative metrics, discusses human and LLM-based evaluation, and references key papers. However, it lacks detailed citations and empirical validation of claims.

Key Moments

Cited Sources

  • GPT-4 as an LLM judge (paper) — Mentioned as a key paper on using LLMs as judges.

Concurring Sources

Contribution & Novelties

The lecture provides a structured overview of LLM evaluation, synthesizing existing knowledge. It emphasizes the importance of evaluation in model development and discusses emerging trends like LLM-as-a-judge and agent-based evaluation. The speaker’s perspective as a researcher adds depth.

Pour aller plus loin :

  • LLM Evaluation — Overview of LLM evaluation on Wikipedia.
  • BLEU — Detailed explanation of the BLEU metric.
  • ROUGE — Detailed explanation of the ROUGE metric.
  • Perplexity — Definition and use in language models.
  • LLM-as-a-Judge — Paper on using GPT-4 as a judge for LLM evaluation.

88 words

Radar Profile

The radar profile shows high scores in information quantity, quality, and technical level, with slightly lower but still strong reliability. This indicates a well-structured and informative lecture with minor gaps in source citation.

Reliability 8/10