Evaluating AI Agents and RAG Systems

Evaluating AI Agents and RAG Systems

🎙 David Okpare 👥 278 📅 November 27, 2025 ⏱ 48 min 👁 49 📄 tutorial 🧭 2026-08-16
Available in: English (current) Français

Keywords

RAGAI agentsevaluationLLMproduction

Summary

This technical session by David Okpare focuses on evaluating AI agents and Retrieval-Augmented Generation (RAG) systems. The speaker emphasizes the importance of evaluation to avoid silent failures in production, citing examples like Taco Bell’s AI drive-thru. He distinguishes between offline and online evaluation, advocating for starting with offline and deploying with online monitoring. The talk covers key metrics for RAG: precision@K and recall@K for retrieval, and faithfulness, correctness, and helpfulness for generation. He demonstrates a practical notebook using BM25 for retrieval and an LLM-as-a-judge for generation evaluation. For agents, he highlights the complexity of multi-step tool calls and the need to evaluate tool selection, parameter extraction, and the path to convergence. He introduces test cases and synthetic data for development. The session concludes with a live coding example and emphasizes the iterative nature of evaluation.

135 words

Critical Evaluation

Value of the Information & Strength of the Argument

The value of the information is high, providing practical guidance on evaluating AI systems. The speaker argues convincingly that evaluation is essential for production reliability, using real-world examples and logical reasoning. He explains the trade-offs between code-based evals and LLM-as-a-judge, and stresses the importance of recall over precision in retrieval. The argumentation is solid, though some parts are rushed due to time constraints.

Scientific Rigor, Source Quality, Title Accuracy

The scientific rigor is moderate; the speaker relies on personal experience and common practices rather than formal citations. The sources mentioned are limited to the datasets and libraries used in the demo. The title accurately reflects the content, which is a tutorial on evaluation techniques. No comments were provided for analysis.

129 words

Title / Content Match

The title accurately reflects the content, which focuses on evaluation techniques for RAG systems and AI agents.

Quality & Reliability

7/10

The speaker is an AI consultant with practical experience, and the content is based on established evaluation methodologies. However, the presentation is a live session with some improvisation and lacks formal citations.

Key Moments

Cited Sources

  • Instructor library — Mentioned as a library for structured LLM outputs.
  • BM25 — Used as a retrieval method in the demo.

Concurring Sources

Contribution & Novelties

The session provides a practical, hands-on approach to evaluating RAG systems and AI agents, emphasizing the importance of recall and the use of LLM-as-a-judge. It offers a clear framework for integrating evaluation into the development lifecycle.

Pour aller plus loin :

  • RAG evaluation metrics — Overview of metrics like faithfulness, answer relevancy, and context precision.
  • Agent evaluation frameworks — LangGraph’s guide on evaluating agents.
  • Synthetic data generation — Paper on generating synthetic data for LLM evaluation.

76 words

Radar Profile

The radar profile shows high scores in information quantity and technical level, with moderate scores in quality and reliability. This indicates a content-rich tutorial with practical insights, but with some limitations in formal rigor and source citation.

Reliability 7/10