Building Conversational AI Agents with Thread-Level Eval Metrics

Building Conversational AI Agents with Thread-Level Eval Metrics

🎙 Tony Kipkemboi & Claire Longo 👥 5K 📅 October 23, 2025 ⏱ 75 min 👁 63 📄 tutorial 🧭 2026-08-15
Available in: English (current) Français

Keywords

CrewAIComet OpikLLM-as-a-Judgehuman-in-the-loopthread-level evaluation

Summary

This talk, recorded at MLOps World 2025, presents a framework for building and evaluating conversational AI agents. Tony Kipkemboi from CrewAI explains the concepts of agents, crews, and flows, using a company analogy to illustrate how multi-agent systems can be orchestrated. Claire Longo from Comet introduces Opik, an open-source tool for LLM observability and evaluation. They demonstrate a customer support chatbot that classifies intents, routes to specialized agents, and uses LLM-as-a-Judge metrics and human feedback to monitor and improve performance. The session emphasizes the importance of thread-level evaluation over single-prompt tracing, and shows how to integrate evaluation into the development lifecycle. The talk is practical, with a live coding demo, and aims to help developers move from prototypes to production-ready agents.

121 words

Critical Evaluation

Value of the Information & Strength of the Argument

The value of the information lies in its practical, actionable guidance for building and evaluating AI agents. The speakers provide a clear conceptual framework (agents, crews, flows) and demonstrate a concrete implementation using CrewAI and Opik. The argumentation is persuasive, using real-world examples of AI failures (Chevy chatbot, McDonald’s drive-thru) to highlight the need for observability and evaluation. However, the argumentation is largely anecdotal and lacks rigorous empirical evidence. The speakers acknowledge limitations of LLM-as-a-Judge (e.g., variability) but argue for its utility in monitoring trends. The session is more of a tutorial than a scientific study, so the argumentation is based on practical experience rather than formal research.

Scientific Rigor, Source Quality, Title Accuracy

The scientific rigor is moderate. The speakers are industry practitioners with relevant expertise, and they reference open-source tools and frameworks. However, they do not cite academic papers or formal studies. The sources cited are primarily the tools themselves (CrewAI, Opik) and the conference website. The title accurately reflects the content, focusing on thread-level evaluation metrics. The talk is well-structured and the technical details are consistent with current practices in the field. The lack of formal citations and reliance on anecdotal evidence reduces the overall rigor.

208 words

Title / Content Match

The title accurately reflects the content, focusing on building conversational AI agents with thread-level evaluation metrics.

Quality & Reliability

7/10

The session provides a practical, hands-on tutorial with concrete examples and references to open-source tools (CrewAI, Comet Opik). The speakers are industry practitioners with relevant expertise. However, the content is largely promotional and lacks formal citations or rigorous scientific validation.

Key Moments

Cited Sources

  • MLOps World — Conference website where the talk was recorded.

Concurring Sources

  • CrewAI — Open-source framework for orchestrating AI agents.
  • Comet Opik — Open-source tool for LLM evaluation and monitoring.

Contribution & Novelties

The talk provides a practical, integrated approach to building and evaluating conversational AI agents, combining orchestration (CrewAI) with observability and evaluation (Comet Opik). It emphasizes thread-level evaluation, which is more holistic than single-prompt tracing. The session offers a reusable template for a customer support chatbot, demonstrating how to implement LLM-as-a-Judge and human-in-the-loop feedback. This is valuable for practitioners looking to move from prototypes to production.

Pour aller plus loin :

97 words

Radar Profile

The radar profile shows high scores in information quantity and quality, reflecting the practical depth of the tutorial. The technical level is moderate, suitable for a technical audience. The overall reliability is moderate, as the content is based on industry experience rather than formal research.

Reliability 6/10