The Hard Truth About AI Agents: Lessons from Running Agents in Production

The Hard Truth About AI Agents: Lessons from Running Agents in Production

🎙 Hannes Hapke 👥 5K 📅 October 23, 2025 ⏱ 33 min 👁 79 📄 expert opinion 🧭 2026-08-15
Available in: English (current) Français

Keywords

AI agentsproductionreliabilityobservabilityguardrails

Summary

Hannes Hapke, Principal ML Engineer at Digits, shares lessons from deploying AI agents in production for financial applications. He defines agents as LLM-driven loops with tool calls, but emphasizes the need for robust infrastructure. Key components include agent memory, retrieval, LLM proxies, observability, and guardrails. He discusses challenges with frameworks, tool integration, and task planning. He highlights the importance of observability using OpenTelemetry and tools like Arize Phoenix. He demonstrates the value of agent memory in personalizing responses. He advises on using guardrails with a different LLM as judge and mentions responsible AI practices. He notes that MCP and A2A are not yet production-ready due to security concerns. He concludes with a live demo of a synchronous agent answering financial questions and emphasizes evaluating frameworks carefully.

126 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable, practical insights from real production experience, which is rare and highly useful. The speaker argues for careful evaluation of frameworks, emphasizing dependency management and the benefits of implementing the core loop directly. He supports his points with concrete examples, such as the vendor hydration use case and the memory demo. The argumentation is coherent and grounded in his experience, though it relies on anecdotal evidence rather than systematic studies.

Scientific Rigor, Source Quality, Title Accuracy

The talk is based on the speaker’s professional experience and his book ‘Generative AI Design Patterns’ (co-authored with Valliappa Lakshmanan). He mentions specific tools and frameworks (e.g., LangChain, CrewAI, Arize Phoenix, Guardrails AI) but does not provide formal citations. The title accurately reflects the content, focusing on practical lessons. The talk is a single presentation, so there are no comments to analyze.

150 words

Title / Content Match

The title accurately reflects the content, which focuses on practical lessons from deploying AI agents in production.

Quality & Reliability

8/10

Talk by a principal ML engineer with production experience, grounded in real-world implementations, but lacks formal citations and is based on anecdotal evidence.

Key Moments

Cited Sources

Concurring Sources

  • Generative AI Design Patterns — Book co-authored by the speaker, providing patterns for GenAI applications.

Contribution & Novelties

The talk provides a candid, field-tested perspective on deploying AI agents in production, highlighting common pitfalls and practical solutions. It emphasizes the importance of observability, memory, and guardrails, and offers concrete advice on tool integration and task planning.

Pour aller plus loin :

  • Generative AI Design Patterns — Book by the speaker and Valliappa Lakshmanan, covering design patterns for GenAI applications.
  • OpenTelemetry — Standard for observability, used by the speaker for agent monitoring.
  • Arize Phoenix — Open-source observability tool for LLM applications, mentioned in the talk.
  • Guardrails AI — Framework for adding guardrails to AI applications, mentioned in the talk.
  • MCP (Model Context Protocol) — Protocol for connecting AI models to external tools, discussed with security concerns.

117 words

Radar Profile

The radar profile shows high scores in information quantity and quality, with moderate technical depth and reliability. This reflects a talk that is rich in practical insights but relies on anecdotal evidence rather than formal research.

Reliability 7/10