Why AI Agents Fail in Production (Real Data)

Why AI Agents Fail in Production (Real Data)

🎙 AI Security Podcast 👥 20K 📅 January 14, 2026 ⏱ 51 min 👁 17K 📄 expert opinion 🧭 2026-08-16
Available in: English (current) Français

Keywords

AI agentsproductiongovernancesecurityMCP

Summary

In this episode of the AI Security Podcast, host Ashish and co-host Caleb interview Dev Rishi, GM of AI at Rubrik and former CEO of Predibase. The conversation centers on why enterprise AI agents often fail to move from experimentation to production. Dev shares data from his conversations with 180 organizations, revealing that 88% have started experimenting with agents, but only 22-23% have reached the formalization phase. He identifies three major fears among IT leaders: shadow agents (uncontrolled proliferation), lack of real-time governance, and the inability to undo catastrophic mistakes. To address these, he introduces the concept of ‘Agent Rewind’, a capability leveraging Rubrik’s backup infrastructure to roll back changes made by rogue agents. The discussion also covers the shift from fine-tuning to foundation models, the importance of automation-oriented use cases, and the debate between MCP and A2A protocols for agent interoperability. Dev argues that traditional anomaly detection is insufficient for AI security, advocating for AI-driven policy enforcement using small language models as judges. The episode concludes with insights on identity management for agents and the technical superiority of A2A despite MCP’s likely market dominance.

185 words

Critical Evaluation

Value of the Information & Strength of the Argument

The episode provides valuable insights into the practical challenges of deploying AI agents in enterprise settings, grounded in the speaker’s direct experience with numerous organizations. The argumentation is coherent, with clear examples such as the coding agent that deleted a production database, illustrating the need for remediation capabilities. Dev’s proposal of ‘Agent Rewind’ is a novel concept that leverages existing backup infrastructure, adding a practical dimension to the discussion. However, the argumentation is somewhat one-sided, as it promotes Rubrik’s solutions without critically examining alternative approaches or potential limitations.

Scientific Rigor, Source Quality, Title Accuracy

The scientific rigor is moderate; the discussion is based on anecdotal evidence and the speaker’s professional observations rather than peer-reviewed research. The sources cited are primarily the podcast’s own website and newsletter, which are not independent references. The title accurately reflects the content, focusing on the failure modes of AI agents in production. The episode does not reference external studies or publications, limiting its scholarly value. The inclusion of specific data points (e.g., 88% experimentation rate) adds credibility but lacks methodological detail. Overall, the title-content alignment is strong, but the sourcing is weak.

196 words

Title / Content Match

The title accurately reflects the core theme of the episode, which focuses on the challenges and failures of AI agents in production environments, supported by real-world data and expert insights.

Quality & Reliability

7/10

The podcast features an industry expert with direct experience in enterprise AI deployment, providing concrete data from 180 organizations. However, the discussion is largely anecdotal and promotional, with limited peer-reviewed sources or independent verification.

Chapters

Cited Sources

Concurring Sources

Contribution & Novelties

The episode offers a fresh perspective on AI agent deployment challenges, particularly the concept of ‘Agent Rewind’ as a remediation layer. It provides real-world data on adoption phases and highlights the governance gap between policy creation and enforcement. The discussion on using SLMs as judges for policy enforcement is an innovative idea that could influence future security architectures.

Pour aller plus loin :

120 words

Radar Profile

The radar profile shows high scores in information quantity and quality, reflecting the episode's rich content and expert insights. The technical level is moderate, suitable for a broad audience. The reliability score is slightly lower due to the lack of independent sources and the promotional nature of the discussion.

Reliability 6/10