Why Great Models Fail: Lessons From 9 Years of Deploying ML Models - Megan Robertson

Why Great Models Fail: Lessons From 9 Years of Deploying ML Models - Megan Robertson

🎙 Megan Robertson 👥 227K 📅 August 13, 2026 ⏱ 59 min 👁 64 📄 expert opinion 🧭 2026-08-13
Available in: English (current) Français

Keywords

model failureproduction MLproject scopingstakeholder alignmentmodel monitoring

Summary

In this talk from NDC Toronto, Megan Robertson shares lessons from nine years of deploying machine learning models in industry. She emphasizes that many models fail not due to technical flaws but due to inadequate scoping and lack of adaptation to changing real-world conditions. She illustrates this with a personal example of a social media topic modeling project that performed well in development but was never adopted because she didn’t consult stakeholders and underestimated data costs. Robertson outlines a four-step scoping process: define the project, understand constraints and risks, identify possible solutions, and plan maintenance and monitoring. She stresses the importance of quantifying success with KPIs and aligning with business value. The second major theme is that the world changes over time, so models degrade; she discusses monitoring strategies and model updating techniques. The talk provides practical advice for data scientists and ML engineers to increase the chances of their models succeeding in production.

154 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk offers valuable insights from real-world experience, highlighting common pitfalls in ML deployment. Robertson’s argumentation is clear and well-structured, using a personal case study to illustrate the consequences of poor scoping. She provides actionable strategies, such as defining quantitative success metrics and involving stakeholders early. The emphasis on value over technical novelty is a strong point, and the discussion of model degradation and monitoring is practical. However, the arguments are largely anecdotal and lack empirical evidence or references to industry studies, which slightly weakens the overall rigor.

Scientific Rigor, Source Quality, Title Accuracy

The talk is based on the speaker’s professional experience, which lends credibility but also limits generalizability. No external sources are cited, and the only links provided are to NDC conferences. The title accurately reflects the content, and the talk stays on topic. The lack of citations is a minor weakness, but the practical nature of the content compensates. The speaker’s expertise is evident, and the advice is grounded in real-world scenarios.

174 words

Title / Content Match

The title accurately reflects the content, which focuses on reasons why ML models fail in production and strategies to mitigate these failures.

Quality & Reliability

7/10

The talk is based on the speaker's 9 years of practical experience in ML deployment, providing concrete examples and actionable advice. However, it lacks formal citations or references to external research, and the evidence is anecdotal.

Key Moments

Cited Sources

  • NDC Conferences — Official NDC conference website, mentioned in the video description.
  • NDC Toronto — Official NDC Toronto conference website, mentioned in the video description.

Concurring Sources

  • Why Machine Learning Models Fail in Production — Article discussing common reasons for ML model failure, aligning with the talk's themes.

Dissenting Sources

  • The Myth of Model Interpretability — This paper argues that interpretability is often overrated, which contrasts with the talk's emphasis on stakeholder alignment and understanding.

Contribution & Novelties

The talk provides a practitioner’s perspective on common reasons for ML model failure in production, emphasizing the importance of project scoping and continuous monitoring. It offers a structured approach to scoping and practical advice for aligning ML projects with business goals. The speaker’s personal example adds authenticity.

Pour aller plus loin :

  • Machine Learning Model Monitoring — Overview of monitoring techniques.
  • Data drift — Concept of data drift and its impact on model performance.
  • MLOps — Practices for deploying and maintaining ML models in production.

85 words

Radar Profile

The radar profile shows balanced scores across information quantity, quality, technical level, and reliability, with a slightly higher emphasis on practical advice and real-world experience. The talk is strong in providing actionable insights but lacks formal citations, which is reflected in the reliability score.

Reliability 6/10

💬 No comments were provided for analysis.