CI/CD Fails for AI & How CC/CD Fixes It

CI/CD Fails for AI & How CC/CD Fixes It

🎙 Aishwarya Naresh Reganti & Sai Kiriti 👥 5K 📅 November 20, 2025 ⏱ 33 min 👁 445 📄 expert opinion 🧭 2026-08-15
Available in: English (current) Français

Keywords

non-determinismagencycontrolcalibrationevaluation

Summary

The talk, recorded at MLOps World | GenAI Summit 2025, introduces the Continuous Calibration / Continuous Development (CC/CD) framework as an alternative to traditional CI/CD for AI systems. The speakers argue that AI systems are fundamentally non-deterministic and involve an agency-control trade-off, making standard software engineering practices insufficient. They propose a loop that starts with low-agency, high-control deployments and gradually increases autonomy as the system earns trust. The framework consists of continuous development (scoping capability, building simple architecture, designing custom evals) and continuous calibration (running evals, analyzing behavior, applying fixes). The talk uses a customer support agent as a concrete example, illustrating how to incrementally increase agency from routing to co-pilot to resolution assistant. The speakers emphasize avoiding overengineering, using data to drive architectural changes, and maintaining rigorous evaluation and monitoring. They draw on their experience with 50+ deployments and teaching at MIT and Oxford. The presentation is practical and aimed at engineering leaders and product managers building AI products for the enterprise.

163 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable insights into the challenges of deploying AI systems in production, particularly the non-deterministic nature and the need for careful control. The CC/CD framework is a practical contribution, offering a structured approach to incrementally increase autonomy. The argumentation is coherent, using the agency-control trade-off as a central concept and illustrating it with real-world examples like coding assistants and customer support. However, the claims are largely based on anecdotal experience rather than rigorous empirical evidence. The speakers do not provide quantitative data or formal evaluations of the framework’s effectiveness. The argument would be stronger with more concrete case studies or metrics.

Scientific Rigor, Source Quality, Title Accuracy

The talk is an expert opinion piece, not a scientific study. The speakers cite their experience and teaching at MIT and Oxford, but no specific sources are mentioned. The description provides a link to the MLOps World conference, but no further references. The title accurately reflects the content, which is a high-level overview of the CC/CD framework. The talk is well-structured and the arguments are logically presented, but the lack of citations and empirical data limits its scientific rigor. The speakers do not engage with existing literature on MLOps or AI reliability, which would have strengthened their case.

216 words

Title / Content Match

The title accurately reflects the content, which contrasts traditional CI/CD with the proposed CC/CD framework for AI systems.

Quality & Reliability

7/10

The speakers are experienced practitioners (ex-AWS, OpenAI) and the talk is based on 50+ real-world deployments. However, the presentation is largely anecdotal and lacks formal citations or empirical data. The framework is presented as a practical methodology rather than a rigorously validated scientific study.

Key Moments

Cited Sources

  • MLOps World — Conference where the talk was recorded.

Concurring Sources

Contribution & Novelties

The talk introduces the CC/CD framework, which is a novel adaptation of CI/CD for AI systems, emphasizing continuous calibration and incremental autonomy. It provides a practical blueprint for managing non-determinism and the agency-control trade-off. The framework is based on the speakers’ experience with 50+ deployments, offering actionable steps for building AI products.

Pour aller plus loin :

84 words

Radar Profile

The radar profile shows high scores in quantity of information and fiability, but moderate in technical level and quality. This reflects a talk that is rich in practical advice but lacks deep technical depth and formal rigor.

Reliability 7/10