TMLS Backstage 01: The State of Agents OPS

TMLS Backstage 01: The State of Agents OPS

🎙 Toronto Machine Learning Society (TMLS) 👥 5K 📅 September 29, 2025 ⏱ 71 min 👁 45 📄 expert opinion 🧭 2026-08-15
Available in: English (current) Français

Keywords

agentic workflowsRPAsecuritytool useMCP

Summary

This panel discussion, hosted by TMLS, brings together Denys Linkov (Head of ML at Wisedocs), Diego Oppenheimer (Head of Product at Hyperparam), and Priyankya Somrah (Principal at Work Bench) to explore the state of agentic operations. They discuss the necessity of agentic workflows compared to traditional RPA, highlighting the shift from deterministic to non-deterministic design. The panel addresses where agents are most valuable, such as in adaptive decision-making and handling unstructured data, while acknowledging that deterministic processes still have a place. They identify the weakest parts of the stack, including security, authentication, logging, and debugging in agentic systems. The conversation covers tool use, the role of MCP, and the challenges of tracing and monitoring in dynamic agent graphs. The panel also touches on enterprise adoption, noting security as a primary blocker, and concludes with a live Q&A session.

138 words

Critical Evaluation

Value of the Information & Strength of the Argument

The discussion provides valuable insights from practitioners with hands-on experience in AI infrastructure and enterprise software. The panelists offer a balanced view, acknowledging both the potential and the limitations of agentic systems. They argue convincingly that agentic workflows are not a replacement for all automation but are suited for tasks requiring adaptability and context awareness. The argumentation is solid, grounded in real-world examples like Excel rollouts and McDonald’s processes, and they critically examine the challenges of security, debugging, and tooling. However, some points are based on anecdotal evidence rather than rigorous data, and the discussion occasionally lacks depth on technical specifics.

Scientific Rigor, Source Quality, Title Accuracy

The panelists reference a ‘Benchmark Report on Agentic Ops’ but do not provide specific citations or data points from it. The discussion is largely based on personal experience and industry observations, which adds practical value but limits scientific rigor. The title accurately reflects the content, as it is a backstage conversation about the state of agentic operations. The sources cited are minimal, with only a link to the episode post in the description. No external sources are explicitly mentioned in the video, and the panelists do not cite specific studies or papers. The lack of formal references reduces the overall scientific credibility, but the practical insights from experienced professionals are still valuable.

228 words

Title / Content Match

Title accurately reflects the content: a backstage discussion on the state of agentic operations.

Quality & Reliability

7/10

Panel of experienced practitioners (ML lead, product head, VC principal) discussing agentic ops based on industry experience and a benchmark report. No formal citations, but practical insights and balanced perspectives. Some claims lack empirical backing, but overall credible.

Key Moments

Cited Sources

Concurring Sources

Contribution & Novelties

The video offers a practitioner’s perspective on the current state of agentic operations, highlighting the practical challenges and opportunities. It provides a nuanced view on where agentic workflows are beneficial versus traditional automation, and emphasizes the importance of security and observability. The discussion on the non-deterministic nature of agent systems and its implications for debugging is particularly insightful.

Pour aller plus loin :

113 words

Radar Profile

The radar profile shows balanced scores across information quantity, quality, technical level, and reliability, indicating a well-rounded discussion. The high scores in information quality and reliability reflect the practical expertise of the panelists, while the moderate technical level suggests the content is accessible to a broad audience.

Reliability 7/10

💬 No comments provided.