Building Effective Agents

Building Effective Agents

🎙 Sushant Mehta 👥 5K 📅 October 20, 2025 ⏱ 35 min 👁 178 📄 expert opinion 🧭 2026-08-15
Available in: English (current) Français

Keywords

agentsworkflowspost-trainingRLHFtool use

Summary

Sushant Mehta, Senior Research Engineer at Google DeepMind, presents a talk on building effective agents, recorded at MLOps World GenAI Summit 2025. He begins by explaining the importance of post-training in aligning LLMs for real-world tasks, outlining a typical three-stage pipeline involving fine-tuning, reward modeling, and reinforcement learning. He then defines agents as LLMs with agency over tools and actions, contrasting them with more deterministic workflows. He emphasizes that agents are best suited for open-ended, unpredictable tasks where dynamic planning is required, while simpler tasks may only need prompt engineering. He discusses core building blocks: augmented LLMs with retrieval and tools, sequential prompts with validation, model routing, and evaluator-optimizer loops. He highlights the importance of clear success criteria, feedback loops, and RL with verifiable rewards for improving agent performance. He also advises on when to avoid agents due to cost and latency, and cautions against over-reliance on frameworks that obscure underlying prompts. He concludes with best practices such as defining stopping criteria and incorporating human intervention points, and mentions customer support as a promising use case.

176 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable, practical insights into building effective agents, drawing on DeepMind’s applied research. The argumentation is solid, grounded in real-world engineering experience and established post-training pipelines. The speaker clearly distinguishes between workflows and agents, offering a decision framework for when to use each. He emphasizes simplicity and composability over over-engineered abstractions, which is a pragmatic and effective approach. The discussion of RL with verifiable rewards is particularly useful for scaling agent training. The advice on when not to use agents, considering cost and latency, is well-reasoned and adds practical value. The talk is well-structured, with clear examples and a logical progression from background to building blocks to best practices.

Scientific Rigor, Source Quality, Title Accuracy

The talk demonstrates scientific rigor by referencing well-known post-training pipelines (Tulu, Llama, Nemotron) and standard RL algorithms (PPO, DPO). However, it does not provide formal citations or URLs, relying instead on the speaker’s expertise and general knowledge. The title accurately reflects the content, which focuses on practical patterns and decision frameworks for building effective agents. The talk is based on the speaker’s experience at Google DeepMind, lending it credibility. The lack of formal citations is a minor weakness, but the content aligns with current industry practices and research.

214 words

Title / Content Match

The title accurately reflects the content, which focuses on practical patterns and decision frameworks for building effective agents.

Quality & Reliability

8/10

Talk by a senior research engineer at Google DeepMind, based on practical experience and established post-training pipelines (Tulu, Llama, Nemotron). No formal citations but references to known open-source models and reports. High credibility due to speaker's expertise and alignment with industry practices.

Key Moments

Cited Sources

  • MLOps World — Conference website where the talk was recorded.

Concurring Sources

Contribution & Novelties

The talk provides a practical, engineering-focused perspective on building effective agents, emphasizing simplicity and composability over complex frameworks. It offers a clear decision framework for when to use agents versus workflows, and highlights the importance of RL with verifiable rewards for scalable agent training. The discussion of common patterns like model routing and evaluator-optimizer loops is valuable for practitioners.

Pour aller plus loin :

100 words

Radar Profile

The radar profile shows high scores in information quantity, quality, and reliability, with a slightly lower technical depth. This indicates a talk that is rich in practical content and credible, but not overly technical, making it accessible to a broad audience.

Reliability 8/10