Agents as Ordinary Software: Principled Engineering for Scale | Linus Lee, Thrive Capital

Agents as Ordinary Software: Principled Engineering for Scale | Linus Lee, Thrive Capital

🎙 Linus Lee 👥 5K 📅 October 23, 2025 ⏱ 28 min 👁 2K 📄 expert opinion 🧭 2026-08-15
Available in: English (current) Français

Keywords

composabilityobservabilitystatelessnesschangeabilityPolymer

Summary

Linus Lee, an Entrepreneur in Residence at Thrive Capital, presents the engineering principles behind ‘Puck’, an internal AI research assistant that autonomously handles thousands of tasks weekly. He emphasizes that scaling AI agents requires applying classical software engineering values: composability, observability, statelessness, and changeability. The system is built on an orchestration library called ‘Polymer’, which uses well-typed ’tasks’ as composable building blocks. Lee details how these principles are implemented: composability through modular tasks, observability via fully replayable logs, statelessness by modeling side effects as data, and changeability through adapters and evaluation suites. He provides real-world metrics, including 2.5 billion tokens of inference per week with 80% in background tasks, and illustrates debugging workflows using trace logs. The talk concludes that treating agents as ordinary software enables maintainability, testability, and scalability, allowing a small team to operate at significant scale.

139 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides high practical value for engineers building AI agent systems, offering concrete design patterns and architectural insights from a production environment. The argumentation is solid, grounded in real-world experience and specific examples, such as the 11-minute pipeline trace and the debugging process. Lee effectively argues that traditional software engineering principles are crucial for scaling AI systems, supporting his claims with metrics and architectural details. The reasoning is coherent and persuasive, though it relies on anecdotal evidence from a single organization rather than broader empirical validation.

Scientific Rigor, Source Quality, Title Accuracy

The talk demonstrates strong scientific rigor in its engineering approach, emphasizing reproducibility and systematic debugging. However, it lacks formal citations or references to external research, relying primarily on the speaker’s experience. The title accurately reflects the content, focusing on engineering principles for AI agents. The presentation is well-structured and technically detailed, suitable for a professional audience. No comments were provided for analysis.

164 words

Title / Content Match

The title accurately reflects the content: the talk focuses on applying traditional software engineering principles to build scalable AI agents, treating them as ordinary software components.

Quality & Reliability

8/10

The talk is a first-hand account from a practitioner at Thrive Capital, detailing the architecture and engineering principles behind their internal agent system 'Puck'. It provides concrete metrics and design patterns, but lacks external validation or peer review. The speaker's credibility is high given his background, but the content is primarily experiential rather than rigorously tested.

Key Moments

Cited Sources

  • MLOps World — Conference website where the talk was recorded.

Concurring Sources

Contribution & Novelties

The talk provides a unique perspective on applying classical software engineering principles to AI agent systems, offering a concrete framework (Polymer) and real-world metrics from a production environment. It bridges the gap between theoretical discussions of agent architecture and practical implementation details.

Pour aller plus loin :

  • Software Engineering at Google — The book referenced by the speaker for the definition of software engineering.
  • LLM Agents — A comprehensive overview of agent architectures and design patterns.
  • Observability in Machine Learning — Discusses the importance of observability in ML systems.

89 words

Radar Profile

The radar profile shows high scores in information quantity and quality, with a slightly lower technical level and reliability. This indicates a talk that is rich in practical insights and well-articulated, but may not delve into the deepest technical details or provide external validation.

Reliability 7/10