
How Anthropic uses Claude Code: Agentic Software Engineering at Scale - Daisy Hollman
Keywords
Summary
216 words
Critical Evaluation
Value of the Information & Strength of the Argument
The talk provides valuable insights into the practical challenges of scaling agentic software engineering, particularly around context window limitations and the need for customization. Hollman’s argumentation is coherent and grounded in her direct experience at Anthropic. She effectively uses analogies (e.g., ED vs. VS Code) to illustrate the current state of agent tooling. The claim that context engineering is becoming the core of software engineering is compelling and well-supported by examples. However, some arguments rely on anecdotal evidence and projections (e.g., Moore’s law of agents) that are not rigorously substantiated.
Scientific Rigor, Source Quality, Title Accuracy
The talk demonstrates strong technical rigor, with clear explanations of tool call mechanics and plugin design. Hollman references METR’s chart and Mozilla’s data, but these are not formally cited with URLs. The title accurately reflects the content, focusing on Anthropic’s use of Claude Code. The talk is an expert opinion rather than a peer-reviewed study, so the scientific rigor is moderate. The description provides links to NDC conferences but no direct references to the mentioned data sources.
182 words
Title / Content Match
The title accurately reflects the content: the talk focuses on how Anthropic uses Claude Code for agentic software engineering at scale, covering context engineering and plugin design.
Quality & Reliability
8/10
The speaker is a senior engineer at Anthropic with deep technical expertise, providing an insider perspective on Claude Code's design and usage. The talk is grounded in practical experience and references real-world data (e.g., Mozilla's bug fixes). However, it is largely anecdotal and lacks peer-reviewed sources, and some claims (e.g., Moore's law of agents) are presented without rigorous evidence.
Key Moments
Markers derived by PSI from the transcript: the creator did not define chapters.
- Introduction: Daisy Hollman introduces herself and her role at Anthropic, working on Claude Code.
- Definition of agents vs. chatbots, and the role of tool calls.
- Discussion of the primitive nature of current agent tools, like the edit tool.
- Presentation of METR's chart on agent time horizons and the 'Moore's law of agents'.
- Mozilla's example of bug fixes shipped in April, illustrating practical impact.
- Thesis: Claude needs access to all your tools and knowledge to do your job.
- Explanation of in-context learning and why customization happens in text space.
- The context window as a fixed-size box and the need for context engineering.
- Introduction to Claude Code plugins and their components.
- Post-tool-use hooks for immediate feedback, like red squiggles for agents.
- Software engineering as teaching agents, and the role of context engineers.
- How Anthropic uses Claude Code internally, including agent teams and scaling.
Cited Sources
- NDC Conferences — Conference organizer for the talk.
- NDC Copenhagen — Specific conference where the talk was recorded.
Concurring Sources
- METR's research on AI task time horizons — Referenced in the talk for the chart on agent capabilities.
- Mozilla's blog on AI-assisted bug fixing — Referenced for the example of bug fixes shipped in April.
Contribution & Novelties
The talk provides an insider perspective on how Anthropic designs and uses Claude Code for large-scale agentic software engineering. It introduces the concept of context engineering as a core discipline, and discusses plugin design and post-tool-use hooks as mechanisms for customization. The emphasis on the stagnation of context window sizes and the need to optimize within that constraint is a novel framing.
Pour aller plus loin :
- In-context learning — Relevant to the discussion of customization via text.
- Tool use in LLMs — Background on tool calling.
- Context window — Explanation of the concept.
94 words
Radar Profile
The radar profile shows high scores in information quantity and quality, reflecting the speaker's expertise and the depth of content. The technical level is high, but the reliability is slightly lower due to the anecdotal nature of some claims. This suggests a talk that is informative and technically rich but may require additional verification for some assertions.