Stanford CS329A Self-Improving AI Agents | Part 7 | Self-Improvement and Deep Research Agents

Stanford CS329A Self-Improving AI Agents | Part 7 | Self-Improvement and Deep Research Agents

🎙 Aakanksha Chowdhery 👥 1.2M 📅 August 3, 2026 ⏱ 72 min 👁 1K 📄 lecture 🧭 2026-08-06
Available in: English (current) Français

Keywords

AlphaCodeAlphaCode2Search-O1Search-R1competitive programming

Summary

This lecture from Stanford’s CS329A course, taught by Aakanksha Chowdhery, explores self-improvement in AI agents through search. It begins with AlphaCode, a system that generates a million code samples per problem, filters and clusters them, and selects a subset for submission to competitive programming platforms. AlphaCode achieved a top 54% ranking among human competitors. The lecture then discusses AlphaCode2, which fine-tunes Gemini Pro with a learned scoring model, reaching the 85th percentile. The discussion highlights how solve rate scales with sample budget and identifies selection and clustering as bottlenecks. The lecture then introduces Search-O1, a method that triggers search queries when a reasoning model expresses uncertainty, outperforming standard RAG on benchmarks like GPQA and HotpotQA. Finally, it compares Search-O1’s prompting-based approach to Search-R1’s reinforcement-learning-based method for teaching models when to search.

131 words

Critical Evaluation

The lecture provides a rigorous and detailed examination of self-improvement through search in AI agents. It is grounded in well-known research papers (AlphaCode, AlphaCode2, Search-O1, Search-R1) and offers a clear explanation of the methodologies, including pre-training, fine-tuning, large-scale sampling, filtering, clustering, and selection. The instructor demonstrates a deep understanding of the subject, providing insights into the bottlenecks and trade-offs involved. The content is technically rich, with specific details on model sizes, sample counts, and benchmark results. The lecture also encourages critical thinking by asking students to consider sources of variance in performance across different contests. The sources cited are the official course materials and the referenced papers, which are appropriate for an academic lecture. The title accurately reflects the content, and the lecture is well-structured, with a logical flow from code generation to deep research agents. The main limitation is that the lecture assumes prior knowledge of LLMs and test-time compute, making it less accessible to a general audience. However, for its intended audience, it is an excellent resource.

169 words

Title / Content Match

The title accurately reflects the content, which covers self-improvement through search and deep research agents.

Quality & Reliability

8/10

Lecture from Stanford CS329A, presented by an adjunct professor with deep expertise in LLMs. Content is based on peer-reviewed research (AlphaCode, AlphaCode2, Search-O1, Search-R1) and includes technical details. The lecture is well-structured and provides critical analysis of the methods.

Key Moments

Cited Sources

Concurring Sources

  • AlphaCode paper — The original AlphaCode paper, which the lecture discusses in detail.
  • AlphaCode2 paper — The AlphaCode2 paper, which the lecture discusses as an improvement over AlphaCode.

Dissenting Sources

Contribution & Novelties

This lecture provides a comprehensive overview of self-improvement through search, covering both code generation (AlphaCode, AlphaCode2) and deep research agents (Search-O1, Search-R1). It offers insights into the scaling behavior of sample budgets and the importance of selection and clustering in large-scale sampling. The comparison between prompting-based and reinforcement-learning-based approaches for teaching models when to search is particularly valuable.

Pour aller plus loin :

  • AlphaCode paper — The original AlphaCode paper, detailing the methods discussed.
  • AlphaCode2 paper — The AlphaCode2 paper, which fine-tunes Gemini Pro.
  • Search-O1 paper — The Search-O1 paper, introducing uncertainty-based search triggering.
  • Search-R1 paper — The Search-R1 paper, using reinforcement learning to teach when to search.

108 words

Radar Profile

The radar profile shows high scores across all dimensions, indicating a well-balanced and informative lecture. The highest scores are in quality of information and technical level, reflecting the depth and accuracy of the content. The lecture is particularly strong in providing detailed technical explanations and critical analysis.

Reliability 8/10