Stanford CS329A Self-Improving AI Agents | Part 2 | Test-Time Compute Scaling

Stanford CS329A Self-Improving AI Agents | Part 2 | Test-Time Compute Scaling

🎙 Azalia Mirhoseini 👥 1.2M 📅 August 3, 2026 ⏱ 63 min 👁 6K 📄 lecture 🧭 2026-08-05
Available in: English (current) Français

Keywords

test-time computeinference scalingpower lawverificationLLM

Summary

This lecture from Stanford’s CS329A course, delivered by Azalia Mirhoseini, focuses on improving LLM performance at inference time without additional training. It begins by contrasting pre-training, fine-tuning, and inference, then introduces the Large Language Monkeys paper, which shows that repeated sampling from a model, combined with a verifier, can significantly boost performance, even for smaller models. The lecture presents a scaling law for test-time compute: coverage (fraction of problems solved) follows a power law with the number of samples. This behavior is explained by a long tail of hard problems that are rarely solved at pass@1. The importance of automated verification is highlighted, with examples like unit tests for code and formal proofs for math. The generation-verification gap is discussed, showing that majority voting and reward models underperform compared to perfect verification. The lecture also covers a paper on optimally scaling test-time compute, comparing parallel sampling and sequential revision, and introduces outcome and process reward models. Finally, the Arkon paper on inference-time architecture search is presented, which combines multiple techniques and models, achieving a 14.1% improvement in pass@1 over GPT-4 and Claude 3.5 Sonnet. The lecture concludes with a discussion on when pre-training still outperforms added test-time compute for the hardest problems.

202 words

Critical Evaluation

The lecture provides a rigorous and well-structured overview of test-time compute scaling, a rapidly evolving area in AI. The speaker, Azalia Mirhoseini, is a recognized expert, and the content is based on recent, peer-reviewed research, lending high credibility. The presentation methodically builds from foundational concepts (repeated sampling) to more advanced topics (scaling laws, verification gaps, architecture search), making it accessible to a graduate-level audience. The use of empirical data and clear visualizations strengthens the arguments. The discussion of the generation-verification gap is particularly insightful, highlighting a critical bottleneck in current approaches. The lecture also thoughtfully addresses the trade-offs between pre-training and test-time compute, avoiding overhype. One minor limitation is that the lecture assumes familiarity with LLM fundamentals, but this is appropriate for the course context. The sources cited are primarily the papers discussed, which are appropriate and verifiable. Overall, this is an excellent, informative lecture that provides a solid foundation for understanding and applying test-time compute scaling.

157 words

Title / Content Match

The title accurately reflects the content: a lecture on test-time compute scaling for self-improving AI agents.

Quality & Reliability

9/10

Lecture from a Stanford graduate course by a recognized expert, presenting peer-reviewed research (Large Language Monkeys, test-time compute scaling, Arkon) with clear methodology and empirical results. High reliability due to academic context and explicit references.

Key Moments

Cited Sources

Concurring Sources

Dissenting Sources

  • No discordant sources identified — The lecture presents a consensus view within the field; no contradictory sources were mentioned.

Contribution & Novelties

This lecture provides a comprehensive and up-to-date synthesis of test-time compute scaling, a key technique for improving LLM performance without additional training. It clearly explains the power law scaling behavior, the generation-verification gap, and the trade-offs between parallel sampling and sequential revision. The discussion of the Arkon paper, which combines multiple inference-time techniques, offers a practical approach to achieving significant gains. The lecture also highlights the importance of automated verification and the potential of test-time compute for self-improving agents.

Pour aller plus loin :

138 words

Radar Profile

The radar profile shows high scores across all dimensions, indicating a well-rounded and reliable lecture. The strong performance in quantity and quality of information, combined with a high technical level and reliability, suggests this is an excellent resource for understanding test-time compute scaling.

Reliability 9/10

💬 No comments were provided for analysis.