Stanford CS229 Machine Learning | Spring 2026 | Lecture 14: Transformers, In-Context Learning

Stanford CS229 Machine Learning | Spring 2026 | Lecture 14: Transformers, In-Context Learning

🎙 Chris Ré, Tengyu Ma 👥 1.2M 📅 July 31, 2026 ⏱ 77 min 👁 4K 📄 lecture 🧭 2026-08-03
Available in: English (current) Français

Keywords

transformersin-context learningtokenizationautoregressivelarge language models

Summary

This lecture from Stanford CS229 introduces large language models, focusing on transformers and in-context learning. The instructor begins with tokenization, explaining why subword tokenization is preferred over character or word level, and mentions byte pair encoding (BPE). He then formalizes the autoregressive decomposition of sequence probability using the chain rule, and outlines the transformer architecture as a parameterization of conditional distributions. The lecture covers key components such as embeddings, attention mechanisms, and the training objective of next-token prediction. In-context learning is discussed as a phenomenon where models adapt to tasks based on context without weight updates. The instructor also touches on practical considerations like vocabulary size and tokenizer differences. The lecture is part of a series and assumes prior knowledge of machine learning basics.

124 words

Critical Evaluation

The lecture provides a solid introduction to transformers and in-context learning, suitable for a graduate-level machine learning course. The explanation of tokenization is clear and well-motivated, with concrete examples. The formalization of autoregressive models using the chain rule is rigorous and sets the stage for understanding transformer training. The discussion of in-context learning is insightful, though it remains at a high level without diving into specific mechanisms or recent research. The lecture is well-structured, but it lacks depth in certain areas, such as the mathematical details of attention and the training dynamics of transformers. The sources cited are limited to course materials, which is appropriate for a lecture but does not provide external references for further reading. Overall, the content is accurate and pedagogically sound, but it may not offer new insights for those already familiar with the topic. The title accurately reflects the content, and the lecture meets its educational objectives.

152 words

Title / Content Match

The title accurately reflects the content, focusing on transformers and in-context learning.

Quality & Reliability

8/10

Lecture from Stanford CS229, taught by established professors. Content is technical and accurate, but limited to introductory material on transformers and in-context learning. No citations to external sources beyond course materials.

Key Moments

Cited Sources

Concurring Sources

Contribution & Novelties

This lecture provides a clear and accessible introduction to transformers and in-context learning, emphasizing the importance of tokenization and the autoregressive framework. It bridges the gap between theoretical foundations and practical applications, making it valuable for students new to large language models.

Pour aller plus loin :

87 words

Radar Profile

The radar profile shows strong scores in quality and reliability, with slightly lower scores in quantity and technical depth. This indicates a well-structured lecture that provides accurate information but may not cover all aspects in exhaustive detail.

Reliability 8/10