Stanford CS336 Language Modeling from Scratch | Spring 2026 | Lecture 3: Architectures

Stanford CS336 Language Modeling from Scratch | Spring 2026 | Lecture 3: Architectures

🎙 Tatsunori Hashimoto 👥 1.2M 📅 April 15, 2026 ⏱ 89 min 👁 44K 📄 lecture 🧭 2026-08-03
Available in: English (current) Français

Keywords

transformerpre-normSwiGLURoPEhyperparameters

Summary

This lecture from Stanford’s CS336 course, taught by Tatsunori Hashimoto, provides a comprehensive survey of modern transformer architectures for language modeling. The instructor emphasizes learning from the collective experience of the field, as training many models is impractical. He reviews the evolution from the original transformer to current models like LLaMA, Qwen, and Gemma, highlighting common architectural choices. Key topics include the placement of layer normalization (pre-norm vs. post-norm), activation functions (ReLU vs. SwiGLU), positional encodings (sinusoidal vs. RoPE), and other hyperparameters. The lecture also covers training stability techniques and the trade-offs involved in architecture design. The goal is to equip students with the knowledge to make informed decisions when building their own models.

114 words

Critical Evaluation

The lecture provides an excellent, up-to-date survey of transformer architectures, drawing on a wide range of recent models. The instructor’s approach of synthesizing common practices and variations is highly valuable for practitioners. The content is technically rigorous, with clear explanations of the motivations behind architectural choices, such as the shift to pre-norm for training stability and the adoption of SwiGLU and RoPE. The lecture is well-structured, starting with the vanilla transformer and progressively introducing modern modifications. The discussion of hyperparameters and training stability tricks is particularly useful. The instructor also highlights the trade-offs involved, such as the balance between model capacity and computational efficiency. The sources cited are primarily the models themselves (e.g., LLaMA, Qwen) and the course materials, which are appropriate for the context. The lecture’s main strength is its practical orientation, preparing students to implement and train their own models. However, it could benefit from more formal citations to specific papers for each architectural choice. Overall, this is an outstanding lecture that effectively bridges theory and practice.

169 words

Title / Content Match

The title accurately describes the content: a lecture on language model architectures, part of the CS336 course.

Quality & Reliability

9/10

Lecture by a Stanford professor, based on a survey of recent language model architectures, with references to specific models and papers. The content is technical and detailed, but the lecture format and lack of formal citations in the transcript reduce the score slightly.

Key Moments

Cited Sources

Concurring Sources

Contribution & Novelties

This lecture provides a comprehensive and current survey of transformer architectures, synthesizing the design choices made in recent language models. It offers practical guidance for implementing and training models, emphasizing the importance of architectural decisions for stability and performance. The lecture also highlights the trade-offs involved and the rationale behind common practices.

Pour aller plus loin :

121 words

Radar Profile

The radar profile shows high scores across all dimensions, indicating a lecture that is both informative and technically deep. The balance between quantity and quality of information is particularly strong, with a slight emphasis on technical depth.

Reliability 9/10