Stanford CME295 Transformers & LLMs | Autumn 2025 | Lecture 3 - Transformers & Large Language Models

Stanford CME295 Transformers & LLMs | Autumn 2025 | Lecture 3 - Transformers & Large Language Models

🎙 Afshine Amidi, Shervine Amidi 👥 1.2M 📅 October 17, 2025 ⏱ 108 min 👁 114K 📄 lecture 🧭 2026-08-06
Available in: English (current) Français

Keywords

LLMTransformerMixture of ExpertsDecodingIn-context learning

Summary

This lecture, part of Stanford’s CME295 course, provides a comprehensive overview of Large Language Models (LLMs). It begins with a recap of Transformer-based architectures, categorizing them into encoder-decoder, encoder-only, and decoder-only models, with GPT as a prime example of the latter. The lecture then defines LLMs, emphasizing their scale in parameters, training data, and compute. A significant portion is dedicated to Mixture of Experts (MoE), explaining dense and sparse variants, and how they reduce computational cost by activating only a subset of experts. The instructors detail where MoE layers are placed in the Transformer (in the feed-forward networks) and discuss training challenges. The second half covers response generation, contrasting greedy decoding and beam search with sampling-based methods, and the impact of temperature on output diversity. Prompting strategies are explored, including in-context learning, chain-of-thought, and self-consistency. Finally, inference optimizations such as KV cache, PagedAttention, and Multi-Query Attention are introduced. The lecture is well-structured, with clear explanations and practical examples, making it suitable for graduate-level students.

164 words

Critical Evaluation

This lecture provides a solid, comprehensive introduction to Large Language Models, suitable for a graduate-level course. The content is accurate and well-structured, covering key concepts from architecture to inference. The instructors, Afshine and Shervine Amidi, demonstrate deep expertise and pedagogical skill, using clear analogies (e.g., the room of experts for MoE) and practical examples. The lecture excels in explaining complex topics like Mixture of Experts, breaking down dense vs. sparse variants and their computational benefits. The discussion on decoding strategies, including temperature and sampling, is particularly valuable for understanding model behavior. The coverage of prompting techniques, such as chain-of-thought and self-consistency, is timely and relevant. However, the lecture lacks explicit citations to primary research papers, relying instead on course materials and general knowledge. While this is acceptable for a lecture, it limits the ability to verify specific claims. The adéquation between title and content is excellent, as the lecture directly addresses Transformers and LLMs. The presentation is engaging, with interactive Q&A, but the video format may not be ideal for all learners. Overall, this is a high-quality educational resource, though it assumes prior knowledge of Transformers from previous lectures. The lack of peer-reviewed sources is a minor weakness, but the content is consistent with established knowledge in the field.

209 words

Title / Content Match

Title accurately reflects the content: a lecture on Transformers and LLMs, covering architecture, MoE, decoding, and prompting.

Quality & Reliability

9/10

Lecture from Stanford University, presented by adjunct lecturers with clear academic structure, covering established concepts in LLMs with references to course materials. Content is accurate and up-to-date, but lacks peer-reviewed citations.

Chapters

Cited Sources

Concurring Sources

Contribution & Novelties

This lecture provides a clear and structured overview of LLMs, with particular strength in explaining Mixture of Experts and decoding strategies. It serves as an excellent educational resource for students and practitioners. The lecture’s contribution lies in its pedagogical clarity and comprehensive coverage of key concepts.

Pour aller plus loin :

122 words

Radar Profile

The radar profile shows high scores across all dimensions, indicating a well-rounded lecture with strong information content, technical depth, and reliability. The lowest score is in technical level, but it remains high, reflecting the advanced nature of the material.

Reliability 9/10