Stanford CME295 Transformers & LLMs | Autumn 2025 | Lecture 4 - LLM Training

Stanford CME295 Transformers & LLMs | Autumn 2025 | Lecture 4 - LLM Training

🎙 Afshine Amidi, Shervine Amidi 👥 1.2M 📅 October 21, 2025 ⏱ 107 min 👁 96K 📄 lecture 🧭 2026-08-06
Available in: English (current) Français

Keywords

pretrainingscaling lawsquantizationLoRAfine-tuning

Summary

This lecture from Stanford’s CME295 course focuses on the training of large language models (LLMs). It begins with a recap of previous lectures on mixture of experts and inference techniques. The core content covers the pretraining phase, emphasizing its high computational cost and the use of massive datasets (e.g., Common Crawl, Wikipedia, GitHub) with token counts in the hundreds of billions to trillions. The lecture introduces key metrics: FLOPs (floating-point operations) and FLOPS (operations per second), and discusses scaling laws, particularly the Chinchilla law, which guides optimal model size and data size for a given compute budget. Training optimizations are then explored, including data parallelism with ZeRO, model parallelism, and Flash Attention. The lecture also covers quantization and mixed precision training to reduce memory and accelerate computation. The second half focuses on supervised fine-tuning (SFT) and instruction tuning, explaining how pretrained models are adapted to specific tasks. Finally, parameter-efficient fine-tuning methods like LoRA and QLoRA are presented as ways to fine-tune large models with limited resources. The lecture is technical and aimed at graduate students, providing a comprehensive overview of LLM training practices.

183 words

Critical Evaluation

The lecture provides a solid, structured overview of LLM training, suitable for a graduate-level course. The instructors, Afshine and Shervine Amidi, demonstrate deep familiarity with the subject, presenting concepts in a logical progression from pretraining to fine-tuning. The content is accurate and reflects current best practices in the field, with references to well-known works such as the Chinchilla scaling laws and FlashAttention. The use of concrete examples (e.g., GPT-3 with 300B tokens, Llama 3 with 15T tokens) helps ground the theoretical concepts. The explanation of FLOPs vs. FLOPS is particularly useful, as these terms are often confused. The lecture also covers practical optimization techniques like ZeRO and mixed precision training, which are essential for training large models. However, as a lecture, it is not exhaustive; some topics are covered at a high level, and the mathematical derivations are simplified. The discussion of scaling laws is brief and does not delve into the nuances of different scaling law formulations. Similarly, the treatment of quantization and LoRA is introductory, without deep dives into the underlying algorithms. The lecture’s strength lies in its clarity and organization, making complex topics accessible. The sources cited are primarily the course syllabus and general references, rather than specific papers, which limits the ability to verify claims directly. Nonetheless, the content aligns with established knowledge in the field. The title accurately reflects the content, and the lecture fulfills its educational purpose effectively.

234 words

Title / Content Match

The title accurately reflects the content: a lecture on LLM training, covering pretraining, optimization, and fine-tuning.

Quality & Reliability

8/10

Lecture from Stanford University's CME295 course, delivered by adjunct lecturers with expertise in machine learning. Content is structured, covers established concepts (scaling laws, quantization, LoRA) and references well-known works (Chinchilla, FlashAttention). However, it is a lecture, not peer-reviewed, and some details may be simplified for teaching.

Chapters

Cited Sources

Concurring Sources

Contribution & Novelties

This lecture provides a comprehensive and up-to-date overview of LLM training, synthesizing key concepts from pretraining to parameter-efficient fine-tuning. It offers a clear framework for understanding the computational and data requirements of LLMs, and introduces practical optimization techniques. The lecture’s contribution lies in its pedagogical clarity, making advanced topics accessible to graduate students.

Pour aller plus loin :

144 words

Radar Profile

The radar profile shows high scores across all dimensions, indicating a well-rounded lecture with substantial information, good quality, technical depth, and reliability. The lecture excels in providing a comprehensive overview of LLM training, with strong emphasis on practical techniques and theoretical foundations.

Reliability 8/10