Stanford CS229 Machine Learning | Spring 2026 | Lecture 9: K-Means and GMM (non-EM)

Stanford CS229 Machine Learning | Spring 2026 | Lecture 9: K-Means and GMM (non-EM)

🎙 Chris Ré, Tengyu Ma 👥 1.2M 📅 July 31, 2026 ⏱ 76 min 👁 645 📄 lecture 🧭 2026-08-03
Available in: English (current) Français

Keywords

k-meansGMMclusteringunsupervisedEM

Summary

This lecture from Stanford CS229 introduces two fundamental unsupervised learning algorithms: k-means and Gaussian Mixture Models (GMM). The instructor, Chris Ré, begins by contrasting supervised and unsupervised learning, emphasizing the challenges of clustering without labels. He then presents the k-means algorithm, explaining its iterative nature: initializing cluster centers, assigning points to the nearest center, and recomputing centers as the mean of assigned points. He discusses convergence and the NP-hardness of finding the optimal clustering. Next, he introduces GMM as a probabilistic extension of k-means, where each cluster is modeled as a Gaussian distribution. The lecture sets the stage for the Expectation-Maximization (EM) algorithm, which is used to estimate parameters in such models. The instructor highlights the importance of modeling assumptions and the trade-offs between stronger assumptions and weaker guarantees in unsupervised learning. The lecture includes interactive questions and visual explanations, aiming to build intuition behind these core algorithms.

148 words

Critical Evaluation

The lecture provides a solid introduction to k-means and Gaussian Mixture Models, two cornerstone unsupervised learning algorithms. The instructor, Chris Ré, is a renowned computer science professor, and the content aligns with the standard CS229 curriculum. The presentation is clear and pedagogical, using visual examples and interactive questions to build intuition. The mathematical foundations are presented rigorously, with careful explanations of the objective functions and update rules. The discussion of the NP-hardness of k-means and the trade-offs between supervised and unsupervised learning adds depth. The lecture also sets up the EM algorithm, which is crucial for GMM, though it is not covered in detail in this session. The sources cited are the official CS229 course page and Stanford’s AI program page, which are authoritative. However, the lecture is introductory and does not delve into advanced topics or recent research. The adéquation between title and content is excellent. Overall, this is a high-quality educational resource, though it is not a research presentation. The lack of peer review is compensated by the institutional credibility. The lecture’s value lies in its clarity and pedagogical effectiveness, making it suitable for students and practitioners seeking a solid foundation in clustering.

195 words

Title / Content Match

Titre clair et précis, correspondant exactement au contenu de la leçon.

Quality & Reliability

8/10

Lecture from Stanford CS229, taught by renowned professors, covering established algorithms (k-means, GMM) with rigorous mathematical foundation. Content is accurate and well-structured, though it is a lecture and not peer-reviewed.

Key Moments

Cited Sources

Concurring Sources

Contribution & Novelties

This lecture provides a clear and rigorous introduction to k-means and GMM, emphasizing the modeling assumptions and trade-offs in unsupervised learning. It effectively sets up the EM algorithm, which is essential for many latent variable models. The pedagogical approach, with visual examples and interactive questions, enhances understanding.

Pour aller plus loin :

82 words

Radar Profile

The radar profile shows high scores in quality and technical level, with slightly lower scores in quantity and reliability, reflecting the lecture's depth and authoritative source but limited scope and lack of peer review.

Reliability 8/10