[ИАД, весна 2026] Введение в машинное обучение. Лекция 7: Вероятностные модели машинного обучения

[ИАД, весна 2026] Введение в машинное обучение. Лекция 7: Вероятностные модели машинного обучения

🎙 Machine Learning – Intelligent Systems 👥 8K 📅 April 2, 2026 ⏱ 118 min 👁 144 📄 lecture 🧭 2026-08-16
Available in: English (current) Français

Keywords

maximum likelihoodEM algorithmGaussian mixtureBayes theoremunsupervised learning

Summary

This lecture, part of a course on machine learning, introduces probabilistic models, focusing on maximum likelihood estimation (MLE) and the Expectation-Maximization (EM) algorithm. The instructor begins by contextualizing the lecture within the broader course, referencing Pedro Domingos’ ‘five schools’ of machine learning and noting the recent success of large language models. The core mathematical foundation is presented: given a sample from an unknown distribution, the goal is to estimate the parameters of a parametric model by maximizing the likelihood. Two analytical examples are worked out: multivariate Gaussian density estimation and discrete distribution estimation, both yielding closed-form solutions. The lecture then addresses the more complex case of mixture models, where the likelihood contains a sum inside the logarithm, making analytical solution intractable. The EM algorithm is derived from Karush-Kuhn-Tucker conditions, revealing an iterative procedure with an E-step (computing posterior probabilities via Bayes’ theorem) and an M-step (maximizing weighted likelihood for each component). A visual demonstration on a 2D Gaussian mixture illustrates the algorithm’s convergence. The lecture concludes by discussing the versatility of mixture models and their connection to latent variable problems.

180 words

Critical Evaluation

Value of the Information & Strength of the Argument

The lecture provides substantial value by offering a rigorous, self-contained derivation of the EM algorithm from first principles, which is rare in introductory treatments. The argumentation is solid: the instructor carefully builds from the likelihood principle, shows the intractability of mixture likelihoods, and then derives the EM update equations using KKT conditions, clarifying the probabilistic interpretation of the auxiliary variables. The use of concrete examples (Gaussian and discrete distributions) and a visual demonstration enhances understanding. The pedagogical approach is effective, though the derivation is mathematically dense and may require prior knowledge of calculus and probability.

Scientific Rigor, Source Quality, Title Accuracy

The lecture demonstrates high scientific rigor: the mathematical derivations are correct and well-explained, and the instructor distinguishes between analytical and numerical solutions. However, no external sources are cited, which limits the verifiability of the content. The title accurately reflects the content, as it is indeed an introductory lecture on probabilistic models. The instructor’s expertise is evident, and the content aligns with standard statistical learning theory. The lack of citations is a minor weakness, but the lecture’s internal consistency and clarity compensate.

191 words

Title / Content Match

The title accurately reflects the content: an introductory lecture on probabilistic models in machine learning, specifically covering maximum likelihood estimation and the EM algorithm.

Quality & Reliability

8/10

The lecture is a formal academic presentation, mathematically rigorous, with clear derivations and references to standard concepts (MLE, EM algorithm). The instructor demonstrates deep expertise and provides a structured pedagogical approach. No external sources are cited, but the content aligns with established statistical learning theory.

Key Moments

Contribution & Novelties

The lecture provides a clear and rigorous derivation of the EM algorithm, emphasizing its probabilistic interpretation and connection to maximum likelihood estimation. It bridges the gap between theoretical foundations and practical application, making it valuable for students and practitioners.

Pour aller plus loin :

70 words

Radar Profile

The radar profile shows balanced scores across all dimensions, with slightly higher scores in technical level and reliability, reflecting the lecture's mathematical rigor and the instructor's expertise. The lower score in information quantity suggests a focused scope rather than a broad survey.

Reliability 8/10