[ИАД, весна 2026] Введение в машинное обучение. Лекция 5: Обучаемая векторизация данных

[ИАД, весна 2026] Введение в машинное обучение. Лекция 5: Обучаемая векторизация данных

🎙 Machine Learning – Intelligent Systems 👥 8K 📅 March 13, 2026 ⏱ 108 min 👁 153 📄 lecture 🧭 2026-08-16
Available in: English (current) Français

Keywords

PCASVDautoencodermatrix factorizationrecommender systems

Summary

This lecture, part of a series on machine learning, focuses on the concept of learnable data vectorization. The instructor begins by revisiting the principle of empirical induction and the minimization of empirical risk, then introduces PCA as a classical method for dimensionality reduction. He explains PCA through the lens of low-rank matrix factorization, connecting it to singular value decomposition (SVD) and showing how it can be seen as a constrained linear autoencoder. The lecture then broadens to general matrix factorization, discussing scenarios where SVD is not applicable, such as missing data and non-negativity constraints. Recommender systems are used as a motivating example, and the instructor demonstrates how stochastic gradient descent can be used to learn latent factor models. He also touches on regularization and the interpretation of latent factors as user interests. The lecture concludes by hinting at the progression to transformers and large language models, emphasizing the importance of learnable vectorization in modern deep learning.

156 words

Critical Evaluation

Value of the Information & Strength of the Argument

The lecture provides substantial value by connecting classical dimensionality reduction techniques (PCA, SVD) with modern deep learning approaches (autoencoders, transformers). The argumentation is solid, with clear mathematical derivations and intuitive explanations. The instructor effectively builds on previous lectures, creating a coherent narrative. He also engages the audience with questions, encouraging active thinking. The use of concrete examples, such as credit scoring and recommender systems, helps illustrate abstract concepts. The discussion of limitations and extensions (e.g., non-negative matrix factorization) adds depth. Overall, the lecture is informative and well-argued, though it assumes prior knowledge of linear algebra and basic ML concepts.

Scientific Rigor, Source Quality, Title Accuracy

The lecture demonstrates scientific rigor through precise mathematical formulations and references to established methods (PCA, SVD, autoencoders). However, it lacks explicit citations to external sources, relying instead on the instructor’s expertise. The title accurately reflects the content, focusing on learnable vectorization. The lecture is part of a structured course, suggesting a systematic approach. The instructor mentions a reference to Chechotka (likely a misspelling of Cichocki) for non-negative matrix factorization, but no specific URLs are provided. The content is consistent with standard ML literature, but without external verification, the overall reliability is moderate.

206 words

Title / Content Match

The title accurately reflects the content: the lecture focuses on learnable data vectorization, covering PCA, autoencoders, and matrix factorization, culminating in transformers and large language models.

Quality & Reliability

8/10

The lecture is a well-structured academic presentation, covering classical methods (PCA, SVD) and modern extensions (autoencoders, matrix factorization) with mathematical rigor. The instructor demonstrates deep knowledge and provides clear derivations. However, it is a single lecture without external citations or peer-reviewed references, and the content is not independently verified.

Key Moments

Cited Sources

  • Chechotka (likely Cichocki) on non-negative matrix factorization — Mentioned in the context of non-negative matrix factorization literature.

Concurring Sources

Contribution & Novelties

The lecture provides a comprehensive overview of learnable data vectorization, bridging classical methods (PCA, SVD) with modern deep learning (autoencoders, transformers). It emphasizes the conceptual shift from fixed feature extraction to learned representations. The discussion of matrix factorization as a unifying framework is particularly insightful, showing how PCA is a special case of autoencoders and how recommender systems can be modeled as latent factor models. The lecture also highlights practical considerations such as missing data and non-negativity constraints, which are often overlooked in introductory treatments.

Pour aller plus loin :

146 words

Radar Profile

The radar profile shows high scores in quantity of information and technical level, indicating a dense, advanced lecture. Quality of information and global reliability are slightly lower, reflecting the lack of external citations and the lecture's informal nature. The overall balance suggests a strong educational resource for those with prior ML knowledge.

Reliability 7/10