[ИАД, осень 2025] Вероятностные тематические модели. Лекция 7

[ИАД, осень 2025] Вероятностные тематические модели. Лекция 7

🎙 Konstantin Vorontsov 👥 8K 📅 October 24, 2025 ⏱ 102 min 👁 140 📄 lecture 🧭 2026-08-16
Available in: English (current) Français

Keywords

topic modelingmultilingualmultimodalEM algorithmregularization

Summary

This lecture, part of a course on probabilistic topic models, focuses on extending standard topic models to handle multimodal data, particularly multilingual collections. The instructor begins by reviewing the basic EM algorithm for topic modeling, then introduces the concept of modalities and how to incorporate them by weighting term frequencies. He then discusses multilingual topic models, where documents in different languages are treated as parallel texts and combined into a single document. He presents two regularization approaches: one using a simple dictionary-based regularizer that encourages topic distributions of translation pairs to be similar, and a more advanced model that introduces a topic-dependent translation probability matrix. The latter is shown to yield a Bayes-consistent formula for aligning translation probabilities across languages. Experimental results on a small Wikipedia collection demonstrate that simply combining parallel texts is the most effective strategy, outperforming dictionary-based regularization, and that a small fraction of parallel texts is sufficient. The lecture also touches on the use of topic models for cross-lingual search and the practical challenges of handling 100 languages, including vocabulary reduction via BPE tokenization. Finally, the instructor briefly introduces the idea of a generative modality, such as categories or authors, which can be incorporated into the probabilistic space to model document generation more richly.

208 words

Critical Evaluation

Value of the Information & Strength of the Argument

The lecture provides a rigorous and detailed exposition of multimodal and multilingual topic models. The instructor carefully derives the EM updates for the proposed regularizers, showing how they lead to interpretable formulas. He supports his claims with experimental evidence, comparing different models and showing the superiority of parallel texts. The argumentation is solid, building on established theory and clearly explaining the intuition behind each model. The value lies in the practical guidance for building multilingual topic models and the insights into the relative importance of parallel data versus dictionaries.

Scientific Rigor, Source Quality, Title Accuracy

The lecture references several sources, including a survey by Ivan Vulić (2015) and a paper by Suvorova et al. (2020) on probabilistic topic models with topic-dependent translation probabilities. The instructor also mentions a study with Antiplagiat company on multilingual scientific text classification, though the main paper was not published. The title accurately reflects the content, and the lecture is well-structured. The instructor acknowledges limitations, such as the small scale of experiments and the lack of follow-up work. Overall, the scientific rigor is high, with clear derivations and honest reporting of results.

195 words

Title / Content Match

The title accurately reflects the content: a lecture on probabilistic topic models, specifically focusing on multimodal and multilingual extensions.

Quality & Reliability

8/10

The lecture is given by a recognized expert in topic modeling, presents formal derivations and references to published work, and includes experimental results. However, some references are not fully specified, and the lecture is a recording of a course session.

Key Moments

Cited Sources

  • Vulić, I. (2015). A Survey of Multilingual Topic Models — Referenced as a comprehensive overview of multilingual topic models.
  • Suvorova, M., et al. (2020). Probabilistic topic models with topic-dependent translation probabilities — Referenced for experiments on topic-dependent translations.
  • AntiPlagiat company research on multilingual scientific text classification — Mentioned as a collaborative project on multilingual search.

Concurring Sources

  • Vulić, I. (2015). A Survey of Multilingual Topic Models — Supports the claim that parallel texts are sufficient for multilingual topic models.
  • Suvorova, M., et al. (2020). Probabilistic topic models with topic-dependent translation probabilities — Provides experimental evidence for the proposed model.

Contribution & Novelties

The lecture provides a clear and thorough exposition of multilingual topic models, highlighting the surprising effectiveness of simply combining parallel texts over dictionary-based regularization. It also introduces a novel regularizer that models topic-dependent translation probabilities, leading to a Bayes-consistent alignment of translation probabilities across languages. This approach offers a principled way to incorporate dictionaries and can be useful for linguists and translation studies.

Pour aller plus loin :

105 words

Radar Profile

The radar profile shows high scores across all dimensions, indicating a well-balanced and informative lecture. The strongest aspects are the quantity and quality of information, as well as the technical depth, reflecting the formal and rigorous treatment of the subject.

Reliability 8/10