Generative AI L9: Skipgram, CBOW (full softmax variant, negative sampling variant), GloVe

Generative AI L9: Skipgram, CBOW (full softmax variant, negative sampling variant), GloVe

🎙 Agha Ali Raza 👥 3K 📅 May 9, 2026 ⏱ 71 min 👁 95 📄 lecture 🧭 2026-08-15
Available in: English (current) Français

Keywords

word2vecskipgramCBOWnegative samplingGloVe

Summary

This lecture, part of the ‘Foundations of Generative AI’ course at LUMS, provides a comprehensive overview of word embedding models. The instructor begins with a recap of data generation for training word-to-vector models, including sliding window protocols and subsampling. He then explains the skipgram model with full softmax, detailing the neural network architecture, the objective function, and the computational expense due to the softmax denominator. He contrasts this with the CBOW model, which predicts the center word from context words, and introduces the concept of aggregating context embeddings. The lecture then transitions to the negative sampling variant, which addresses the computational inefficiency of full softmax by approximating the denominator. Finally, the instructor introduces GloVe embeddings, which are based on global word-word co-occurrence statistics. Throughout, he emphasizes the mathematical formulations and practical considerations, and provides examples and notes for further study.

140 words

Critical Evaluation

Value of the Information & Strength of the Argument

The lecture provides a solid theoretical foundation for understanding word embedding models. The instructor carefully derives the objective functions and explains the computational challenges, making a strong argument for the need for negative sampling. He also discusses the linearity of the hidden layer and the reasoning behind it, which is often glossed over. The comparison between skipgram and CBOW is clear, and the introduction of GloVe adds a broader perspective. The argumentation is logical and well-structured, with a focus on both intuition and mathematical rigor.

Scientific Rigor, Source Quality, Title Accuracy

The lecture is scientifically rigorous, with clear derivations and explanations. The instructor references the original papers and course materials, which are available online. The title accurately reflects the content, covering all the mentioned models. The video is part of a structured course, and the instructor’s expertise is evident. The sources cited include the course website and playlist, which provide access to slides and assessments. The lecture does not cite external sources beyond the course materials, but the content is consistent with established literature on word embeddings.

186 words

Title / Content Match

The title accurately reflects the content: the lecture covers Skipgram, CBOW (both full softmax and negative sampling variants), and GloVe embeddings.

Quality & Reliability

8/10

Lecture from a graduate course at LUMS, covering theoretical foundations and practical variants of word embedding models. The content is structured, mathematically rigorous, and includes derivations and comparisons. The instructor is an academic, and the course materials are openly available. The video is part of a series, and the presentation is clear, though the audio is in Urdu/English mix, which may limit accessibility.

Chapters

Cited Sources

Concurring Sources

  • Word2Vec paper — The lecture's content on skipgram and CBOW aligns with the original Word2Vec paper.
  • GloVe paper — The lecture's introduction to GloVe matches the original paper's approach.

Contribution & Novelties

The lecture provides a thorough pedagogical walkthrough of word embedding models, emphasizing the mathematical derivations and computational trade-offs. It clarifies the differences between full softmax and negative sampling, and introduces GloVe as an alternative approach. The instructor’s explanations of the linear hidden layer and the aggregation in CBOW are particularly insightful.

Pour aller plus loin :

97 words

Radar Profile

The radar profile shows high scores in technical level and information quality, indicating a rigorous and detailed lecture. The quantity of information is also high, but the global reliability is slightly lower, possibly due to the informal presentation style and the mix of languages.

Reliability 8/10