Generative AI L25: GPT pre-training objective, span corruption, comparisons of various models

Generative AI L25: GPT pre-training objective, span corruption, comparisons of various models

🎙 Agha Ali Raza 👥 3K 📅 May 20, 2026 ⏱ 51 min 👁 58 📄 lecture 🧭 2026-08-15
Available in: English (current) Français

Keywords

GPTT5span corruptionpre-trainingmodel comparison

Summary

This lecture, part of a graduate course on Generative AI, focuses on pre-training objectives for transformer models. The instructor begins by recapping the MLM objective used in BERT, explaining the masking strategy and the loss calculation on all corrupted tokens. He then introduces the T5 model, which uses span corruption as a pre-training objective, where entire spans of text are replaced with sentinel tokens and the model must predict the missing spans. The lecture also covers decoder-only models like GPT, emphasizing their causal language modeling objective. A significant portion is dedicated to comparing various models (BERT, T5, GPT) in terms of architecture, pre-training data size, and computational cost. The instructor uses an analogy to help students appreciate the scale of pre-training data, calculating that reading LLaMA-1’s 1.4 trillion tokens would take a human 7,991 years. He also discusses the strengths of encoder-only, decoder-only, and encoder-decoder architectures, and touches on the importance of instruction tuning for models to follow prompts.

159 words

Critical Evaluation

Value of the Information & Strength of the Argument

The lecture provides valuable insights into the pre-training objectives of major transformer models, with clear explanations and concrete examples. The instructor effectively argues for the efficiency of span corruption over full text prediction, citing empirical results from the T5 paper. He also emphasizes the scale of pre-training data, using a compelling analogy to make the numbers tangible. The argumentation is solid, grounded in established research, and the instructor’s teaching style is engaging.

Scientific Rigor, Source Quality, Title Accuracy

The lecture demonstrates scientific rigor by referencing original papers (T5, BERT, GPT) and providing detailed technical explanations. The instructor is transparent about assumptions (e.g., reading speed, token-to-word ratio) when making calculations. The title accurately reflects the content, covering the stated topics. The course materials and slides are available online, enhancing credibility. No comments were provided for analysis.

144 words

Title / Content Match

The title accurately reflects the content, covering GPT pre-training objectives, span corruption, and model comparisons.

Quality & Reliability

8/10

Lecture from a graduate course at LUMS, with clear explanations of technical concepts, references to original papers (T5, BERT, GPT), and detailed data on model training. The instructor is an academic, and the content is well-structured. Minor limitations: no external citations in the video itself, but slides and course materials are provided.

Chapters

Cited Sources

Concurring Sources

  • T5 paper — The lecture's explanation of span corruption aligns with the T5 paper's methodology.
  • BERT paper — The lecture's description of MLM and NSP matches the original BERT paper.

Contribution & Novelties

This lecture provides a clear and detailed explanation of span corruption as used in T5, contrasting it with BERT’s MLM and BART’s full text prediction. It also offers a unique perspective on the scale of pre-training data, using a human reading analogy to make the numbers relatable. The comparison of various models (BERT, T5, GPT) in terms of architecture and data requirements is valuable for understanding the evolution of transformer models.

Pour aller plus loin :

  • T5 paper (Raffel et al., 2019) — The original paper introducing the text-to-text framework and span corruption.
  • BERT paper (Devlin et al., 2018) — The original paper introducing MLM and NSP.
  • GPT-2 paper (Radford et al., 2019) — The paper introducing the decoder-only architecture and large-scale pre-training.

123 words

Radar Profile

The radar profile shows high scores across all dimensions, indicating a well-rounded lecture with substantial information, high technical depth, and strong reliability. The balance between quantity and quality of information is notable, making it a valuable resource for understanding pre-training objectives.

Reliability 8/10