Reinforcement Learning 2026 - Session 9

Reinforcement Learning 2026 - Session 9

🎙 Robust and Interpretable Machine Learning Lab 👥 1K 📅 July 12, 2026 ⏱ 86 min 👁 7 📄 lecture 🧭 2026-08-16
Available in: English (current) Français

Keywords

policy gradientvariance reductioncausality trickbaselineactor-critic

Summary

This session of a reinforcement learning course begins with a clarification on the distinction between Double Q-learning and Double Deep Q-Network (Double DQN). The instructor explains that Double Q-learning uses two separate models, while Double DQN uses a target network (theta-minus) that is a lagged copy of the main network. The main topic is variance reduction in policy gradient methods. The instructor reviews the policy gradient objective and its gradient, emphasizing the high variance of Monte Carlo return estimates. They introduce four techniques to reduce variance: causality trick, discount factor, baseline, and actor-critic. The causality trick involves using only future rewards for each action, justified by showing that past rewards have zero expected value in the gradient. The instructor provides a theoretical derivation using conditional expectations and the log-derivative trick. They also discuss the variance reduction property of subtracting a baseline, using the identity Var(X+W) = Var(X) + Var(W) + 2Cov(X,W). The session ends with a discussion on the independence of the baseline and the gradient term.

167 words

Critical Evaluation

Value of the Information & Strength of the Argument

The lecture provides a solid theoretical foundation for variance reduction in policy gradient methods. The instructor clearly explains the motivation and derives the causality trick step-by-step, using conditional expectations and the log-derivative trick. The argumentation is rigorous, with a clear connection to the variance reduction property. The clarification on Double Q-learning vs Double DQN adds value by correcting a common misconception.

Scientific Rigor, Source Quality, Title Accuracy

The instructor references the Double DQN paper and explains its relation to Double Q-learning. The derivation of the causality trick is mathematically sound. The title accurately reflects the content. No external sources are cited beyond the paper mentioned.

114 words

Title / Content Match

The title accurately reflects the content, which is a session on reinforcement learning.

Quality & Reliability

8/10

The lecture is based on established reinforcement learning theory, with clear derivations and references to the Double DQN paper. The instructor corrects a previous misconception, demonstrating scientific rigor. However, the video has very low viewership and no external validation.

Key Moments

Cited Sources

  • Double DQN paper — Referenced for the distinction between Double Q-learning and Double DQN

Concurring Sources

  • Double DQN paper — The lecture's explanation aligns with the paper's content.

Contribution & Novelties

The lecture provides a clear and rigorous derivation of the causality trick for variance reduction in policy gradient methods. It also clarifies the difference between Double Q-learning and Double DQN, which is often confused. The session is part of a course, so it offers pedagogical value.

Pour aller plus loin :

75 words

Radar Profile

The radar profile shows high scores in quantity and quality of information, with a moderate level of technical depth. The overall reliability is good, but the low viewership and lack of external validation slightly reduce the score.

Reliability 7/10