Augustine Mavor Parker on the noisy tv problem | FAI CDT

Augustine Mavor Parker on the noisy tv problem | FAI CDT

🎙 Augustine Mavor-Parker 👥 3K 📅 November 5, 2025 ⏱ 42 min 👁 96 📄 expert opinion 🧭 2026-08-15
Available in: English (current) Français

Keywords

reinforcement learningcuriosityexplorationaleatoric uncertaintyintrinsic rewards

Summary

In this interview, Augustine Mavor-Parker, a recent PhD graduate from UCL’s Foundational AI CDT, discusses his research on sample-efficient reinforcement learning (RL). He explains the challenge of sparse rewards in games like Montezuma’s Revenge, where agents receive no feedback for most actions. To address this, he explores intrinsic reward functions based on curiosity, where agents are rewarded for prediction errors, encouraging exploration. However, this approach suffers from the ’noisy TV problem’: agents can become trapped by unpredictable, stochastic elements in the environment, endlessly seeking novelty without making progress. Mavor-Parker’s key contribution is a method that estimates aleatoric uncertainty (irreducible randomness) and subtracts it from the curiosity reward, preventing agents from being distracted by noisy TVs. He validates this approach with experiments where agents can choose to watch random images instead of playing the game, showing his method avoids this distraction. He also discusses his internships: at Capernic AI, he worked on depth perception for spherical images in VR; at Oxford with Claire Lyle, he studied sinusoidal activation functions for RL generalization; and at Redwood Research, he worked on automating circuit discovery in large language models for interpretability. The interview provides insights into the challenges and practical applications of RL research.

200 words

Critical Evaluation

Value of the Information & Strength of the Argument

The interview provides valuable insights into the noisy TV problem and a novel solution based on aleatoric uncertainty estimation. Mavor-Parker clearly explains the intuition behind curiosity-driven exploration and the failure mode of noisy TVs. He argues convincingly that subtracting aleatoric uncertainty from the intrinsic reward can prevent agents from being distracted by unpredictable elements. The discussion of his internships adds practical context, showing how RL techniques are applied in industry and safety research. The argumentation is coherent and well-supported by his own research and experiments, though the informal format limits the depth of technical detail.

Scientific Rigor, Source Quality, Title Accuracy

The speaker is a credible researcher with a PhD from UCL, and his work is published in a paper titled ‘How to Stay Curious While Avoiding Noisy TVs Using Aleatoric Uncertainty Estimation’. He references relevant prior work, such as DeepMind’s Deep Q-Learning and curiosity-driven approaches, and discusses collaborations with established researchers like Claire Lyle. The title accurately reflects the content, focusing on the noisy TV problem and his other projects. The interview is scientifically rigorous, though it lacks formal citations or references to specific papers beyond his own work.

199 words

Title / Content Match

The title accurately reflects the main topic discussed, which is the noisy TV problem in reinforcement learning, along with the speaker's other PhD projects.

Quality & Reliability

8/10

The speaker is a PhD graduate in reinforcement learning, and the content is based on his own research and internships at reputable institutions. The explanations are technically accurate and well-articulated, though the format is an informal interview rather than a peer-reviewed presentation.

Key Moments

Cited Sources

  • How to Stay Curious While Avoiding Noisy TVs Using Aleatoric Uncertainty Estimation — Mavor-Parker's paper on the noisy TV problem, mentioned in the description.

Concurring Sources

Contribution & Novelties

The interview highlights Mavor-Parker’s original contribution of using aleatoric uncertainty estimation to mitigate the noisy TV problem in curiosity-driven RL. This is a novel approach that distinguishes between irreducible randomness and epistemic uncertainty, allowing agents to avoid being trapped by unpredictable elements. The discussion also touches on his work in AI safety and interpretability, providing a broader perspective on RL research.

Pour aller plus loin :

110 words

Radar Profile

The radar profile shows high scores in quality of information and technical level, indicating a technically sound and informative interview. The quantity of information is moderate, and the global reliability is high, reflecting the speaker's expertise. The overall assessment is positive, with a note of 4 out of 5.

Reliability 8/10

💬 No comments were provided for analysis.