[M2L 2025] 3.4 AI Safety & Ethics in Practice - Julia Haas

[M2L 2025] 3.4 AI Safety & Ethics in Practice - Julia Haas

🎙 Julia Haas 👥 3K 📅 November 12, 2025 ⏱ 48 min 👁 55 📄 expert opinion 🧭 2026-08-15
Available in: English (current) Français

Keywords

AI safetyethicsLLMtrolley problemOpenAI model spec

Summary

Julia Haas, a researcher at DeepMind, presents a practical session on AI safety and ethics, focusing on the application of trolley problems to large language models (LLMs). She begins by explaining her background in studying human moral cognition and how it led her to reinforcement learning and eventually to DeepMind, where she now works on measuring normative capabilities in LLMs. The core of the session involves examining real-world ’trolley problems’ that have surfaced on Reddit, such as comparing Elon Musk’s tweets to Hitler’s impact or asking whether one should misgender Caitlyn Jenner to prevent a nuclear apocalypse. Haas highlights the challenges these cases pose for LLMs, including the tendency to ‘both sides’ the issue or to refuse to answer, which is often criticized. She then discusses how OpenAI addresses these problems through its model spec, which prioritizes helpfulness, minimizing harm, and following a chain of command, leading to unambiguous answers even in controversial cases. The session includes a group activity where participants are asked to design their own ethical principles for LLM responses, and the results are discussed, revealing diverse approaches. Haas emphasizes the difficulty of balancing competing principles and the importance of explicit reasoning in AI ethics. The talk concludes with a discussion on the broader implications for AI safety and the need for careful consideration of ethical frameworks in AI development.

223 words

Critical Evaluation

Value of the Information & Strength of the Argument

The video provides valuable insights into the practical challenges of AI ethics, particularly in the context of LLMs. Haas effectively uses real-world examples to illustrate the complexity of ethical decision-making in AI, and she encourages critical thinking by involving the audience in a hands-on exercise. The argumentation is solid, drawing on her expertise and referencing OpenAI’s model spec as a concrete example. However, the talk is more of an expert opinion than a rigorous academic analysis, and some points could benefit from more detailed evidence or citations. The interactive format adds value but also limits the depth of the discussion.

Scientific Rigor, Source Quality, Title Accuracy

The talk demonstrates scientific rigor through its structured approach and reference to OpenAI’s model spec, which is a credible source. However, the speaker does not provide a comprehensive list of sources, and some claims are based on anecdotal evidence from Reddit. The title accurately reflects the content, and the session is well-organized. The lack of formal citations is a minor weakness, but the speaker’s expertise and the practical nature of the talk compensate for this.

190 words

Title / Content Match

The title accurately reflects the content: a practical session on AI safety and ethics, led by Julia Haas, focusing on trolley problems and their application to LLMs.

Quality & Reliability

8/10

The speaker is a researcher at DeepMind with a background in moral cognition and reinforcement learning, providing expert insights into AI safety and ethics. The content is well-structured, grounded in real cases and references to OpenAI's model spec, and encourages critical thinking. However, it is primarily an opinion/expert talk rather than a peer-reviewed study, and some claims are not fully sourced.

Key Moments

Cited Sources

  • OpenAI Model Spec — Referenced as the framework OpenAI uses to govern model behavior, including handling of trolley problems.

Concurring Sources

  • OpenAI Model Spec — The talk's discussion of OpenAI's approach aligns with the content of the model spec.

Contribution & Novelties

The video offers a unique perspective on AI ethics by focusing on the practical application of trolley problems to LLMs, highlighting the challenges of balancing competing principles. It provides a hands-on exercise that encourages critical thinking about ethical frameworks. The discussion of OpenAI’s model spec offers a concrete example of how AI labs address these issues.

Pour aller plus loin :

108 words

Radar Profile

The radar profile shows high scores in information quantity, quality, and reliability, with a slightly lower technical level, reflecting the talk's balance between expert insight and accessibility. The overall high scores indicate a valuable resource for understanding AI ethics in practice.

Reliability 8/10

💬 No comments were provided for analysis.