Why do AI models need to be safe?

Why do AI models need to be safe?

🎙 IBM Research 👥 120K 📅 September 26, 2025 ⏱ 27 min 👁 944 📄 expert opinion 🧭 2026-08-16
Available in: English (current) Français

Keywords

AI safetygenerative AIsteerabilityalignmentrisk management

Summary

In this interview, IBM Fellow Kush Varshney discusses the importance of AI safety, emphasizing the need to reduce harm in AI systems. He explains IBM’s approach, which involves mapping, measuring, and managing risks. The conversation covers the creation of a risk atlas to catalog potential harms, the Granite Guardian model for detecting risks, and the distinction between alignment (the goal) and steerability (the methods to achieve it). Varshney highlights various steering techniques, including fine-tuning, prompt engineering, activation steering, and decoding adjustments. He introduces the concept of generative computing, which integrates generative AI into broader programming paradigms, using intrinsic functions to call AI only when appropriate. The discussion also touches on the challenges of keeping pace with rapid innovation, the importance of customization for different contexts, and the future of AI safety research, particularly in the realm of autonomous agents. Throughout, Varshney stresses the need for responsible innovation and the balance between speed and control.

154 words

Critical Evaluation

Value of the Information & Strength of the Argument

The video provides valuable insights into IBM’s practical approaches to AI safety, such as the risk atlas and Granite Guardian, which are not widely known. The argumentation is coherent, with Varshney clearly explaining concepts like alignment vs. steerability and the rationale behind generative computing. However, the discussion remains at a high level, with limited technical depth or empirical evidence. The value lies in the expert perspective and the introduction of IBM’s tools, but the argumentation could be strengthened with more concrete examples or data.

93 words

Title / Content Match

The title accurately reflects the content, which focuses on the necessity of AI safety and IBM's approaches to it.

Quality & Reliability

8/10

The video features an IBM Fellow discussing AI safety, drawing on IBM's research and products. It provides a high-level overview with some technical depth, but lacks detailed citations or empirical evidence. The information is credible given the speaker's expertise, but it is primarily an expert opinion rather than a rigorous scientific presentation.

Key Moments

Cited Sources

Concurring Sources

Contribution & Novelties

The video offers a unique perspective from an IBM Fellow on AI safety, highlighting IBM’s specific tools and frameworks such as the risk atlas, Granite Guardian, and the AI Steerability 360 toolkit. It introduces the concept of generative computing and intrinsic functions, which are not widely discussed in mainstream AI discourse. The discussion of steerability as a spectrum and the importance of agency in AI-human collaboration adds depth to the conversation.

Pour aller plus loin :

122 words

Radar Profile

The radar profile shows high scores in quality and reliability, reflecting the expert nature of the content, but lower scores in quantity and technical level, indicating that the video provides a broad overview rather than deep technical details.

Reliability 8/10