Towards building safe and secure AI: Lessons and Open Challenges

Towards building safe and secure AI: Lessons and Open Challenges

🎙 Dawn Song 👥 5K 📅 August 11, 2026 ⏱ 29 min 👁 8 📄 expert opinion 🧭 2026-08-15
Available in: English (current) Français

Keywords

AI safetyAI securityagentic AIcybersecurityLLM agents

Summary

Dawn Song, a professor at UC Berkeley, delivers a talk on building safe and secure AI, focusing on agentic AI systems. She emphasizes the importance of considering adversarial settings, as AI systems become more capable and integrated into real-world applications. The talk is divided into two main parts: securing agentic AI systems against attacks, and mitigating the misuse of frontier AI, particularly in cybersecurity. For securing agentic AI, she discusses the need for robust evaluation and risk assessment, highlighting automatic red-teaming frameworks developed by her group, such as LeakAgent, AV2, AgentVision, and Agent Exploits. She also stresses the importance of defense-in-depth and presents ProAgent, a programmable privilege control system. In the second part, she analyzes how frontier AI could reshape cybersecurity, presenting benchmarks like CyberGym and ExploitBench that measure AI capabilities in vulnerability discovery and exploit generation. She notes that attackers currently benefit more from AI than defenders, and advocates for proactive defense through secure-by-construction approaches and formal verification. She also introduces AgentBeast, a platform for standardized agent evaluation, and the Agent’s Last Exam benchmark for economically valuable tasks. The talk concludes with an invitation to the Agent Summit at UC Berkeley.

192 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable insights into the current landscape of AI security, particularly for agentic systems. It synthesizes multiple research efforts and benchmarks, offering a comprehensive overview of attack vectors and defense strategies. The argumentation is well-structured, moving from general risks to specific solutions and benchmarks. However, due to time constraints, many technical details are omitted, and the talk serves more as a high-level survey than an in-depth analysis. The speaker’s expertise lends credibility, but the lack of detailed evidence for some claims limits the depth of the argumentation.

Scientific Rigor, Source Quality, Title Accuracy

The talk demonstrates strong scientific rigor, referencing multiple peer-reviewed papers and benchmarks developed by the speaker’s group and collaborators. The sources are credible and directly relevant to the topic. The title accurately reflects the content, which covers lessons and open challenges in AI safety and security. The talk is well-organized and the speaker’s credentials enhance the reliability of the information presented.

165 words

Title / Content Match

The title accurately reflects the content, which covers lessons and open challenges in building safe and secure AI.

Quality & Reliability

8/10

The speaker is a renowned expert in AI security with extensive credentials. The talk presents a comprehensive overview of AI risks and security challenges, referencing multiple research works and benchmarks. However, it is a high-level talk with limited technical depth, and some claims are not fully substantiated in the video.

Key Moments

Cited Sources

  • International AI Report — Comprehensive overview of AI risks
  • Overview paper on attack and defense landscape of Agentic AI — Laying out the design space and security risks
  • LeakAgent — RL-based red-teaming agent for privacy leakage
  • AV2 — End-to-end red-teaming system using blackbox optimization
  • AgentVision — Automatic attacks on real-world web agents
  • Agent Exploits — Multi-agent system for automatic red-teaming of white-box agents
  • ProAgent — Programmable privilege control and guard for agents
  • CyberGym — Benchmark for AI capabilities in vulnerability discovery
  • ExploitBench — Benchmark for AI capabilities in exploit generation
  • AgentBeast — Platform for standardized agent evaluation
  • Agent's Last Exam — Benchmark for economically valuable tasks

Concurring Sources

  • International AI Report — Comprehensive overview of AI risks
  • Overview paper on attack and defense landscape of Agentic AI — Laying out the design space and security risks

Contribution & Novelties

The talk provides a comprehensive overview of the current state of AI safety and security, particularly for agentic AI systems. It highlights the importance of considering adversarial settings and presents several novel benchmarks and frameworks developed by the speaker’s group, such as CyberGym, ExploitBench, and AgentBeast. The talk also emphasizes the need for proactive defense through secure-by-construction approaches and formal verification.

Pour aller plus loin :

92 words

Radar Profile

The radar profile shows high scores in quality of information and global reliability, reflecting the speaker's expertise and the credibility of the sources. The quantity of information is moderate, as the talk covers many topics but at a high level. The technical level is moderate, suitable for a general technical audience but not deeply technical.

Reliability 8/10