
Day 3: Shafi Goldwasser - Trustworthy AI: Robustness and Alignment | ADIA Lab Symposium 2025
Keywords
Summary
130 words
Critical Evaluation
Value of the Information & Strength of the Argument
The talk provides valuable insights into the formal foundations of trustworthy AI, particularly the application of cryptographic methods to machine learning. Goldwasser’s argumentation is rigorous, grounded in her own research and established theoretical frameworks. She clearly explains the limitations of current approaches and the need for provable guarantees. The discussion of backdoors and their undetectability is compelling, as is the proposed defense via post-processing. The argument that filtering is fundamentally flawed is well-illustrated with a clever construction. Overall, the value lies in bridging cryptography and AI, offering a fresh perspective on robustness and alignment.
Scientific Rigor, Source Quality, Title Accuracy
The talk is scientifically rigorous, referencing several published papers (e.g., on undetectable backdoors and oblivious defense) and the work of others (e.g., Boaz Barak on inference-time compute). However, the transcript does not include explicit citations or URLs, so the sources are not directly verifiable from the video. The title accurately reflects the content, focusing on robustness and alignment. The speaker’s authority and the technical depth contribute to high reliability, but the lack of detailed references in the talk itself slightly reduces the score.
192 words
Title / Content Match
The title accurately reflects the content: a focused talk on robustness and alignment in trustworthy AI, delivered by Shafi Goldwasser at the ADIA Lab Symposium.
Quality & Reliability
8/10
Presentation by a Turing Award laureate with deep expertise in cryptography and AI security. The talk is based on published research, but it is a high-level overview without detailed proofs or citations in the transcript. The speaker clearly distinguishes between established results and open questions.
Key Moments
Markers derived by PSI from the transcript: the creator did not define chapters.
- Introduction by host, highlighting Shafi Goldwasser's credentials.
- Goldwasser poses critical questions about AI reliability and control.
- Overview of trustworthy AI challenges: privacy, verification, robustness, alignment, model stealing.
- Cryptographic recipe for trustworthy AI: model adversary, define trust, build solution, prove.
- Discussion of robustness to distribution shifts, adversarial examples, and insider adversaries.
- Presentation of undetectable backdoors in neural networks and their implications.
- Defense against backdoors via random self-reducibility and post-processing.
- Introduction to alignment and the challenge of formalizing human values.
- Discussion of jailbreaks and the inadequacy of simple filtering.
- Cryptographic impossibility result for filtering harmful prompts.
Cited Sources
- Planting Undetectable Backdoors in Machine Learning Models — Discussed in the talk as a method to plant undetectable backdoors.
- Oblivious Defense in Machine Learning Models — Discussed as a defense against backdoors using post-processing.
- Puzzle Jailbreaking LLMs through Word Based Puzzles — Mentioned as an example of jailbreaking via puzzles.
Concurring Sources
- Turing Award — Shafi Goldwasser is a recipient, supporting her expertise.
Contribution & Novelties
The talk offers a unique cryptographic perspective on trustworthy AI, emphasizing worst-case adversarial modeling and formal proofs. It presents novel results on undetectable backdoors and defenses, and highlights the fundamental impossibility of simple filtering for alignment. The discussion of inference-time compute as a safety measure is timely.
Pour aller plus loin :
- Cryptography and Machine Learning — Background on cryptographic principles applied to AI.
- Adversarial Machine Learning — Overview of attacks and defenses.
- AI Alignment — General concept of aligning AI with human values.
84 words
Radar Profile
The radar profile shows high scores in quality and reliability, reflecting the speaker's authority and rigorous approach. The quantity of information is moderate, as the talk is focused and not exhaustive. The technical level is high, suitable for an expert audience.