Day 3: Shafi Goldwasser - Trustworthy AI: Robustness and Alignment | ADIA Lab Symposium 2025

Day 3: Shafi Goldwasser - Trustworthy AI: Robustness and Alignment | ADIA Lab Symposium 2025

🎙 Shafi Goldwasser 👥 824 📅 November 5, 2025 ⏱ 29 min 👁 110 📄 expert opinion 🧭 2026-08-16
Available in: English (current) Français

Keywords

robustnessalignmentbackdoorsjailbreakscryptography

Summary

Shafi Goldwasser, a Turing Award laureate, delivers a technical talk on trustworthy AI, focusing on robustness and alignment. She begins by highlighting the paradox: AI is widely deployed despite our lack of understanding and control. She then outlines key challenges: privacy, verification, robustness, alignment, and model stealing. Using a cryptographic perspective, she emphasizes worst-case adversarial modeling and formal proofs. She discusses her work on planting undetectable backdoors in neural networks, showing that any model can be compromised, and then presents a defense method based on random self-reducibility and post-processing. For alignment, she explains the difficulty of formalizing human values and shows that simple input/output filtering is insufficient, as adversaries can craft prompts that bypass filters. She concludes by suggesting that inference-time computation and cryptographic techniques may offer more robust solutions.

130 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides valuable insights into the formal foundations of trustworthy AI, particularly the application of cryptographic methods to machine learning. Goldwasser’s argumentation is rigorous, grounded in her own research and established theoretical frameworks. She clearly explains the limitations of current approaches and the need for provable guarantees. The discussion of backdoors and their undetectability is compelling, as is the proposed defense via post-processing. The argument that filtering is fundamentally flawed is well-illustrated with a clever construction. Overall, the value lies in bridging cryptography and AI, offering a fresh perspective on robustness and alignment.

Scientific Rigor, Source Quality, Title Accuracy

The talk is scientifically rigorous, referencing several published papers (e.g., on undetectable backdoors and oblivious defense) and the work of others (e.g., Boaz Barak on inference-time compute). However, the transcript does not include explicit citations or URLs, so the sources are not directly verifiable from the video. The title accurately reflects the content, focusing on robustness and alignment. The speaker’s authority and the technical depth contribute to high reliability, but the lack of detailed references in the talk itself slightly reduces the score.

192 words

Title / Content Match

The title accurately reflects the content: a focused talk on robustness and alignment in trustworthy AI, delivered by Shafi Goldwasser at the ADIA Lab Symposium.

Quality & Reliability

8/10

Presentation by a Turing Award laureate with deep expertise in cryptography and AI security. The talk is based on published research, but it is a high-level overview without detailed proofs or citations in the transcript. The speaker clearly distinguishes between established results and open questions.

Key Moments

Cited Sources

  • Planting Undetectable Backdoors in Machine Learning Models — Discussed in the talk as a method to plant undetectable backdoors.
  • Oblivious Defense in Machine Learning Models — Discussed as a defense against backdoors using post-processing.
  • Puzzle Jailbreaking LLMs through Word Based Puzzles — Mentioned as an example of jailbreaking via puzzles.

Concurring Sources

  • Turing Award — Shafi Goldwasser is a recipient, supporting her expertise.

Contribution & Novelties

The talk offers a unique cryptographic perspective on trustworthy AI, emphasizing worst-case adversarial modeling and formal proofs. It presents novel results on undetectable backdoors and defenses, and highlights the fundamental impossibility of simple filtering for alignment. The discussion of inference-time compute as a safety measure is timely.

Pour aller plus loin :

84 words

Radar Profile

The radar profile shows high scores in quality and reliability, reflecting the speaker's authority and rigorous approach. The quantity of information is moderate, as the talk is focused and not exhaustive. The technical level is high, suitable for an expert audience.

Reliability 8/10