Hacking AI Systems: How to (Still) Trick Artificial Intelligence • Katharine Jarmul • GOTO 2025

Hacking AI Systems: How to (Still) Trick Artificial Intelligence • Katharine Jarmul • GOTO 2025

🎙 Katharine Jarmul 👥 1.1M 📅 January 8, 2026 ⏱ 35 min 👁 2K 📄 expert opinion 🧭 2026-08-02
Available in: English (current) Français

Keywords

adversarial attacksmachine learningsecurityprivacyLLM

Summary

In this GOTO Copenhagen 2025 talk, Katharine Jarmul explores the field of adversarial AI/ML, focusing on how attackers can exploit weaknesses in deep learning systems. She begins by adopting the mindset of an attacker, outlining goals such as data exfiltration, service disruption, and brand damage. She then describes common AI system architectures, from large-scale distributed systems to simple API calls and local setups, highlighting shared components like models, data, and infrastructure. The core of the talk examines potential weaknesses in models, starting with training data, which often contains sensitive or inappropriate content scraped from the internet. She explains how embeddings and decision boundaries create vulnerabilities that can be exploited. The talk covers various attack vectors, including prompt injection, data poisoning, and model inversion, and emphasizes the importance of understanding coding theory and information theory to grasp these concepts. Finally, she provides a quick primer on protective measures, such as monitoring, guardrails, and robust design. The talk is practical, with references to real-world examples and resources for further learning.

168 words

Critical Evaluation

The talk provides a comprehensive overview of adversarial AI/ML, effectively bridging theoretical concepts with practical attack scenarios. Jarmul’s expertise in privacy and security is evident, and she communicates complex ideas in an accessible manner without oversimplifying. The strength of the talk lies in its structured approach: starting with attacker mindset, mapping system architectures, and then systematically identifying vulnerabilities. The use of real-world examples, such as the LAION dataset containing sensitive images, grounds the discussion in concrete issues. However, the talk lacks formal citations to specific research papers, relying instead on general knowledge and the speaker’s experience. While this is acceptable for a conference talk, it limits the ability to verify claims. The discussion of coding theory and embeddings is insightful but could be deepened for a technical audience. The protective measures section is brief, serving as a primer rather than a comprehensive guide. Overall, the talk is valuable for practitioners seeking to understand AI security threats and is well-aligned with its title. The adéquation between title and content is strong, as the talk indeed addresses how to trick AI systems and offers defensive insights.

184 words

Title / Content Match

The title accurately reflects the content, which focuses on adversarial attacks on AI systems and how to protect against them.

Quality & Reliability

8/10

The talk is delivered by a recognized expert in privacy and AI security, with practical examples and references to real-world datasets and research. The content is technically sound and well-structured, though it lacks formal citations and is based on the speaker's experience and general knowledge.

Chapters

Cited Sources

Concurring Sources

External References

Contribution & Novelties

The talk provides a practical, attacker-centric perspective on AI security, emphasizing the importance of understanding model weaknesses and the role of training data. It offers a clear framework for analyzing AI systems from a security standpoint, which is valuable for practitioners. The inclusion of real-world examples and resources for further learning enhances its utility.

Pour aller plus loin :

95 words

Radar Profile

The radar profile shows high scores in information quantity, quality, and reliability, with a slightly lower technical depth. This indicates a well-balanced talk that is informative and credible, though it may not delve into the most advanced technical details.

Reliability 8/10