Attacking AI - Jason Haddix - NDC Security 2026

Attacking AI - Jason Haddix - NDC Security 2026

🎙 Jason Haddix 👥 227K 📅 March 26, 2026 ⏱ 55 min 👁 66K 📄 expert opinion 🧭 2026-08-13
Available in: English (current) Français

Keywords

prompt injectionAI red teamLLM assessmentRAGagentic systems

Summary

Jason Haddix, a seasoned offensive security expert, presents a comprehensive methodology for assessing AI-enabled systems, based on his experience at Arcanum. He emphasizes the shift from traditional web app testing to AI-specific challenges, highlighting the non-deterministic nature of LLMs (the ‘first try fallacy’) and the need for repeated testing. The talk outlines a seven-point methodology: identifying inputs, attacking the ecosystem, attacking the front-end model, attacking prompt engineering, attacking data, attacking the application, and pivoting to show impact. He illustrates this with real case studies, including an Amazon Rufus chatbot bypass using ASCII encoding, a healthcare system compromised via malicious file uploads leading to blind XSS and credential theft, and an automotive internal app with RAG and DevOps integration. He also discusses the importance of attacking the entire ecosystem, including agents and supporting web apps, and provides resources like the prompt injection taxonomy and CTFs like Gandalf. The talk is practical, aimed at security professionals, and stresses the need for holistic testing.

161 words

Critical Evaluation

Value of the Information & Strength of the Argument

The talk provides high value by sharing real-world case studies and a practical methodology that is often missing in academic discussions. The argumentation is solid, grounded in the speaker’s extensive experience and concrete examples. He effectively demonstrates the vulnerabilities in AI systems and the importance of holistic testing, addressing both technical and procedural aspects. The ‘first try fallacy’ concept is particularly insightful, explaining why AI testing differs from traditional security testing. The case studies are compelling and illustrate the methodology in action, making the argument persuasive.

Scientific Rigor, Source Quality, Title Accuracy

The talk demonstrates scientific rigor through its structured methodology and real-world validation. However, it lacks formal citations to academic sources, relying instead on the speaker’s expertise and anecdotal evidence. The sources cited are limited to conference links, which are not directly related to the content. The title accurately reflects the content, and the talk is well-organized. The speaker’s credibility and the practical nature of the examples enhance the overall reliability, though the lack of external references is a minor weakness.

181 words

Title / Content Match

The title accurately reflects the content, which focuses on attacking AI systems with practical methodology and case studies.

Quality & Reliability

8/10

The talk is based on extensive practical experience from real-world AI security assessments, with concrete case studies and a clear methodology. The speaker is a recognized expert in offensive security. However, claims are not backed by formal citations or peer-reviewed sources, and some details are anecdotal.

Key Moments

Cited Sources

Concurring Sources

Contribution & Novelties

The talk provides a unique, practitioner-focused methodology for AI security assessments, filling a gap between academic research and real-world testing. It introduces the ‘first try fallacy’ and emphasizes holistic testing of AI ecosystems, including agents and supporting infrastructure. The case studies offer concrete examples of attacks and their impact.

Pour aller plus loin :

82 words

Radar Profile

The radar profile shows high scores in quantity and quality of information, with moderate technical depth and high reliability. This indicates a talk that is rich in practical content and credible, but not highly technical in terms of formal theory.

Reliability 8/10