How "Polite" Prompts Trick AI Into Deleting Your Files

How "Polite" Prompts Trick AI Into Deleting Your Files

🎙 Eva Benn 👥 101K 📅 January 14, 2026 ⏱ 13 min 👁 342 📄 expert opinion 🧭 2026-08-16
Available in: English (current) Français

Keywords

agentic browserprompt injectiontool misuseexcessive agencyOWASP GenAI

Summary

In this video, Eva Benn and Amanda Rousseau (Malware Unicorn) discuss a real-world attack on Perplexity’s Comet browser, where a single crafted email can cause the AI agent to delete files from Google Drive without any malware or user clicks. The attack exploits the agent’s excessive agency and its tendency to treat polite, sequential instructions as legitimate tasks. The video breaks down the attack mechanics, highlighting how the agent uses connectors to access Gmail and Google Drive, and how the email’s polite tone and ownership-shifting verbs lower resistance. Amanda explains that the attack is a form of prompt injection and tool misuse, aligning with OWASP’s Top 10 risks for agentic AI. The discussion covers defense strategies, emphasizing the need to redefine trust boundaries, restrict tool permissions at the code level, and carefully evaluate inputs and outputs. The video also touches on the future of AI security, mentioning the AIBOM project for supply chain governance. Overall, it provides a clear, accessible overview of a significant emerging threat in AI security.

169 words

Critical Evaluation

Value of the Information & Strength of the Argument

The video provides valuable insights into a real and emerging threat: agentic AI attacks. The demonstration of a zero-click attack on Google Drive is compelling and highlights the practical risks of excessive agency. The argumentation is solid, supported by expert commentary from Amanda Rousseau, who explains the attack mechanics clearly. The connection to OWASP’s Top 10 risks for agentic AI adds credibility and helps contextualize the threat. However, the video lacks deep technical details and does not show the full attack demonstration, which might leave some viewers wanting more. The discussion is well-structured, moving from attack description to defense strategies, and the emphasis on polite prompts as a vector is a novel and important point.

Scientific Rigor, Source Quality, Title Accuracy

The video references specific research from Straiker (blogs linked in description) and the OWASP GenAI Security Project, which are reputable sources. The title accurately reflects the content, focusing on the use of polite prompts to trick AI. The video does not cite any peer-reviewed papers but relies on industry research and expert opinion, which is appropriate for the topic. The presence of Amanda Rousseau, a well-known malware researcher, adds to the credibility. The video does not include any obvious misinformation, but it is important to note that the attack demonstration is not fully shown, and the video is more of a high-level overview than a detailed technical analysis. Overall, the sources are relevant and trustworthy, and the title-content alignment is good.

251 words

Title / Content Match

The title accurately reflects the content, which focuses on how polite prompts can trick AI agents into performing destructive actions.

Quality & Reliability

7/10

The video presents a real-world demonstration of an agentic AI attack, with expert commentary from Amanda Rousseau, a respected malware researcher. The claims are based on research from Straiker, and references to OWASP GenAI Security Project provide authoritative context. However, the video is largely a high-level overview without deep technical details, and the demonstration is not fully shown in the transcript.

Key Moments

Cited Sources

Concurring Sources

External References

Contribution & Novelties

The video provides a practical demonstration of a zero-click agentic AI attack, highlighting the danger of excessive agency and the effectiveness of polite prompts. It bridges the gap between theoretical OWASP risks and real-world exploitation, offering actionable insights for security teams. The discussion with Amanda Rousseau adds expert perspective on attack techniques and defense strategies.

Pour aller plus loin :

97 words

Radar Profile

The radar profile shows a balanced score across all dimensions, with slightly higher scores in information quantity and quality, and lower in technical depth. This reflects the video's strength as an accessible overview rather than a deep technical dive.

Reliability 7/10

💬 No comments were provided for analysis.