Google Webinar | Risk Classification with LLMs

Google Webinar | Risk Classification with LLMs

🎙 Sharon Crew 👥 3K 📅 December 15, 2025 ⏱ 41 min 👁 46 📄 expert opinion 🧭 2026-08-16
Available in: English (current) Français

Keywords

LLMrisk classificationsecurityground truthprompt engineering

Summary

This webinar, presented by Sharon Crew from Google Cloud CISO security engineering, addresses the challenge of classifying security risks at scale using large language models (LLMs). The speaker begins by defining ‘risk insights’ as the transformation of raw vulnerability data into actionable themes and patterns. She explains why LLMs are chosen over manual classification: to handle the volume, inconsistency, and lack of normalization in security data from various sources. The core of the presentation is a step-by-step process for building an LLM-based classification system: defining the use case, developing an initial prompt, collecting ground truth data with expert consensus, measuring accuracy, iterating on prompts, and monitoring for drift. Key lessons learned include the importance of a clear taxonomy, using temperature zero for determinism, employing thinking models for better reasoning, and using multilabel classification. The speaker emphasizes that ground truth collection is the most challenging part and that LLMs are not magic; they reflect the quality of the taxonomy and data. The webinar concludes with a Q&A session addressing questions about determinism, multimodal models, model selection, data poisoning, and open-source platforms.

180 words

Critical Evaluation

Value of the Information & Strength of the Argument

The value of the information is high for practitioners seeking to implement LLM-based classification in security. The speaker provides a concrete, actionable framework and shares real-world lessons learned, such as the need for a clear taxonomy, the importance of ground truth, and the benefits of temperature zero and thinking models. The argumentation is solid, based on practical experience rather than theoretical claims. The speaker acknowledges limitations and emphasizes that LLMs are not magic, which adds credibility. However, the presentation lacks quantitative evidence (e.g., specific accuracy numbers) and does not compare with alternative methods in depth.

Scientific Rigor, Source Quality, Title Accuracy

The scientific rigor is moderate. The speaker references the MITRE CWE-25 taxonomy and mentions internal Google practices but provides no formal citations or links to external sources. The title accurately reflects the content. The webinar is an expert opinion based on practical experience, not a peer-reviewed study. The lack of detailed data and external references reduces the overall scientific rigor, but the practical insights are valuable.

176 words

Title / Content Match

The title accurately reflects the content: a webinar on risk classification using LLMs.

Quality & Reliability

7/10

The webinar provides a practical, experience-based overview of using LLMs for security risk classification. It includes concrete methodology, lessons learned, and acknowledges limitations. However, it lacks detailed quantitative results, formal citations, and is based on a single practitioner's perspective.

Key Moments

Cited Sources

Concurring Sources

Contribution & Novelties

The webinar provides a practical, experience-based methodology for using LLMs in security risk classification, emphasizing the importance of ground truth and prompt iteration. It offers specific lessons learned, such as using temperature zero, thinking models, and multilabel classification. The speaker shares a real-world use case from Google Cloud, which adds credibility.

Pour aller plus loin :

103 words

Radar Profile

The radar profile shows a balanced performance across all dimensions, with slightly higher scores in information quantity and quality, reflecting the practical and detailed nature of the webinar. The technical level is moderate, suitable for a broad security audience, and the reliability is good given the expert source.

Reliability 7/10