SIEM vs. Data Lake: Why We Ditched Traditional Logging?

SIEM vs. Data Lake: Why We Ditched Traditional Logging?

🎙 Cloud Security Podcast 👥 39K 📅 December 2, 2025 ⏱ 46 min 👁 16K 📄 expert opinion 🧭 2026-08-16
Available in: English (current) Français

Keywords

SIEMdata lakeS3AthenaOCSF

Summary

In this episode of the Cloud Security Podcast, host Ashish Rajan interviews Cliff Crosland, CEO and co-founder of Scanner.dev, about the shift from traditional SIEMs to data lakes for security log management. Crosland shares his experience at a previous startup where escalating log volumes made their SIEM license cost more than the entire engineering budget, prompting them to redirect logs to S3 and query with Amazon Athena. However, this approach proved impractical: Athena queries were slow and expensive for large datasets, and the data lake became a ‘black hole’ of unsearchable data. He discusses the engineering lift required to build and maintain a data lake, including schema normalization and the challenges of custom log sources. The conversation covers the evolution of logging from SQL-based systems to full-text search, the role of Amazon Security Lake and OCSF, and the potential of AI agents to automate detection engineering and schema management. Crosland emphasizes that while data lakes offer cost-effective, scalable storage, they require significant engineering effort and are not yet as user-friendly as traditional SIEMs. He predicts that future tools will embrace ‘messy’ logs and make data lakes more accessible, potentially replacing SIEMs for many organizations.

194 words

Critical Evaluation

Value of the Information & Strength of the Argument

The episode provides valuable practical insights into the economic and technical drivers for moving from SIEMs to data lakes. Crosland’s firsthand account of the cost pressures and the limitations of using Athena for security investigations is compelling and grounded in real experience. He articulates the trade-offs between cost, visibility, and usability, and highlights the often-overlooked challenge of schema normalization for custom logs. The argumentation is coherent and persuasive, though it is largely anecdotal and lacks quantitative data or comparative analysis. The discussion is balanced, acknowledging both the benefits and the significant engineering effort required, which adds credibility. However, the guest’s role as CEO of a data lake startup introduces a potential bias toward promoting data lake solutions, though he does candidly discuss the difficulties.

Scientific Rigor, Source Quality, Title Accuracy

The episode is an expert opinion piece rather than a formal scientific presentation. The guest cites specific tools and technologies (e.g., Splunk, Amazon Athena, Amazon Security Lake, OCSF) and mentions an example from Apple’s security team, but no formal sources or studies are referenced. The description provides links to the podcast’s website, bootcamp, newsletter, and LinkedIn, but these are not sources for the technical claims. The title accurately reflects the content, which focuses on the rationale for moving away from traditional SIEMs. The discussion is technically accurate in its descriptions of data lake architectures and query performance, but it would benefit from citations to official documentation or industry reports to enhance rigor.

251 words

Title / Content Match

The title accurately reflects the core discussion: the reasons for moving from traditional SIEMs to data lakes, based on the guest's experience.

Quality & Reliability

7/10

The episode features a practitioner with direct experience building and operating a data lake for security logging, providing practical insights and candid lessons. However, claims are largely anecdotal and not backed by formal studies or citations, and the discussion is oriented toward promoting the guest's startup, which may introduce bias.

Chapters

Cited Sources

Concurring Sources

  • Amazon Athena documentation — Confirms Athena's SQL-based querying and performance characteristics for large datasets.
  • OCSF official site — Describes the Open Cybersecurity Schema Framework, which is central to the discussion of log normalization.

Contribution & Novelties

The episode offers a candid, practitioner-driven perspective on the challenges and benefits of building a security data lake, particularly the economic breaking point where SIEM costs become prohibitive and the practical limitations of SQL-based query engines like Athena for security investigations. It highlights the often-underestimated engineering effort required for schema normalization and the potential of AI agents to automate detection engineering. The discussion of Amazon Security Lake and OCSF provides a current overview of managed solutions.

Pour aller plus loin :

144 words

Radar Profile

The radar profile shows high scores in information quantity and quality, reflecting the depth of practical insights, but lower scores in technical depth and reliability, indicating that the content is more anecdotal than rigorously technical. The overall balance suggests a valuable practitioner perspective with room for more formal evidence.

Reliability 6/10

💬 No comments were provided for analysis.