Every Major AWS Outage (And Why They Keep Happening)

Every Major AWS Outage (And Why They Keep Happening)

🎙 freeCodeCamp.org 👥 11.8M 📅 July 3, 2026 ⏱ 16 min 👁 34K 📄 documentary 🧭 2026-08-03
Available in: English (current) Français

Keywords

AWSoutageUS-East-1cloud reliabilitycascading failure

Summary

The video chronicles six major AWS outages that occurred between 2011 and 2025, all centered on the US-East-1 region in Northern Virginia. It begins with the 2011 network upgrade mistake that triggered a feedback loop, followed by the 2012 storm that caused power failures, the 2017 S3 typo that took down the index subsystem, the 2020 Kinesis dependency cascade, the 2021 outage that affected Amazon’s own services, and the 2025 DynamoDB DNS automation failure. Each incident is explained with technical detail, highlighting how human error, hidden dependencies, and complexity lead to cascading failures. The video emphasizes that US-East-1 has become a critical chokepoint for the internet, and despite repeated postmortems and safeguards, outages continue to occur due to unforeseen failure modes. It concludes with the uncomfortable truth that complex systems fail in unpredictable ways, and the concentration of traffic in one region amplifies the impact.

145 words

Critical Evaluation

The video offers a compelling and well-researched overview of major AWS outages, effectively illustrating the fragility of centralized cloud infrastructure. It excels in storytelling, making technical concepts accessible without oversimplifying the underlying causes. The narrative is supported by specific dates, technical terms, and references to public postmortems, which lends credibility. However, the video lacks direct citations to primary sources, relying instead on general knowledge and possibly secondary accounts. While the technical explanations are generally accurate, some details may be condensed for narrative effect, such as the exact nature of the 2025 DNS automation failure. The video also implicitly critiques AWS’s reliability, but it does not delve into broader discussions about cloud monopolies or regulatory implications. The inclusion of Netflix’s Chaos Monkey as a resilience strategy is a highlight, demonstrating proactive failure engineering. The adéquation between title and content is strong, as the video systematically covers each major outage. Overall, the video is a valuable educational resource for understanding cloud reliability challenges, though it could benefit from more explicit sourcing and a deeper exploration of systemic solutions.

176 words

Title / Content Match

The title accurately reflects the content, which systematically covers major AWS outages and their root causes.

Quality & Reliability

8/10

The video provides a well-structured historical account of major AWS outages, referencing specific incidents and technical details. It includes expert commentary and references to public postmortems, though it lacks direct citations to primary sources. The narrative is engaging and generally accurate, but some details may be simplified for a general audience.

Chapters

Cited Sources

  • freeCodeCamp News — General resource for programming articles, potentially including AWS-related content.
  • Scrimba — Sponsor link, not directly related to the content.
  • freeCodeCamp — Main website of the channel, offering free coding education.

Concurring Sources

Dissenting Sources

Contribution & Novelties

The video synthesizes publicly known information about AWS outages into a coherent narrative, highlighting the systemic issue of concentration in US-East-1. It provides a historical perspective that is often scattered across individual postmortems. The inclusion of Netflix’s Chaos Monkey as a resilience example offers a practical lesson. The video’s main contribution is raising awareness about the fragility of centralized cloud infrastructure and the importance of designing for failure.

Pour aller plus loin :

116 words

Radar Profile

The radar profile shows high scores in quantity and quality of information, with a moderate technical level, indicating a well-balanced educational video. The reliability score is also high, reflecting the use of known incidents and public postmortems.

Reliability 8/10