Every major Azure incident, analyzed: a timestamped timeline, the root cause in plain language, the business impact, and whether it earned SLA credits.
An allowlist configuration error on Azure Storage scale units made storage unavailable across an availability zone in Central US. Because Virtual Machines cannot run without their disks, the failure cascaded into VMs, dependent Azure services, and Microsoft 365 for many hours - landing in the same week as the unrelated CrowdStrike incident.
Layer-7 DDoS attacks - attributed by Microsoft to the actor it tracks as Storm-1359 - flooded the web front ends of the Azure Portal, Outlook on the web, and OneDrive across several days in early June 2023. Authentication and portal access failed intermittently, showing how an attack on web tiers reads as an identity outage to the business.
During a planned router addition to Microsoft’s global WAN, a command with unintended effects caused routers to forward packets incorrectly worldwide. Azure, Teams, and Outlook were disrupted globally - worst for roughly 90 minutes - with full recovery within hours after the change was rolled back.
A name-server delegation change made during a planned Azure DNS migration disrupted resolution for Microsoft services globally. Because DNS sits in front of everything, customers could not reach Storage, SQL Database, and Microsoft 365 even though those services were themselves healthy - for roughly three hours.
7 min read
Recent incident log
Smaller incidents from the live feed (last 90 days). Major events graduate into full post-mortems above.