Microsoft Blames Automated Maintenance Bug for Five-Hour Azure and Microsoft 365 Outage in West US
A maintenance-system bug stripped IP routes from Microsoft's West US datacenter on July 23, knocking out Teams, SharePoint, and Azure services for about five hours.
Editor's Note ·
- Clarification:
- The article presents "IP routes were removed from more devices than intended between Microsoft's West US datacenter and its wide-area network" as a direct quote from Microsoft, citing BleepingComputer. BleepingComputer's report does not put this sentence in quotation marks; it is the reporter's own paraphrase of Microsoft's post-incident review, not a verbatim company statement. The underlying fact is accurate and is genuinely a verbatim Microsoft quote as reported by The Register ("a bug in the request conversion system incorrectly marked additional devices as a part of the maintenance event and caused a set of IP routes to be removed from more devices than intended"), which the article correctly quotes with proper attribution in the same paragraph.
Overview
A bug in Microsoft’s automated network maintenance system triggered a roughly five-hour outage across Microsoft 365 and Azure services on July 23, after the system removed IP routes from more devices than intended between Microsoft’s West US datacenter and its wide-area network, according to BleepingComputer and The Register.
What We Know
The disruption began at 10:44 AM ET on July 23, when what The Register described as routine device maintenance in the West US datacenter went wrong, according to The Register. Within a minute, multiple Azure services began showing degradation, and by 11:11 AM ET, Downdetector had logged 2,403 outage reports compared with a normal baseline of 29, as reported by BleepingComputer. SharePoint accounted for 78 percent of those complaints, Excel for 11 percent, and the Microsoft 365 Admin Center for 6 percent, per the same report.
Microsoft attributed the failure to a bug in its maintenance-request system: “IP routes were removed from more devices than intended between Microsoft’s West US datacenter and its wide-area network,” the company said, according to BleepingComputer. The Register’s independent account of Microsoft’s explanation matches almost word for word: “A bug in the request conversion system incorrectly marked additional devices as a part of the maintenance event and caused a set of IP routes to be removed from more devices than intended,” with the routes removed “between our datacenter and wide-area network, impacting traffic entering or exiting the region,” The Register reported. The Register also linked the trigger to recent fiber maintenance activity near the affected devices.
The outage’s reach extended well beyond a single product. BleepingComputer listed OneDrive, SharePoint Online, Microsoft Teams, the Microsoft 365 Admin Center, Power Automate, Copilot Chat, Microsoft Loop, Fabric, Power BI, Power Apps, Copilot Studio, Windows 365, and Microsoft Defender among the affected services, with users reporting intermittent access failures, error messages, and degraded chat functionality with images failing to load in Teams. On the Azure side, App Service, Application Gateway, and Azure AD B2C were among the services affected, per the same report. The Register put the total at 27 affected Azure services in the West US region alone.
Microsoft began rolling back the change at 1:45 PM ET (17:45 UTC), completed the rollback and restored the wide-area network connection at 2:26 PM ET (18:26 UTC), and confirmed all services were fully recovered by 3:41 PM ET (19:41 UTC) — a timeline reported consistently by both BleepingComputer and The Register, putting the total disruption at approximately five hours.
What We Don’t Know
- Neither report specifies how many customers or organizations outside Microsoft’s own telemetry and Downdetector sample were affected, or the total financial impact of the outage.
- Neither source names the Microsoft engineering team or individual spokesperson behind the public statement; it is attributed to the company generally.
- It is not yet clear from either report what specific changes, if any, Microsoft has made to its maintenance-request conversion system to prevent a recurrence.
Analysis
The incident is notable less for its cause — a misconfigured maintenance job is a familiar failure mode in large cloud networks — than for its blast radius. Because the affected IP routes sat between an entire regional datacenter and Microsoft’s wide-area network, the failure cascaded across more than a dozen customer-facing Microsoft 365 apps and dozens of underlying Azure services simultaneously, rather than degrading a single product in isolation. That a single automated maintenance job could ripple across Teams, SharePoint, Power Platform, and core Azure infrastructure at once underscores how tightly coupled Microsoft’s productivity suite has become with its cloud backbone.