A routine maintenance task at Google Cloud’s us-central1-b region escalated into a multi-hour outage after an engineer inadvertently disconnected every fiber-optic cable in the affected zone. The incident, which lasted from 07:41 to 11:52 PT on September 1, left virtual machines unreachable and caused complete traffic flow disruption for resources in the impacted area.
What happened
During what was intended to be a standard hardware maintenance procedure, an engineer sequentially unplugged 100% of the fiber paths connecting network devices within 13 minutes. Google’s service health report noted that the disconnection bypassed built-in redundancy safeguards, which are designed to withstand single or even multiple device failures without affecting customer traffic. The company’s infrastructure relies on physical separation of devices and fiber paths, along with diverse power sources, to prevent such disruptions. However, the rapid and sequential disconnection of all cables prevented automated warnings from reaching the engineer before the damage was done.
Once the cables were unplugged, virtual machines in one zone of us-central1-b lost external connectivity entirely, though they retained internal communication. Google’s response involved rerouting traffic to healthy capacity elsewhere in the region, followed by physically reseating the disconnected fibers. Normal traffic flow resumed once the links were restored, but the outage lasted over four hours, with peak disruption reaching 100% packet loss for affected resources.
Why redundancy failed
Google’s cloud infrastructure is engineered to handle multiple simultaneous failures without customer impact. The system’s resilience is based on physical separation of routing devices and fiber paths, ensuring that no single point of failure can take down an entire zone. However, the incident revealed a critical vulnerability: human error during maintenance can override these safeguards if procedures are not followed precisely. In this case, the engineer’s actions sequentially disabled all redundant paths, effectively isolating the zone from the broader network.
The company’s preliminary root cause analysis emphasized that the disconnection was not a technical failure but a procedural one. Google’s maintenance protocols are designed to prevent such outcomes, but the engineer’s deviation from these protocols—whether due to oversight or miscommunication—resulted in the outage. The speed of the disconnection, completed within 13 minutes, further limited the system’s ability to trigger corrective alerts in time.
What Google is doing
Google has not disclosed specific corrective actions but indicated that the incident would prompt a review of maintenance procedures and training. The company’s service health report acknowledged that while the infrastructure itself performed as designed once the issue was identified, the procedural failure highlighted a need for additional safeguards. Technicians were able to restore service by physically reseating the fibers and redirecting traffic, but the incident underscored the risks of human error in highly redundant systems.
For professionals: Operators should review maintenance protocols to ensure redundant paths cannot be disabled sequentially. Automated alerts for rapid disconnections may help prevent similar incidents, particularly in environments where physical separation is relied upon for resilience.
Broader implications
The outage serves as a reminder that even the most robust cloud infrastructure can be compromised by simple human mistakes. While Google’s redundancy measures are effective against hardware failures, they are not foolproof against procedural errors. The incident may prompt other hyperscale providers to reassess their own maintenance workflows, particularly in regions where physical redundancy is a cornerstone of reliability.
For customers, the event reinforces the importance of multi-region deployments for critical workloads. Even a single zone outage can have significant operational impacts, and relying on a single region—no matter how resilient—introduces avoidable risk.
Companies mentioned
Automated pipeline · Security
Synthesized from 1 industry feed on 4 Sep 2026. Passed independent editor verification (score 95/100) before publication. Style guide v1.4.
Sources
Decision trail
- Checking for duplicates — New story No recent or in-pipeline article covers this specific Google Cloud outage incident.
- Checking for duplicates — New story pre_write:; No recent or in-pipeline article covers this specific G-Cloud outage incident.
- Writing the article — Draft created article_id=508 slug=google-cloud-outage-triggered-by-engineer-unplugging-fibers
-
Editor review — Approved
- Score: 95/100
- Factual grounding: The draft states the outage lasted 'over four hours' while the source specifies the duration as 07:41 to 11:52 PT (4 hours and 11 minutes). The phrasing 'over four hours' is technically correct but slightly imprecise. This is minor as the exact duration is provided elsewhere in the text.
- Style compliance: The standfirst uses the phrase 'cut off a US region' while the source specifies 'a portion of the G-Cloud region' or 'one zone of us-central1-b'. The standfirst overstates the scope of the outage. This is minor as the body clarifies the affected area accurately.
- No copied phrasing: The phrase 'redundancy across multiple routing devices' and 'physically separated in each datacenter, with diverse power sources' closely mirrors the source wording. While the ideas are paraphrased, the phrasing is too similar to the source. This is minor as the facts are correctly attributed.
- Generating reader Q&A — Generated 4 items
- Assigning hero image — Reused library image reused image #121
- Linking related stories — Linked 2 relations from 441 candidates
- Publishing — Published google-cloud-outage-triggered-by-engineer-unplugging-fibers
- Mastodon — Posted https://mstdn.social/@hostingpaper/117211213417017833



Discussion · coming soon
Be the first to join the thread when community discussion launches.