Skip to main content
gcpdown

Is Google Cloud down right now?

Know before it hits your SLAs.

All Services Operational

Data from the official Google Cloud Service Health · independent, not affiliated with Google.

Post-mortems

The latest outage analyses

What actually broke, why, and what it means for your platform - every major Google Cloud incident, explained clearly.

Cloud IAMcritical

The June 2025 Google Cloud Global Outage: How Service Control Crash-Looped

A policy update containing unintended blank fields reached a code path in Service Control - the global layer that checks quota and policy on nearly every Google Cloud API call - that had a null-pointer defect and no feature-flag protection. Service Control crash-looped globally, returning 503s across dozens of products for roughly three hours; us-central1 took longest to recover due to a retry herd.

9 min read

Cloud Load Balancingcritical

The November 2021 Google Cloud Load Balancer Outage: A Config Race Condition

A race condition in the pipeline that propagates Google Cloud Load Balancer configuration caused a corrupt configuration to be pushed globally. Requests to sites fronted by GCLB began returning 404 errors even though the backends were healthy - taking down high-profile sites for roughly two hours until the configuration was rolled back.

7 min read

Virtual Private Cloudcritical

The June 2019 Google Cloud Network Congestion Event

A routine maintenance configuration change was applied to a far wider scope than intended, descheduling the network control-plane jobs in several US regions. The network fell back to preserving only high-priority traffic (like the control plane’s own signaling) while starving everything else, so Compute Engine, YouTube, Gmail, and Snapchat degraded severely for about four hours - worst in the US East.

8 min read

Cloud Storagemajor

The March 2019 Google Cloud Storage Outage: An Internal Blob Failure

Degradation in an internal blob-storage service that underpins Google Cloud Storage caused elevated error rates and latency for GCS operations, and for Google’s own products that store objects on the same layer - Gmail and Photos attachments among them - for roughly four hours.

6 min read

Browse the full post-mortem archive →

Guides

Build for the next outage

Evergreen platform-engineering playbooks written from real incident patterns.

Architecture

Multi-Region GCP Architecture Patterns

What zones and regions protect you from on Google Cloud, which products are global-by-default (and therefore fail globally), and the concrete multi-region patterns - storage, Spanner, and Cloud Run behind a global load balancer - with a cost breakdown.

10 min read

Kubernetes

GKE High Availability: The Complete Guide

What the GKE SLA actually covers (the control plane, not your workloads), regional vs zonal clusters, node-pool spread and PodDisruptionBudgets, and how to keep pods serving when the API server - or a global Google layer - is down.

13 min read

Know before your customers do

Instant email alerts the moment a Google Cloud incident is detected - with the affected services and regions, not a vague status tweet.