Skip to main content
gcpdown

Google Cloud reliability guides

The evergreen playbooks behind our outage analyses - how to architect so the next Service Control or GCLB event is a log line, not an incident channel.

Architecture

Multi-Region GCP Architecture Patterns

What zones and regions protect you from on Google Cloud, which products are global-by-default (and therefore fail globally), and the concrete multi-region patterns - storage, Spanner, and Cloud Run behind a global load balancer - with a cost breakdown.

10 min read

Kubernetes

GKE High Availability: The Complete Guide

What the GKE SLA actually covers (the control plane, not your workloads), regional vs zonal clusters, node-pool spread and PodDisruptionBudgets, and how to keep pods serving when the API server - or a global Google layer - is down.

13 min read