The evergreen playbooks behind our outage analyses - how to architect so the next Service Control or GCLB event is a log line, not an incident channel.
What zones and regions protect you from on Google Cloud, which products are global-by-default (and therefore fail globally), and the concrete multi-region patterns - storage, Spanner, and Cloud Run behind a global load balancer - with a cost breakdown.
What the GKE SLA actually covers (the control plane, not your workloads), regional vs zonal clusters, node-pool spread and PodDisruptionBudgets, and how to keep pods serving when the API server - or a global Google layer - is down.