Skip to main content
gcpdown

Google Cloud reliability guides

The evergreen playbooks behind our outage analyses - how to architect so the next Service Control or GCLB event is a log line, not an incident channel.

A network patch panel with colored cables
Architecture

GCP Load Balancing and Multi-Region Failover

How the global HTTP(S) load balancer routes traffic to the nearest healthy region, how to configure health checks and failover, and its global-dependency risk.

9 min read

A network of light over the Earth at night
Architecture

Multi-Region GCP Architecture Patterns

What zones and regions protect you from on Google Cloud, which products are global-by-default (and therefore fail globally), and the concrete multi-region patterns - storage, Spanner, and Cloud Run behind a global load balancer - with a cost breakdown.

10 min read

Shipping containers at a port
Kubernetes

GKE High Availability: The Complete Guide

What the GKE SLA actually covers (the control plane, not your workloads), regional vs zonal clusters, node-pool spread and PodDisruptionBudgets, and how to keep pods serving when the API server - or a global Google layer - is down.

13 min read