Skip to main content

Geographic Redundancy

Geographic Redundancy

This page describes how the CM Marketplace platform achieves geographic redundancy within its deployment region. For detailed HA configurations (database failover, retry policies, Polly patterns, graceful shutdown), see Redundancy & High Availability.


1. Deployment Region​

All production infrastructure runs in Google Cloud Platform europe-west4 (Eemshaven, Netherlands). This region consists of 3 availability zones: europe-west4-a, europe-west4-b, and europe-west4-c.


2. Zone-Level Redundancy​

Database — Multi-Zone (Regional HA)​

The production Cloud SQL instance uses availability_type = "REGIONAL", which means:

  • A primary instance runs in europe-west4-c
  • A synchronous standby replica runs in a different zone within europe-west4
  • Failover is automatic and transparent — applications reconnect without configuration changes
  • Data is synchronously replicated — zero data loss on failover
# infra/environments/production/database.tf

db_configuration = {
availability_type = "REGIONAL" # Multi-zone HA
location_preference_zone = "europe-west4-c"
region = "europe-west4"
}

Recovery capabilities:

  • Automatic failover: ~60 seconds, handled by Cloud SQL
  • Point-in-time recovery: Restore to any second within last 7 days
  • Daily backups: Stored in the eu multi-region location (covers all EU regions)
backup_configuration {
enabled = true
location = "eu" # Multi-region backup storage
point_in_time_recovery_enabled = true
transaction_log_retention_days = 7
backup_retention_settings {
retained_backups = 7
}
}

Cloud Run — Multi-Zone by Default​

Cloud Run in a given region automatically distributes instances across all available zones:

  • When minScale = 1, the instance can be in any zone
  • When minScale = 2 (livechat), instances are spread across zones
  • On zone failure, Cloud Run automatically places instances in surviving zones
  • No configuration needed — this is built into Cloud Run's scheduler

GKE Autopilot — Regional Cluster​

resource "google_container_cluster" "primary" {
location = local.region # Regional = multi-zone
enable_autopilot = true
}

A regional GKE Autopilot cluster runs the control plane and nodes across all 3 zones. If one zone fails, pods are automatically rescheduled to nodes in surviving zones.

Redis — Single Zone​

# infra/modules/redis/main.tf

tier = "BASIC"
location_id = "${local.region}-c" # europe-west4-c only

Redis currently runs in a single zone (BASIC tier). This is acceptable because:

  • Cache is ephemeral — all data can be regenerated from the database
  • Cache misses result in a database read, not an error
  • Application code handles null cache responses gracefully

Upgrade path for zone redundancy: Switch to STANDARD tier with read_replicas_mode = "READ_REPLICAS_ENABLED" for cross-zone replication.


3. Network Redundancy​

VPC Access Connector​

# infra/modules/vpc_network/access_connector.tf

min_instances = 2 # At least 2 connector instances always running
max_instances = 10

With a minimum of 2 instances, the connector survives single-instance failures.

Cloud NAT​

# infra/modules/vpc_network/router.tf

resource "google_compute_router_nat" "shared-nat" {
nat_ip_allocate_option = "MANUAL_ONLY"
nat_ips = [google_compute_address.static-marketplace.id]
}

Cloud NAT is a managed, distributed service — it does not have a single point of failure. The static IP is regional and survives zone failures.


4. Messaging Redundancy​

Pub/Sub — Global Service​

Google Cloud Pub/Sub is a globally distributed service:

  • Messages are replicated across multiple zones within the region
  • Pub/Sub has a 99.95% SLA
  • Message retention is 7 days — survives extended zone outages
  • Dead-letter topic (Events-error) provides a safety net for failed messages

Cloud Tasks — Regional Service​

Cloud Tasks stores tasks durably within the region across multiple zones. Tasks survive zone failures and are dispatched from surviving infrastructure.


5. Backup Strategy — Cross-Region​

DataBackup LocationRetentionRecovery
Cloud SQLeu (multi-region)7 daily backupsPITR to any second in 7 days
Pub/Sub messagesRegional (auto)7 daysAutomatic re-delivery
Cloud TasksRegional (auto)Until executionAutomatic retry
Redis cacheNone (ephemeral)—Regenerated from DB
SecretsRegionalVersionedAutomatic
Container imagesArtifact Registry (eu)All versionsRedeploy any version

Database backups are stored in the eu multi-region location, meaning they survive even a complete region failure and can be used to restore in another European region.


6. Summary — What Survives What​

FailureImpactRecovery
Single zone failureZero downtime — all components have zone redundancy (except Redis, which degrades gracefully to DB reads)Automatic
Region failureFull outage — current architecture is single-regionManual: restore DB from EU backups in new region, redeploy from Artifact Registry
Pub/Sub zone failureNo impact — Pub/Sub replicates across zonesAutomatic
NAT instance failureNo impact — Cloud NAT is distributedAutomatic
Single Cloud Run instance failureNo impact — minScale ensures other instances serve trafficAutomatic (sub-second)

For multi-region expansion strategy, see Multi-Region Readiness. For detailed HA configurations, see Redundancy & High Availability.

Last updated: May 2026