Geographic Redundancy
Geographic Redundancy
This page describes how the CM Marketplace platform achieves geographic redundancy within its deployment region. For detailed HA configurations (database failover, retry policies, Polly patterns, graceful shutdown), see Redundancy & High Availability.
1. Deployment Region
All production infrastructure runs in Google Cloud Platform europe-west4 (Eemshaven, Netherlands). This region consists of 3 availability zones: europe-west4-a, europe-west4-b, and europe-west4-c.
2. Zone-Level Redundancy
Database — Multi-Zone (Regional HA)
The production Cloud SQL instance uses availability_type = "REGIONAL", which means:
- A primary instance runs in europe-west4-c
- A synchronous standby replica runs in a different zone within europe-west4
- Failover is automatic and transparent — applications reconnect without configuration changes
- Data is synchronously replicated — zero data loss on failover
# infra/environments/production/database.tf
db_configuration = {
availability_type = "REGIONAL" # Multi-zone HA
location_preference_zone = "europe-west4-c"
region = "europe-west4"
}
Recovery capabilities:
- Automatic failover: ~60 seconds, handled by Cloud SQL
- Point-in-time recovery: Restore to any second within last 7 days
- Daily backups: Stored in the
eumulti-region location (covers all EU regions)
backup_configuration {
enabled = true
location = "eu" # Multi-region backup storage
point_in_time_recovery_enabled = true
transaction_log_retention_days = 7
backup_retention_settings {
retained_backups = 7
}
}
Cloud Run — Multi-Zone by Default
Cloud Run in a given region automatically distributes instances across all available zones:
- When
minScale = 1, the instance can be in any zone - When
minScale = 2(livechat), instances are spread across zones - On zone failure, Cloud Run automatically places instances in surviving zones
- No configuration needed — this is built into Cloud Run's scheduler
GKE Autopilot — Regional Cluster
resource "google_container_cluster" "primary" {
location = local.region # Regional = multi-zone
enable_autopilot = true
}
A regional GKE Autopilot cluster runs the control plane and nodes across all 3 zones. If one zone fails, pods are automatically rescheduled to nodes in surviving zones.
Redis — Single Zone
# infra/modules/redis/main.tf
tier = "BASIC"
location_id = "${local.region}-c" # europe-west4-c only
Redis currently runs in a single zone (BASIC tier). This is acceptable because:
- Cache is ephemeral — all data can be regenerated from the database
- Cache misses result in a database read, not an error
- Application code handles
nullcache responses gracefully
Upgrade path for zone redundancy: Switch to STANDARD tier with read_replicas_mode = "READ_REPLICAS_ENABLED" for cross-zone replication.
3. Network Redundancy
VPC Access Connector
# infra/modules/vpc_network/access_connector.tf
min_instances = 2 # At least 2 connector instances always running
max_instances = 10
With a minimum of 2 instances, the connector survives single-instance failures.
Cloud NAT
# infra/modules/vpc_network/router.tf
resource "google_compute_router_nat" "shared-nat" {
nat_ip_allocate_option = "MANUAL_ONLY"
nat_ips = [google_compute_address.static-marketplace.id]
}
Cloud NAT is a managed, distributed service — it does not have a single point of failure. The static IP is regional and survives zone failures.
4. Messaging Redundancy
Pub/Sub — Global Service
Google Cloud Pub/Sub is a globally distributed service:
- Messages are replicated across multiple zones within the region
- Pub/Sub has a 99.95% SLA
- Message retention is 7 days — survives extended zone outages
- Dead-letter topic (
Events-error) provides a safety net for failed messages
Cloud Tasks — Regional Service
Cloud Tasks stores tasks durably within the region across multiple zones. Tasks survive zone failures and are dispatched from surviving infrastructure.
5. Backup Strategy — Cross-Region
| Data | Backup Location | Retention | Recovery |
|---|---|---|---|
| Cloud SQL | eu (multi-region) | 7 daily backups | PITR to any second in 7 days |
| Pub/Sub messages | Regional (auto) | 7 days | Automatic re-delivery |
| Cloud Tasks | Regional (auto) | Until execution | Automatic retry |
| Redis cache | None (ephemeral) | — | Regenerated from DB |
| Secrets | Regional | Versioned | Automatic |
| Container images | Artifact Registry (eu) | All versions | Redeploy any version |
Database backups are stored in the eu multi-region location, meaning they survive even a complete region failure and can be used to restore in another European region.
6. Summary — What Survives What
| Failure | Impact | Recovery |
|---|---|---|
| Single zone failure | Zero downtime — all components have zone redundancy (except Redis, which degrades gracefully to DB reads) | Automatic |
| Region failure | Full outage — current architecture is single-region | Manual: restore DB from EU backups in new region, redeploy from Artifact Registry |
| Pub/Sub zone failure | No impact — Pub/Sub replicates across zones | Automatic |
| NAT instance failure | No impact — Cloud NAT is distributed | Automatic |
| Single Cloud Run instance failure | No impact — minScale ensures other instances serve traffic | Automatic (sub-second) |
For multi-region expansion strategy, see Multi-Region Readiness. For detailed HA configurations, see Redundancy & High Availability.
Last updated: May 2026