Skip to main content

Redundancy & High Availability

Redundancy & High Availability — CM Marketplace Platform

This document describes how the CM Marketplace platform achieves redundancy, fault tolerance, and high availability across all layers of the stack. Every component — from the database to the application services to the messaging infrastructure — is designed for production resilience.


1. Architecture Overview​

The Marketplace platform runs on Google Cloud Platform (GCP) in the europe-west4 region (Netherlands). Infrastructure is fully managed as code via Terraform with modular design.

LayerTechnologyHA Mechanism
ComputeCloud Run (Knative)Auto-scaling, min/max instances
OrchestrationGKE AutopilotManaged Kubernetes, auto-repair
DatabaseCloud SQL (PostgreSQL 14)Regional HA with automatic failover
CacheMemorystore (Redis)Direct peering, VPC-integrated
MessagingPub/Sub + Cloud TasksDead-letter queues, retry with backoff
NetworkingVPC + Cloud NATStatic IP, shared NAT, access connectors
SecretsSecret ManagerPer-service isolation
MonitoringCloud Monitoring5xx alerting, email notifications

Microservices​

The platform consists of 13 independently deployable services:

ServicePurposeMin Instances
configurationCentral configuration & tenant management1
crmCRM adapter (Salesforce, Dynamics, Shopify, SAP, HubSpot)1
livechatLivechat adapter (Salesforce, Genesys, CXOne, Zendesk)2
cdpCustomer Data Platform1
cmsContent Management System1
commsCommunications / Messaging service1
product-feedProduct catalog management1
integrationsThird-party integrations1
jsltserviceJSLT transformation engine1
marketplace-frontendAngular frontend application1
marketplace-app-uiUI component library1
app-bgserviceBackground services (SSE, WebSocket workers)—
app-paymentPayment processing (PayHub adapter)—

2. Database — Regional High Availability​

The production database uses Cloud SQL PostgreSQL 14 with REGIONAL availability, meaning it maintains a synchronous standby replica in a different zone within the same region. Failover is automatic and transparent to the application.

Production Database Configuration​

# infra/environments/production/database.tf

module "database" {
source = "../../modules/database"
common_variables = local.common_variables

db_configuration = {
location_preference_zone = "europe-west4-c"
db_disk_size = 20
availability_type = "REGIONAL"
tier = "db-custom-1-3840"
database_version = "POSTGRES_14"
instance_type = "CLOUD_SQL_INSTANCE"
name = "connectcoredb"
start_time = "12:00"
region = "europe-west4"
}
}

Database Module — Backup & Recovery​

# infra/modules/database/main.tf

resource "google_sql_database_instance" "connect-core-db" {
database_version = var.db_configuration.database_version
deletion_protection = true
name = var.db_configuration.name
region = var.db_configuration.region

settings {
activation_policy = "ALWAYS"
availability_type = var.db_configuration.availability_type
deletion_protection_enabled = true
disk_autoresize = true
disk_autoresize_limit = 0
disk_type = "PD_SSD"

backup_configuration {
enabled = true
location = "eu"
point_in_time_recovery_enabled = true
start_time = var.db_configuration.start_time
transaction_log_retention_days = 7

backup_retention_settings {
retained_backups = 7
retention_unit = "COUNT"
}
}
}

lifecycle {
prevent_destroy = true
}
}

Key HA features:

  • REGIONAL availability: Synchronous standby in another zone — automatic failover with no data loss
  • Point-in-time recovery (PITR): Restore to any second within the last 7 days
  • Automated backups: Daily at 12:00 UTC, stored in the EU, 7 retained
  • SSD disks with auto-resize: No manual capacity planning needed
  • Deletion protection: Both at the Terraform and GCP API level
  • prevent_destroy lifecycle: Terraform cannot accidentally destroy the instance

Connection Pooling​

All services connect with a maximum pool size of 1024 to handle burst traffic without exhausting connections:

ConnectionStrings: "Host=...;Pooling=true;Maximum Pool Size=1024;"

3. Compute — Cloud Run Auto-Scaling​

Every service runs on Google Cloud Run with Knative autoscaling. Each service maintains a minimum of 1 instance (livechat maintains 2) so there is always a warm container ready to serve traffic. Scaling up to 10 instances happens automatically based on request concurrency.

Production Cloud Run Configuration​

# infra/environments/production/cloudrun.tf

module "cloudrun" {
source = "../../modules/cloudrun"
common_variables = local.common_variables

cloud_run_services = [
{
name = "crm"
container_concurrency = 50
timeout_seconds = 300
cpu_limit = "1"
memory_limit = "1Gi"
annotations = {
"autoscaling.knative.dev/minScale" = "1"
"autoscaling.knative.dev/maxScale" = "10"
"run.googleapis.com/cloudsql-instances" = "project:europe-west4:connectcoredb"
"run.googleapis.com/vpc-access-connector" = module.vpc_network.access_connector_ids.shared-nat
"run.googleapis.com/vpc-access-egress" = "all-traffic"
}
},
{
name = "livechat"
container_concurrency = 50
timeout_seconds = 300
cpu_limit = "1"
memory_limit = "1Gi"
annotations = {
"autoscaling.knative.dev/minScale" = "2"
"autoscaling.knative.dev/maxScale" = "10"
...
}
},
{
name = "product-feed"
container_concurrency = 50
timeout_seconds = 3600
cpu_limit = "1"
memory_limit = "2Gi"
...
},
# ... 11 services total
]
}

Cloud Run Module — Traffic & Revision Management​

# infra/modules/cloudrun/main.tf

resource "google_cloud_run_service" "cloud_run_services" {
for_each = { for service in var.cloud_run_services : service.name => service }

template {
spec {
container_concurrency = each.value.container_concurrency
timeout_seconds = each.value.timeout_seconds

containers {
image = each.value.image
resources {
limits = {
cpu = each.value.cpu_limit
memory = each.value.memory_limit
}
}
}
}
}

traffic {
percent = 100
latest_revision = true
}
}

Key HA features:

  • Minimum instances always warm: No cold-start latency for first request
  • Concurrency limit per instance: 50 requests (20 for CMS/background) — prevents overload
  • Automatic scale-out: Cloud Run adds instances when concurrency threshold is approached
  • Revision-based deploys: Previous revisions remain available for instant rollback
  • 100% traffic to latest: Zero-downtime rolling deployments

Service Resource Summary​

ServiceMemoryCPUConcurrencyTimeoutMin Scale
crm1Gi1 vCPU50300s1
livechat1Gi1 vCPU50300s2
product-feed2Gi1 vCPU503600s1
cms1Gi1 vCPU203600s1
cdp1Gi1 vCPU50300s1
configuration1Gi1 vCPU50300s1
comms1Gi1 vCPU50300s1
integrations512Mi1 vCPU50300s1
jsltservice1Gi1 vCPU50300s1
frontend1Gi1 vCPU50300s1
app-ui1Gi1 vCPU50300s1

4. GKE Autopilot — Managed Kubernetes​

For workloads that require long-running containers (background services, WebSocket connections), we use GKE Autopilot which is fully managed by Google — automatic node provisioning, repair, and security patching.

# infra/modules/gke/main.tf

resource "google_container_cluster" "primary" {
name = "${local.google_project}-gke-autopilot"
location = local.region

enable_autopilot = true
}

Key HA features:

  • Autopilot mode: Google manages node pools, scaling, security patches, and repair
  • Regional cluster: Nodes spread across multiple zones
  • Automatic pod rescheduling: If a node fails, pods move to healthy nodes

5. Caching Layer — Redis (Memorystore)​

Distributed caching is handled by Google Cloud Memorystore for Redis, connected via VPC direct peering for low-latency access from all services.

Redis Infrastructure​

# infra/modules/redis/main.tf

resource "google_redis_instance" "connectcache" {
authorized_network = var.vpc_network_id
connect_mode = "DIRECT_PEERING"
location_id = "${local.region}-${var.redis_zone}"
memory_size_gb = var.memory_size_gb
name = "connectcache"
tier = "BASIC"
read_replicas_mode = "READ_REPLICAS_DISABLED"
transit_encryption_mode = "DISABLED"

persistence_config {
persistence_mode = "DISABLED"
}

lifecycle {
prevent_destroy = true
}
}

Cache Provider — Application Layer​

// src/lib-shared/Cache/CacheProvider.cs

public sealed class CacheProvider : ICacheProvider
{
private readonly IDistributedCache _cache;
private readonly RedisConfiguration _config;

public async Task<T?> GetCacheAsync<T>(string key) where T : class
{
var searchKey = $"{_config.AppName}_{key}";
var result = await _cache.GetStringAsync(searchKey);
if (string.IsNullOrEmpty(result)) return null;
return JsonSerializer.Deserialize<T>(result);
}

public async Task SetCacheAsync<T>(string key, T value, bool isSliding = false)
{
var options = new DistributedCacheEntryOptions();
if (isSliding)
options.SetSlidingExpiration(
TimeSpan.FromMinutes(_config.SlidingExpirationInMinutes));
else
options.SetAbsoluteExpiration(
TimeSpan.FromMinutes(_config.AbsoluteExpirationInMinutes));

var serializedValue = JsonSerializer.Serialize(value);
var searchKey = $"{_config.AppName}_{key}";
await _cache.SetStringAsync(searchKey, serializedValue, options);
}

public async Task UpdateCacheAsync<T>(string key, T value, bool isSliding = false)
{
var searchKey = $"{_config.AppName}_{key}";
var cache = await _cache.GetStringAsync(searchKey);
if (string.IsNullOrWhiteSpace(cache)) return;

// Preserve existing TTL on update
var redis = await ConnectionMultiplexer.ConnectAsync(_config.Location);
var db = redis.GetDatabase();
var expiry = await db.KeyTimeToLiveAsync(searchKey);
// ... reapply same expiration ...
await _cache.SetStringAsync(searchKey, serializedValue, options);
}
}

Cache Configuration​

// src/lib-shared/Cache/RedisConfiguration.cs

public sealed class RedisConfiguration
{
[Required] public required string Location { get; init; }
[Required] public required int SlidingExpirationInMinutes { get; init; } = 15;
[Required] public required int AbsoluteExpirationInMinutes { get; init; } = 1440;
public string AppName { get; init; } = "";
}

Key features:

  • Namespace isolation: Each service prefixes keys with {AppName}_ (e.g., CRM_, LIVECHAT_)
  • Dual expiration strategies: Sliding (15 min default) for session data, Absolute (24 hours) for config data
  • TTL preservation on updates: UpdateCacheAsync reads and reapplies the existing TTL
  • Validated at startup: [Required] attributes enforce configuration correctness

6. Redis-Backed Message Queue​

For real-time workload distribution (e.g., livechat handover queuing), we use a Redis Sorted Set-based queue that provides ordered, position-tracked message delivery.

// src/lib-shared/Queue/RedisQueueService.cs

public class RedisQueueService(
IDatabase database,
IOptions<QueueConfiguration> options,
ILogger<RedisQueueService> logger) : IQueueService
{
public async Task<EnqueueResult> EnqueueAsync(
Guid tenantAdapterId, string channel, string serializedPayload)
{
var channelKey = $"{tenantAdapterId}_{channel}";
var masterKey = $"{tenantAdapterId}_{_config.KeyPrefix}";

var score = DateTimeOffset.UtcNow.ToUnixTimeMilliseconds();
await database.SortedSetAddAsync(
$"{_config.KeyPrefix}:{channelKey}", serializedPayload, score);
await database.SetAddAsync(masterKey, channelKey);

var length = await database.SortedSetLengthAsync(
$"{_config.KeyPrefix}:{channelKey}");
var rank = await database.SortedSetRankAsync(
$"{_config.KeyPrefix}:{channelKey}", serializedPayload);

return new EnqueueResult
{
Success = true,
QueuePosition = (rank ?? length - 1) + 1,
QueueLength = length
};
}

public async Task<string?> DequeueAsync(
Guid tenantAdapterId, string channelKey)
{
var entry = await database.SortedSetPopAsync(
$"{_config.KeyPrefix}:{channelKey}", Order.Ascending);
if (entry is null)
{
await database.SetRemoveAsync(
$"{tenantAdapterId}_{_config.KeyPrefix}", channelKey);
return null;
}

// Auto-cleanup: remove channel from master when empty
var remaining = await database.SortedSetLengthAsync(
$"{_config.KeyPrefix}:{channelKey}");
if (remaining == 0)
await database.SetRemoveAsync(
$"{tenantAdapterId}_{_config.KeyPrefix}", channelKey);

return entry.Value.Element.ToString();
}
}

Key features:

  • Ordered processing: ZSET with Unix timestamp scoring ensures FIFO delivery
  • Position tracking: Callers know their exact queue position
  • Auto-cleanup: Empty queues are automatically removed from the master index
  • Multi-tenant isolation: Keys are prefixed with {tenantAdapterId}

7. Event-Driven Architecture — Pub/Sub with Dead Letter Queues​

All inter-service communication uses Google Cloud Pub/Sub with attribute-based routing. Failed messages are automatically routed to a dead-letter topic after 5 delivery attempts.

Pub/Sub Configuration​

# infra/environments/production/pubsub.tf

module "pubsub" {
source = "../../modules/pubsub"

topics = [
{ name = "Events" },
{ name = "Events-error" } # Dead-letter topic
]
error_topic = "Events-error"

subscriptions = [
{
name = "CRMEvents"
filter = "attributes.integration = \"CRM\""
topic_name = "Events"
push_config = {
endpoint = "${module.cloudrun.run_urls.crm}/v1/events"
no_wrapper = { write_metadata = true }
}
dead_letter_policy = {
dead_letter_topic = "projects/.../topics/Events-error"
max_delivery_attempts = 5
}
retry_policy = {
minimum_backoff = "10s"
maximum_backoff = "600s"
}
},
{
name = "LivechatInternalEvent"
filter = "attributes.integrationName= \"Livechat\""
push_config = {
endpoint = "${module.cloudrun.run_urls.livechat}/v1/internal-event"
}
dead_letter_policy = { ... }
retry_policy = {
minimum_backoff = "10s"
maximum_backoff = "600s"
}
},
# ... 11 subscriptions total, plus dead-letter subscriber
]
}

All Production Subscriptions​

SubscriptionFilterPush Endpoint
TokenRefreshFailedType = "TokenRefresh"(pull)
CRMEventsintegration = "CRM"crm/v1/events
CRMEventsToCDPAdapterKey = "CDP"cdp/v1/events
CRMSyncEventsToCDPAdapterKey = "CDP-SYNC"cdp/v1/sync
CRMInternalEventintegrationName = "CRM"crm/v1/internal-event
CDPInternalEventintegrationName = "Ads"cdp/v1/internal-event
CDPInternalEventDataTransferintegrationName = "Data Transfer"cdp/v1/internal-event
LivechatInternalEventintegrationName = "Livechat"livechat/v1/internal-event
CommsInternalEventintegrationName = "Internal Comm"comms/v1/internal-event
ProductfeedInternalEventintegrationName = "Product Feed"product-feed/v1/internal-event
CMSInternalEventintegrationName = "CMS"cms/v1/internal-event
dead-letter(none — all from Events-error)(pull)

Key HA features:

  • Attribute-based filtering: Each subscription only receives messages it cares about
  • Dead-letter queue: Failed messages (after 5 attempts) route to Events-error for investigation
  • Exponential backoff retry: 10s minimum to 600s (10 min) maximum backoff
  • 7-day message retention: Messages survive extended service outages
  • Push delivery with OIDC auth: Secure, authenticated delivery to Cloud Run services
  • At-least-once delivery: Pub/Sub guarantees no message is lost

8. Async Task Queues — Cloud Tasks​

For work that needs controlled dispatch rates and guaranteed execution, we use Google Cloud Tasks with 4 dedicated queues:

# infra/environments/production/cloudtask.tf

module "cloudtasks" {
cloudtask_list = [
{
name = "slack_genai"
max_concurrent = 1000
max_dispatches = 500
max_attempts = 3
max_backoff = "3600s"
max_doublings = 16
min_backoff = "3s"
},
{
name = "event_queue"
max_concurrent = 1000
max_dispatches = 500
max_attempts = 3
max_backoff = "3600s"
max_doublings = 16
min_backoff = "3s"
},
{
name = "product_feed_queue"
max_concurrent = 1000
max_dispatches = 500
max_attempts = 3
...
},
{
name = "livechat_queue"
max_concurrent = 1000
max_dispatches = 500
max_attempts = 3
...
}
]
}

Cloud Tasks Module​

# infra/modules/cloudtasks/main.tf

resource "google_cloud_tasks_queue" "cloudtasks" {
for_each = { for cloudtask in var.cloudtask_list : cloudtask.name => cloudtask }

rate_limits {
max_concurrent_dispatches = each.value.max_concurrent
max_dispatches_per_second = each.value.max_dispatches
}
retry_config {
max_attempts = each.value.max_attempts
max_backoff = each.value.max_backoff
max_doublings = each.value.max_doublings
min_backoff = each.value.min_backoff
}
}

Key HA features:

  • Rate limiting: 1000 concurrent / 500 per second prevents downstream overload
  • 3 retry attempts with exponential backoff (3s to doubled up to 16 times to max 1 hour)
  • Separate queues per domain: Failures in one queue don't block another

9. Application Resilience Patterns​

Polly Retry with Exponential Backoff​

The Livechat Salesforce adapter uses Polly for retry with exponential backoff on HTTP failures:

// src/app-livechat/.../Services/AsyncRetryService.cs

public class AsyncRetryService(
IOptions<SalesforcePartnerMessagingConfiguration> configuration,
ILogger<AsyncRetryService> logger) : IAsyncRetryService
{
public AsyncRetryPolicy<Response<object>> GetRetryPolicyAsync()
{
TryParse(_pollyMaxRetries, out var maxRetries);

return Policy<Response<object>>
.Handle<HttpRequestException>(
ex => ex.StatusCode != HttpStatusCode.OK)
.WaitAndRetryAsync(
maxRetries,
attemptNumber => TimeSpan.FromSeconds(attemptNumber * 2),
(exception, sleepDuration, attemptNumber, context) =>
{
logger.LogInformation(
"Retrying in {SleepDuration}. {AttemptNumber} / {MaxRetries}",
sleepDuration, attemptNumber, maxRetries);
});
}
}

Retry schedule: Attempt 1 = 2s, Attempt 2 = 4s, Attempt 3 = 6s, etc.

SSE Connection — Infinite Retry with State Persistence​

The Salesforce Enhanced Chat background service maintains persistent SSE connections with infinite retry and state persistence on shutdown:

// src/app-bgservice/.../Services/SalesforceSseBackgroundService.cs

public class SalesforceSseBackgroundService : BackgroundService
{
private readonly ConcurrentDictionary<string, SalesforceSseConnectionState>
_activeConnections = new();

private async Task StartSseConnectionAsync(
string tenantAdapterId,
ConfigurationSettings configurationSettings,
string esDeveloperName,
CancellationToken stoppingToken)
{
// Retry forever with 5-second intervals
var retryPolicy = Policy
.Handle<Exception>()
.WaitAndRetryForeverAsync(
_ => TimeSpan.FromSeconds(5),
(exception, retryCount, timeSpan) =>
{
logger.LogInformation(
"SSE connection retry {RetryCount} for adapter {AdapterId}: {Message}",
retryCount, tenantAdapterId, exception.Message);
});

await retryPolicy.ExecuteAsync(async () =>
{
await ConnectAndListenSseAsync(
tenantAdapterId, configurationSettings,
esDeveloperName, stoppingToken);
});
}
}

Graceful Shutdown — Event State Persistence​

When the service shuts down, all active SSE connection states are persisted to cache so they can be resumed:

// src/app-bgservice/.../Services/SalesforceSseBackgroundService.cs

public override async Task StopAsync(CancellationToken cancellationToken)
{
logger.LogInformation(
"StopAsync called — persisting all active SSE connection states");
await PersistAllConnectionStatesAsync();
await base.StopAsync(cancellationToken);
logger.LogInformation("All SSE connection states persisted. Service stopped");
}

private async Task PersistAllConnectionStatesAsync()
{
var tasks = _activeConnections.Values
.Where(state => !string.IsNullOrEmpty(state.CurrentEventId))
.Select(async state =>
{
await PersistLastEventIdAsync(
state.CurrentEventId,
state.EsDeveloperName,
state.OrganizationId,
state.TenantAdapterSession);
});

await Task.WhenAll(tasks);
}

Key resilience features:

  • WaitAndRetryForeverAsync: SSE connections never give up — they reconnect indefinitely
  • Last-Event-Id tracking: On reconnection, the SSE stream resumes from the last processed event
  • Graceful shutdown: All in-flight event IDs are persisted before the container stops
  • Token refresh: Expired tokens trigger automatic re-authentication without losing the event position
  • ConcurrentDictionary: Thread-safe tracking of multiple simultaneous SSE connections

10. Networking & Security​

VPC & Cloud NAT​

All backend services route egress traffic through a shared NAT gateway with a static IP address, enabling IP-based allowlisting by external partners.

# infra/modules/vpc_network/network.tf

resource "google_compute_network" "default" {
auto_create_subnetworks = true
routing_mode = "REGIONAL"
}
# infra/modules/vpc_network/router.tf

resource "google_compute_router_nat" "shared-nat" {
name = var.vpc_network.nat_name
nat_ip_allocate_option = "MANUAL_ONLY"
nat_ips = [google_compute_address.static-marketplace.id]
min_ports_per_vm = var.vpc_network.min_ports_per_vm
tcp_established_idle_timeout_sec = 1200
tcp_time_wait_timeout_sec = 120
tcp_transitory_idle_timeout_sec = 30
}
# infra/modules/vpc_network/access_connector.tf

resource "google_vpc_access_connector" "shared-nat" {
machine_type = "f1-micro"
max_instances = 10
min_instances = 2
max_throughput = 1000
min_throughput = 200
}

Key features:

  • Static outbound IP: All services share a single public IP for partner allowlisting
  • VPC Access Connector: 2-10 instances, 200-1000 Mbps throughput, auto-scales
  • Regional routing: Traffic stays within europe-west4

Service Account Isolation (IAM)​

Each service runs under its own service account with least-privilege roles:

# infra/environments/production/main.tf

sa_roles = {
crm = ["roles/cloudsql.client", "roles/cloudtasks.enqueuer",
"roles/pubsub.publisher", "roles/redis.dbConnectionUser"]
livechat = ["roles/cloudscheduler.admin", "roles/cloudtasks.enqueuer",
"roles/redis.dbConnectionUser"]
configuration = ["roles/cloudsql.client", "roles/cloudtasks.enqueuer",
"roles/pubsub.publisher", "roles/redis.dbConnectionUser"]
product-feed = ["roles/cloudtasks.enqueuer", "roles/storage.objectUser",
"roles/redis.dbConnectionUser"]
cdp = ["roles/redis.dbConnectionUser", "roles/storage.objectUser",
"roles/cloudscheduler.admin"]
cms = ["roles/cloudscheduler.admin", "roles/cloudtasks.enqueuer",
"roles/redis.dbConnectionUser", "roles/storage.objectUser"]
comms = ["roles/cloudtasks.enqueuer", "roles/redis.dbConnectionUser"]
}

Secrets Management​

Secrets are stored in Google Secret Manager with per-service categories:

secret_list = ["common", "product_feed", "configuration",
"crm", "livechat", "cms", "comms", "integrations", "cdp"]

Secrets are injected as environment variables at Cloud Run startup — never stored in code or images.


11. Monitoring & Alerting​

5xx Error Alert​

# infra/modules/alerting/main.tf

resource "google_monitoring_alert_policy" "cloudrun_5xx_alert" {
display_name = "Cloud Run - Error Alert"
combiner = "OR"
enabled = true

conditions {
display_name = "Cloud Run service has 5xx errors"

condition_threshold {
filter = <<-EOT
resource.type = "cloud_run_revision"
AND metric.type = "run.googleapis.com/request_count"
AND metric.labels.response_code_class = "5xx"
EOT
duration = "60s"
comparison = "COMPARISON_GT"
threshold_value = 2

aggregations {
alignment_period = "60s"
per_series_aligner = "ALIGN_DELTA"
group_by_fields = [
"resource.labels.service_name",
"metric.labels.response_code"
]
}
}
}

alert_strategy {
auto_close = "1800s"
}
}

Alerting behavior:

  • Triggers when any Cloud Run service returns more than 2 5xx errors in 60 seconds
  • Groups by service name and response code for targeted triage
  • Auto-closes after 30 minutes of no further errors
  • Sends email notification to the Marketplace team

12. Container Security​

All services run as a non-root user inside Docker containers:

# src/app-crm/.../Dockerfile

FROM proget2.sharedservices.cmgroep.local/dockerfeed/trusted/dotnet/aspnet:10.0

RUN groupadd --gid 1001 appuser \
&& useradd --uid 1001 --gid 1001 -m appuser

USER appuser
WORKDIR /app
COPY ${BUILD_DIR} ./
EXPOSE 80
EXPOSE 443
ENTRYPOINT ["dotnet", "CM.Marketplace.Adapters.CRM.Web.API.dll"]

13. CI/CD — Zero-Downtime Deployments​

Deployment Pipeline​

Push to main -> Build & Test -> Docker Build -> Push to Artifact Registry
-> Auto-deploy to Acceptance -> Manual trigger -> Deploy to Production

Production Deployment Workflow​

# .github/workflows/deploy-prod.yaml

name: Deploy to Production
on:
workflow_dispatch:
repository_dispatch:
types: [artifact_committed_prod]

jobs:
deploy:
runs-on: engage-hosted
environment: Production
concurrency:
group: ${{ github.workflow }}-${{ github.ref }}
cancel-in-progress: false

steps:
- name: Checkout
uses: actions/checkout@v4

- name: Install Terraform
uses: hashicorp/setup-terraform@v3

- name: Google Auth
uses: google-github-actions/auth@v2

- name: Run Terraform init
run: terraform init

- name: Create Terraform plan
run: terraform plan -var-file="production.tfvars.json" -out=plan

- name: Apply Terraform
run: terraform apply --auto-approve "plan"

Key HA features:

  • Concurrency control: Only one deploy runs at a time (cancel-in-progress: false)
  • Environment protection: Production requires manual approval
  • Terraform plan + apply: Changes are previewed before applying
  • Cloud Run traffic management: 100% traffic shifts to the new revision only after health checks pass
  • Instant rollback: Previous Cloud Run revisions remain available

14. Rollback Procedures​

This section documents step-by-step rollback procedures for every layer of the platform. All rollback commands use gcloud CLI or GCP Console and should be executed by the Engineering Leader or an authorized team member.

14.1 Cloud Run Service Rollback​

Use this when a new deployment causes errors (5xx alerts, broken functionality). Cloud Run keeps all previous revisions available.

  1. Navigate to Revisions of specified cloud run service

    Cloud Run revisions list

  2. Click on the specified revision in which you want to route the traffic to and then click actions of that revision and then click Manage traffic.

    Cloud Run revision actions menu — Manage traffic

  3. Click on Send all traffic to one revision, and select the specified revision you want to route.

    Cloud Run Manage traffic — Send all traffic to one revision

  4. Click Save

Recovery time: Traffic shift is immediate.


14.2 Database Rollback​

Option A — Point-in-Time Recovery (PITR)​

Use this when a bad migration or data corruption has occurred. Restores the database to any second within the last 7 days.

Step 1 — Identify the target recovery time:

Determine the timestamp just before the incident (check deploy logs, migration logs, or Cloud Logging).

Step 2 — Create a clone from PITR:

# Restore to a specific point in time (creates a NEW instance)
gcloud sql instances clone connectcoredb connectcoredb-recovery \
--project=<PROJECT_ID> \
--point-in-time="2026-05-14T08:30:00.000Z"

Step 3 — Validate the recovered data:

Connect to the recovery instance and verify the data is correct before switching.

# Connect via Cloud SQL Proxy
cloud-sql-proxy <PROJECT_ID>:europe-west4:connectcoredb-recovery

Step 4 — Switch application to recovered instance:

Update the Terraform configuration or connection strings to point to the recovery instance, then redeploy.

Step 5 — Clean up:

Once validated, the old instance can be decommissioned and the recovery instance renamed if needed.

Recovery time: 10-30 minutes depending on database size.

Option B — Restore from Daily Backup​

Use this when PITR is insufficient (e.g., corrupted transaction logs).

# List available backups
gcloud sql backups list \
--instance=connectcoredb \
--project=<PROJECT_ID>

# Restore from a specific backup (creates a NEW instance)
gcloud sql instances clone connectcoredb connectcoredb-restored \
--project=<PROJECT_ID> \
--source-backup-id=<BACKUP_ID>

Recovery time: 15-45 minutes depending on backup size.

Option C — Database Migration Rollback​

Use this when a bad EF Core migration needs to be reverted. Database migrations are run via the reusable workflow gcp-postgres-migration.yml using dotnet ef database update.

# Roll back to a specific migration (run from CI/CD or locally via Cloud SQL Proxy)
dotnet ef database update <PreviousMigrationName> \
--project "<migration-project>" \
--startup-project "<startup-project>" \
--connection "<connection-string>" \
--context "<DbContext>"

Important: EF Core migrations should be written to be reversible (include a Down() method). Always verify the migration rollback in the Acceptance environment before applying to Production.

Recovery time: 1-5 minutes for schema-only migrations; longer for data migrations.


14.3 Infrastructure Rollback (Terraform)​

Use this when a Terraform apply introduces a bad infrastructure change.

Option A — Revert the Git commit and re-deploy:

# 1. Identify the last good commit
git log --oneline infra/environments/production/

# 2. Revert the bad commit
git revert <BAD_COMMIT_SHA>

# 3. Push and trigger production deploy workflow
# This runs terraform plan + apply with the reverted configuration

Option B — Re-apply from a previous Terraform state:

# 1. Check Terraform state history (if using remote state with versioning)
# 2. Re-run deploy-prod.yaml from the last known good commit
gh workflow run deploy-prod.yaml --ref <LAST_GOOD_COMMIT>

Option C — Manual Terraform rollback:

# Navigate to production environment
cd infra/environments/production

# Initialize and plan with the reverted config
terraform init
terraform plan -var-file="../../../deploy/app/config/production.tfvars.json" -out=rollback-plan

# Review the plan carefully, then apply
terraform apply rollback-plan

Recovery time: 2-10 minutes depending on the resource type being changed.


14.4 Pub/Sub Dead Letter Queue Recovery​

Use this when messages have been routed to the dead-letter queue (Events-error) and need to be reprocessed.

Step 1 — Inspect failed messages:

# Pull messages from the dead-letter subscription (without acknowledging)
gcloud pubsub subscriptions pull dead-letter \
--project=<PROJECT_ID> \
--limit=10 \
--auto-ack=false

Step 2 — Identify and fix the root cause:

Check the message attributes and payload to understand why delivery failed. Fix the underlying service issue first.

Step 3 — Republish messages to the main topic:

# Republish a message to the Events topic for reprocessing
gcloud pubsub topics publish Events \
--project=<PROJECT_ID> \
--message="<MESSAGE_DATA>" \
--attribute="integration=CRM,integrationName=CRM"

Step 4 — Acknowledge the dead-letter messages:

Once messages are reprocessed successfully, acknowledge them in the dead-letter subscription to prevent reprocessing.

Recovery time: Varies — depends on message volume and root cause fix time.


14.5 Redis Cache Rollback​

Redis cache is ephemeral — there is no rollback. If cache becomes corrupted or stale:

Step 1 — Flush specific keys:

# Connect to Redis via VPC and flush keys for a specific service
redis-cli -h <REDIS_HOST> -p 6379
> KEYS CRM_* # List keys for a service
> DEL CRM_<key> # Delete specific stale keys

Step 2 — Or flush all keys (full cache reset):

redis-cli -h <REDIS_HOST> -p 6379
> FLUSHALL

All services are designed to handle cache misses gracefully — they fall back to the database. A full flush causes temporary performance degradation (more DB queries) but no functional impact.

Recovery time: Immediate (flush). Cache repopulation happens automatically as requests flow in.


14.6 Rollback Decision Matrix​

SymptomLikely CauseRollback ProcedureRecovery Time
5xx errors after deployBad application code14.1 Cloud Run RollbackUnder 30 seconds
5xx errors + DB errors in logsBad migration14.2C Migration Rollback1-5 minutes
Data corruption discoveredBad data migration or bug14.2A PITR Recovery10-30 minutes
Infrastructure broken (networking, scaling)Bad Terraform change14.3 Terraform Rollback2-10 minutes
Messages stuck / not processingService bug or config issue14.4 DLQ RecoveryVaries
Stale data shown to usersCache corruption14.5 Redis FlushImmediate
Multiple systems affectedBad Terraform + app deploy14.1 + 14.3 combined5-15 minutes

15. Production vs Acceptance Comparison​

AspectProductionAcceptance
Database AvailabilityREGIONAL (multi-zone HA)ZONAL (single zone)
Database Tierdb-custom-1-3840 (3.8GB RAM)db-f1-micro
Database Disk20GB SSD, auto-resize10GB SSD
Cloud Run Memory1Gi-2Gi512Mi
Cloud Run Max Scale101-5
Cloud Run Min Scale1-20-1
Redis Memory1GB1GB
Pub/Sub Retention7 days7 days
Cloud Tasks Max Retry33
Monitoring AlertsEnabledEnabled

16. Summary of HA Guarantees​

Failure ScenarioMitigationRollback Procedure
Database zone failureRegional HA — automatic failover to standby in another zoneAutomatic (Cloud SQL)
Database data lossPITR + 7 daily backups in EU14.2A PITR or 14.2B Backup Restore
Bad database migrationEF Core reversible migrations14.2C Migration Rollback
Service instance crashCloud Run auto-replaces; min 1-2 instances always runningAutomatic (Cloud Run)
Traffic spikeAuto-scale up to 10 instances per serviceAutomatic (Knative)
Bad application deploymentCloud Run revision-based traffic management14.1 Cloud Run Rollback
Message delivery failurePub/Sub retries with exponential backoff (10s-600s), DLQ after 5 attempts14.4 DLQ Recovery
Task execution failureCloud Tasks retries 3x with backoff (3s-1hr)Automatic (Cloud Tasks)
SSE connection dropPolly WaitAndRetryForever with 5s intervals; Last-Event-Id resumeAutomatic (Polly)
Service shutdown during SSEGraceful shutdown persists all event IDs to Redis cacheAutomatic (BackgroundService)
External API failurePolly retry with exponential backoff (2s, 4s, 6s...)Automatic (Polly)
Network egress failureVPC Access Connector with 2-10 instances, auto-scalesAutomatic (GCP)
Bad infrastructure changeTerraform plan/apply pipeline with environment protection14.3 Terraform Rollback
5xx errors in productionAutomated alerting within 60 seconds, email to team14.1 Cloud Run Rollback
Cache corruptionEphemeral cache with graceful miss handling14.5 Redis Flush
Secret compromisePer-service secret categories in Secret Manager; env-var injectionRotate in Secret Manager, redeploy

Last updated: May 2026