Skip to main content

Scaling Strategy

Scaling Strategy

This page describes how each component of the CM Marketplace platform scales to handle varying load.


1. Cloud Run — Horizontal Auto-Scaling​

All API services run on Google Cloud Run with Knative autoscaling. Scaling is driven by request concurrency — when the number of concurrent requests per instance approaches the configured limit, Cloud Run automatically adds more instances.

Scaling Configuration per Service​

# infra/environments/production/cloudrun.tf (excerpt)

# Each service defines min/max instances and concurrency
annotations = {
"autoscaling.knative.dev/minScale" = "1" # Always-warm instances
"autoscaling.knative.dev/maxScale" = "10" # Maximum burst capacity
}
container_concurrency = 50 # Requests per instance before scale-out

Production Scaling Parameters​

ServiceMin InstancesMax InstancesConcurrency/InstanceWhy
crm11050Standard API traffic
livechat21050Real-time chat requires extra warm capacity
cdp11050Standard API traffic
cms11020Long-running content sync (3600s timeout)
comms11050Standard API traffic
configuration11050Central config lookups
product-feed11050Batch product sync (3600s timeout, 2Gi memory)
integrations11050Lightweight Node.js service (512Mi)
jsltservice11050Stateless transformation engine
frontend11050Static Angular SPA serving
app-ui11050Static UI component serving

How Auto-Scaling Works​

Request arrives
|
v
Cloud Run checks: current_requests / container_concurrency
|
v
If ratio approaches 1.0 -> spin up new instance (up to maxScale)
If ratio drops -> scale down (but never below minScale)
|
v
New instances are ready in ~1-2 seconds (container already built)

Key behaviors:

  • Scale-to-zero is disabled: minScale = 1 (or 2 for livechat) ensures no cold starts
  • Concurrency-based: Not CPU/memory — a single instance handles 50 concurrent requests before triggering scale-out
  • Sub-second scaling: Cloud Run pre-warms containers for fast scaling
  • Per-service isolation: Each service scales independently based on its own traffic

2. Database Scaling​

Vertical Scaling​

Cloud SQL scales vertically via the tier parameter:

# Production
tier = "db-custom-1-3840" # 1 vCPU, 3.84 GB RAM

# Can be upgraded to e.g.:
# tier = "db-custom-2-7680" # 2 vCPU, 7.68 GB RAM
# tier = "db-custom-4-15360" # 4 vCPU, 15.36 GB RAM

Storage Auto-Scaling​

# infra/modules/database/main.tf

disk_autoresize = true # Automatically grow disk
disk_autoresize_limit = 0 # No upper limit
disk_type = "PD_SSD" # SSD for performance

Storage grows automatically as data increases — no manual intervention needed.

Connection Scaling​

ConnectionStrings: "...Pooling=true;Maximum Pool Size=1024;"

Each service instance maintains a pool of up to 1024 database connections, ensuring burst traffic doesn't exhaust connections.


3. Redis Cache Scaling​

Currently configured as a 1GB BASIC tier Memorystore instance:

# infra/modules/redis/main.tf

memory_size_gb = 1
tier = "BASIC"
read_replicas_mode = "READ_REPLICAS_DISABLED"

Scaling path:

  • Vertical: Increase memory_size_gb (up to 300GB)
  • Read replicas: Switch to STANDARD tier and enable READ_REPLICAS_ENABLED for read-heavy workloads
  • Application-level: Cache expiration (15 min sliding / 24 hr absolute) keeps memory usage bounded

4. Pub/Sub — Automatic Scaling​

Google Cloud Pub/Sub scales automatically with no configuration needed:

  • Throughput: Pub/Sub can handle millions of messages per second
  • Storage: 7-day message retention per subscription
  • Delivery: Push subscriptions deliver to Cloud Run, which auto-scales to process messages
# Each subscription has independent retry and delivery settings
retry_policy = {
minimum_backoff = "10s"
maximum_backoff = "600s"
}

Pub/Sub is a fully managed service — there are no instances to scale.


5. Cloud Tasks — Rate-Controlled Scaling​

Cloud Tasks provides controlled scaling to prevent downstream overload:

# infra/environments/production/cloudtask.tf

max_concurrent = 1000 # Max parallel task executions
max_dispatches = 500 # Tasks dispatched per second
QueueMax ConcurrentMax Dispatches/secPurpose
slack_genai1000500Slack + GenAI processing
event_queue1000500General event dispatch
product_feed_queue1000500Product feed operations
livechat_queue1000500Livechat task processing

This acts as a backpressure mechanism — even if 10,000 tasks are enqueued, only 1000 run concurrently and 500 new tasks start per second.


6. VPC Access Connector Scaling​

The serverless VPC connector that bridges Cloud Run to the VPC (for Redis, Cloud SQL, NAT) also scales:

# infra/modules/vpc_network/access_connector.tf

min_instances = 2 # Always available
max_instances = 10 # Scale for throughput
min_throughput = 200 # Mbps baseline
max_throughput = 1000 # Mbps maximum

7. GKE Autopilot — Fully Managed Scaling​

Background services on GKE Autopilot scale automatically:

# infra/modules/gke/main.tf

enable_autopilot = true # Google manages all scaling
  • Google provisions and removes nodes as needed
  • Pod scheduling is automatic based on resource requests
  • No manual node pool management

8. Workflow Batch Scaling​

Data sync workflows (Product Feed, CDP) handle scale via batching:

# Workflow pattern: create batches -> parallel sync

- create_batches_request: # Splits data into manageable chunks
- sync_batches:
for:
value: v
in: $${files} # Process each batch sequentially
steps:
- sync_batch:
retry:
max_retries: 2
backoff:
initial_delay: 2
max_delay: 10

The workflow engine handles batch orchestration, retries, and error tracking — individual batches that fail are logged and reported without blocking the remaining batches.


9. Traffic Load Balancing​

All traffic entering the platform is properly load-balanced between the minimum safe number of locations for the redundancy mechanism used. Every request passes through multiple load-balancing layers before reaching the application.

Full Request Path​

Client (Browser / API consumer)
|
v
Cloudflare (Edge — Global Anycast Network)
| - TLS termination
| - DDoS protection
| - WAF (Web Application Firewall)
| - Global load balancing across edge PoPs
|
v
Google Cloud Run (per-service HTTP Load Balancer)
| - Distributes across instances within europe-west4
| - Multi-zone instance placement
|
+--- Instance 1 (zone a)
+--- Instance 2 (zone b)
+--- Instance 3 (zone c)
|
+--- VPC Access Connector (2-10 instances) --+
| |
v v
Cloud SQL (Regional HA) Redis (Direct Peering)
|
v
Cloud NAT (static IP) --> External APIs

9.1 Cloudflare — Edge Load Balancing & Protection​

All inbound traffic to the CM Marketplace platform passes through Cloudflare before reaching GCP. Cloudflare acts as the outermost load-balancing and protection layer.

What Cloudflare provides:

CapabilityDescription
Global Anycast NetworkRequests are routed to the nearest Cloudflare edge PoP (Point of Presence) from 300+ data centers worldwide, reducing latency
TLS TerminationCloudflare terminates the client TLS connection at the edge, then establishes a separate encrypted connection to the GCP origin
DDoS ProtectionAutomatic Layer 3/4/7 DDoS mitigation — malicious traffic is absorbed at the edge before it reaches GCP
Web Application Firewall (WAF)Filters malicious requests (SQL injection, XSS, etc.) at the edge
Load BalancingDistributes traffic across healthy origins; performs health checks on backend services
Connection PoolingCloudflare maintains persistent connections to the origin, reducing TLS handshake overhead
Automatic FailoverIf a backend becomes unhealthy, Cloudflare routes traffic to healthy endpoints

Cloudflare timeout handling in application code:

The platform explicitly handles Cloudflare-specific behavior. For example, the Product Feed Shopify adapter handles Cloudflare's HTTP 524 ("A Timeout Occurred") response:

// src/app-product-feed/.../Services/ShopifyV2/ShopifyV2Service.cs

// 524 is Cloudflare's "A Timeout Occurred" — the Product Feed service is still processing
// the batch but Cloudflare closed the connection before it finished. The batch will complete
// server-side, so we return 200 to GCP Workflow to prevent an unnecessary retry of this batch.
if (response.StatusCode == (HttpStatusCode)524)
return new Response<object> { StatusCode = HttpStatusCode.OK };

This demonstrates that:

  1. All HTTP traffic flows through Cloudflare
  2. Long-running operations (like product feed sync with 3600s timeout) may exceed Cloudflare's connection timeout
  3. The application is designed to handle this gracefully — Cloudflare closing the connection does not cancel the server-side processing

Cloudflare's role in the load-balancing chain:

Client --> Cloudflare Edge PoP (nearest) --> Cloudflare Origin LB --> GCP Cloud Run LB --> Service Instance

Cloudflare performs the first level of load balancing by routing each client to the nearest edge PoP. It then forwards the request to the GCP origin, where Cloud Run's built-in load balancer handles the second level of distribution across service instances.

9.2 Cloud Run — Built-in HTTP Load Balancing​

Google Cloud Run includes a fully managed HTTP(S) load balancer that automatically distributes incoming requests across all running instances of a service. No separate load balancer resource is needed.

Cloudflare --> Cloud Run HTTPS Endpoint
|
v
Cloud Run Load Balancer (automatic, per-service)
|
+--- Instance 1 (zone a)
+--- Instance 2 (zone b)
+--- Instance 3 (zone c)
+--- ... up to maxScale (10)

How it works:

  • Each Cloud Run service gets a dedicated HTTPS endpoint (e.g., https://crm-xxxxx.run.app)
  • Google's Front End (GFE) receives the request from Cloudflare and routes it to the Cloud Run load balancer
  • The load balancer distributes requests across all healthy instances using a least-connections algorithm
  • Instances are spread across multiple availability zones within europe-west4 (zones a, b, c)
  • When minScale >= 2 (e.g., livechat), instances are guaranteed to be in different zones for redundancy

Traffic configuration in Terraform:

# infra/modules/cloudrun/main.tf

traffic {
percent = 100
latest_revision = true # All traffic goes to the latest healthy revision
}

This means:

  • 100% of traffic is routed to the latest deployed revision
  • Previous revisions remain available but receive 0% traffic unless explicitly rerouted (for rollback)
  • Traffic shifts to a new revision only after health checks pass

Per-service load balancing guarantees:

ServiceMin InstancesMin ZonesLoad Balancing
crm11+Automatic across all instances
livechat22 (spread)Automatic, multi-zone guaranteed
cdp11+Automatic across all instances
cms11+Automatic across all instances
comms11+Automatic across all instances
configuration11+Automatic across all instances
product-feed11+Automatic across all instances
integrations11+Automatic across all instances
jsltservice11+Automatic across all instances
frontend11+Automatic across all instances
app-ui11+Automatic across all instances

When traffic increases and instances scale from 1 to 10, the load balancer automatically includes the new instances in the rotation — no configuration change required.

9.3 GKE Background Service — Kubernetes LoadBalancer​

The background service (bgservice) runs on GKE Autopilot and is exposed via a Kubernetes LoadBalancer service, which provisions a GCP Network Load Balancer in front of the pods.

# infra/environments/production/workload.tf

module "workloads" {
workloads = [
{
name = "bgservice"
replicas = 1
container_port = 8080
service_port = 80
service_type = "LoadBalancer" # Provisions a GCP Network LB
}
]
}
# infra/modules/workload/main.tf

resource "kubernetes_service" "this" {
for_each = { for w in var.workloads : w.name => w }
spec {
selector = {
app = each.value.app_label
}
port {
port = each.value.service_port
target_port = each.value.container_port
}
type = each.value.service_type # "LoadBalancer"
}
}

How it works:

  • Kubernetes creates a GCP Network Load Balancer with an external IP
  • The load balancer distributes traffic across all pods matching the app=bgservice selector
  • GKE Autopilot schedules pods across multiple zones in the regional cluster
  • If a pod fails health checks, the load balancer stops routing traffic to it

9.4 VPC Access Connector — Distributed Egress​

Outbound traffic from Cloud Run services to internal resources (Redis, Cloud SQL) and external APIs (via NAT) is load-balanced across the VPC Access Connector instances:

# infra/modules/vpc_network/access_connector.tf

resource "google_vpc_access_connector" "shared-nat" {
min_instances = 2 # Minimum 2 instances for redundancy
max_instances = 10 # Scale up under load
min_throughput = 200 # 200 Mbps baseline
max_throughput = 1000 # 1000 Mbps maximum
}
  • Minimum 2 instances ensures egress traffic is always distributed across at least 2 connector instances
  • Traffic is balanced automatically — no sticky sessions
  • If one connector instance fails, the remaining instance(s) absorb the traffic

9.5 Cloud SQL — Connection Distribution​

While Cloud SQL itself is not load-balanced (single writer, with standby for failover), application-level connection pooling distributes query load efficiently:

ConnectionStrings: "...Pooling=true;Maximum Pool Size=1024;"
  • Each Cloud Run instance maintains its own connection pool (up to 1024 connections)
  • With 10 instances of a service running, up to 10,240 concurrent connections can be distributed across the pool
  • Cloud SQL's internal connection handling distributes queries across available compute resources

9.6 Pub/Sub — Event Distribution​

Pub/Sub push subscriptions deliver messages to Cloud Run endpoints, which are themselves load-balanced:

Publisher -> Pub/Sub Topic (Events)
|
v
Subscription filter (attribute-based routing)
|
v
Push delivery -> Cloud Run endpoint (load-balanced across instances)
  • Each push delivery is an independent HTTP POST to the Cloud Run service URL
  • Cloud Run's load balancer distributes these deliveries across all service instances
  • Multiple messages are delivered in parallel — Pub/Sub does not wait for one delivery to complete before sending the next

9.7 Cloud NAT — Distributed Egress​

# infra/modules/vpc_network/router.tf

resource "google_compute_router_nat" "shared-nat" {
nat_ip_allocate_option = "MANUAL_ONLY"
nat_ips = [google_compute_address.static-marketplace.id]
}

Cloud NAT is a managed, distributed service — it is not a single VM or instance. Google distributes NAT processing across its infrastructure automatically. All outbound connections share the static IP but are processed in parallel with no single point of failure.


Summary​

ComponentScaling TypeMinMaxTriggerLoad Balancing
Cloudflare (edge)Global anycast300+ PoPsUnlimitedAutomaticGlobal edge LB, DDoS protection
Cloud Run servicesHorizontal auto-scale1-2 instances10 instancesRequest concurrencyBuilt-in HTTP LB, multi-zone
GKE bgserviceKubernetes-managed1 podAutopilotPod resourcesGCP Network LoadBalancer
Cloud SQLVertical (tier upgrade)db-custom-1-3840ConfigurableManualN/A (single writer + failover)
Cloud SQL storageAutomatic disk growth20 GB SSDUnlimitedDisk usageN/A
RedisVertical (memory resize)1 GB300 GBManualN/A (single instance)
Pub/SubFully automatic—UnlimitedMessage volumeBuilt-in (managed service)
Cloud TasksRate-limited dispatch—1000 concurrentTask queue depthBuilt-in (managed service)
VPC ConnectorHorizontal auto-scale2 instances10 instancesNetwork throughputAutomatic across instances
Cloud NATDistributed (managed)——AutomaticDistributed (no SPOF)

Last updated: May 2026