Scaling Strategy
Scaling Strategy
This page describes how each component of the CM Marketplace platform scales to handle varying load.
1. Cloud Run — Horizontal Auto-Scaling
All API services run on Google Cloud Run with Knative autoscaling. Scaling is driven by request concurrency — when the number of concurrent requests per instance approaches the configured limit, Cloud Run automatically adds more instances.
Scaling Configuration per Service
# infra/environments/production/cloudrun.tf (excerpt)
# Each service defines min/max instances and concurrency
annotations = {
"autoscaling.knative.dev/minScale" = "1" # Always-warm instances
"autoscaling.knative.dev/maxScale" = "10" # Maximum burst capacity
}
container_concurrency = 50 # Requests per instance before scale-out
Production Scaling Parameters
| Service | Min Instances | Max Instances | Concurrency/Instance | Why |
|---|---|---|---|---|
| crm | 1 | 10 | 50 | Standard API traffic |
| livechat | 2 | 10 | 50 | Real-time chat requires extra warm capacity |
| cdp | 1 | 10 | 50 | Standard API traffic |
| cms | 1 | 10 | 20 | Long-running content sync (3600s timeout) |
| comms | 1 | 10 | 50 | Standard API traffic |
| configuration | 1 | 10 | 50 | Central config lookups |
| product-feed | 1 | 10 | 50 | Batch product sync (3600s timeout, 2Gi memory) |
| integrations | 1 | 10 | 50 | Lightweight Node.js service (512Mi) |
| jsltservice | 1 | 10 | 50 | Stateless transformation engine |
| frontend | 1 | 10 | 50 | Static Angular SPA serving |
| app-ui | 1 | 10 | 50 | Static UI component serving |
How Auto-Scaling Works
Request arrives
|
v
Cloud Run checks: current_requests / container_concurrency
|
v
If ratio approaches 1.0 -> spin up new instance (up to maxScale)
If ratio drops -> scale down (but never below minScale)
|
v
New instances are ready in ~1-2 seconds (container already built)
Key behaviors:
- Scale-to-zero is disabled:
minScale = 1(or 2 for livechat) ensures no cold starts - Concurrency-based: Not CPU/memory — a single instance handles 50 concurrent requests before triggering scale-out
- Sub-second scaling: Cloud Run pre-warms containers for fast scaling
- Per-service isolation: Each service scales independently based on its own traffic
2. Database Scaling
Vertical Scaling
Cloud SQL scales vertically via the tier parameter:
# Production
tier = "db-custom-1-3840" # 1 vCPU, 3.84 GB RAM
# Can be upgraded to e.g.:
# tier = "db-custom-2-7680" # 2 vCPU, 7.68 GB RAM
# tier = "db-custom-4-15360" # 4 vCPU, 15.36 GB RAM
Storage Auto-Scaling
# infra/modules/database/main.tf
disk_autoresize = true # Automatically grow disk
disk_autoresize_limit = 0 # No upper limit
disk_type = "PD_SSD" # SSD for performance
Storage grows automatically as data increases — no manual intervention needed.
Connection Scaling
ConnectionStrings: "...Pooling=true;Maximum Pool Size=1024;"
Each service instance maintains a pool of up to 1024 database connections, ensuring burst traffic doesn't exhaust connections.
3. Redis Cache Scaling
Currently configured as a 1GB BASIC tier Memorystore instance:
# infra/modules/redis/main.tf
memory_size_gb = 1
tier = "BASIC"
read_replicas_mode = "READ_REPLICAS_DISABLED"
Scaling path:
- Vertical: Increase
memory_size_gb(up to 300GB) - Read replicas: Switch to
STANDARDtier and enableREAD_REPLICAS_ENABLEDfor read-heavy workloads - Application-level: Cache expiration (15 min sliding / 24 hr absolute) keeps memory usage bounded
4. Pub/Sub — Automatic Scaling
Google Cloud Pub/Sub scales automatically with no configuration needed:
- Throughput: Pub/Sub can handle millions of messages per second
- Storage: 7-day message retention per subscription
- Delivery: Push subscriptions deliver to Cloud Run, which auto-scales to process messages
# Each subscription has independent retry and delivery settings
retry_policy = {
minimum_backoff = "10s"
maximum_backoff = "600s"
}
Pub/Sub is a fully managed service — there are no instances to scale.
5. Cloud Tasks — Rate-Controlled Scaling
Cloud Tasks provides controlled scaling to prevent downstream overload:
# infra/environments/production/cloudtask.tf
max_concurrent = 1000 # Max parallel task executions
max_dispatches = 500 # Tasks dispatched per second
| Queue | Max Concurrent | Max Dispatches/sec | Purpose |
|---|---|---|---|
| slack_genai | 1000 | 500 | Slack + GenAI processing |
| event_queue | 1000 | 500 | General event dispatch |
| product_feed_queue | 1000 | 500 | Product feed operations |
| livechat_queue | 1000 | 500 | Livechat task processing |
This acts as a backpressure mechanism — even if 10,000 tasks are enqueued, only 1000 run concurrently and 500 new tasks start per second.
6. VPC Access Connector Scaling
The serverless VPC connector that bridges Cloud Run to the VPC (for Redis, Cloud SQL, NAT) also scales:
# infra/modules/vpc_network/access_connector.tf
min_instances = 2 # Always available
max_instances = 10 # Scale for throughput
min_throughput = 200 # Mbps baseline
max_throughput = 1000 # Mbps maximum
7. GKE Autopilot — Fully Managed Scaling
Background services on GKE Autopilot scale automatically:
# infra/modules/gke/main.tf
enable_autopilot = true # Google manages all scaling
- Google provisions and removes nodes as needed
- Pod scheduling is automatic based on resource requests
- No manual node pool management
8. Workflow Batch Scaling
Data sync workflows (Product Feed, CDP) handle scale via batching:
# Workflow pattern: create batches -> parallel sync
- create_batches_request: # Splits data into manageable chunks
- sync_batches:
for:
value: v
in: $${files} # Process each batch sequentially
steps:
- sync_batch:
retry:
max_retries: 2
backoff:
initial_delay: 2
max_delay: 10
The workflow engine handles batch orchestration, retries, and error tracking — individual batches that fail are logged and reported without blocking the remaining batches.
9. Traffic Load Balancing
All traffic entering the platform is properly load-balanced between the minimum safe number of locations for the redundancy mechanism used. Every request passes through multiple load-balancing layers before reaching the application.
Full Request Path
Client (Browser / API consumer)
|
v
Cloudflare (Edge — Global Anycast Network)
| - TLS termination
| - DDoS protection
| - WAF (Web Application Firewall)
| - Global load balancing across edge PoPs
|
v
Google Cloud Run (per-service HTTP Load Balancer)
| - Distributes across instances within europe-west4
| - Multi-zone instance placement
|
+--- Instance 1 (zone a)
+--- Instance 2 (zone b)
+--- Instance 3 (zone c)
|
+--- VPC Access Connector (2-10 instances) --+
| |
v v
Cloud SQL (Regional HA) Redis (Direct Peering)
|
v
Cloud NAT (static IP) --> External APIs
9.1 Cloudflare — Edge Load Balancing & Protection
All inbound traffic to the CM Marketplace platform passes through Cloudflare before reaching GCP. Cloudflare acts as the outermost load-balancing and protection layer.
What Cloudflare provides:
| Capability | Description |
|---|---|
| Global Anycast Network | Requests are routed to the nearest Cloudflare edge PoP (Point of Presence) from 300+ data centers worldwide, reducing latency |
| TLS Termination | Cloudflare terminates the client TLS connection at the edge, then establishes a separate encrypted connection to the GCP origin |
| DDoS Protection | Automatic Layer 3/4/7 DDoS mitigation — malicious traffic is absorbed at the edge before it reaches GCP |
| Web Application Firewall (WAF) | Filters malicious requests (SQL injection, XSS, etc.) at the edge |
| Load Balancing | Distributes traffic across healthy origins; performs health checks on backend services |
| Connection Pooling | Cloudflare maintains persistent connections to the origin, reducing TLS handshake overhead |
| Automatic Failover | If a backend becomes unhealthy, Cloudflare routes traffic to healthy endpoints |
Cloudflare timeout handling in application code:
The platform explicitly handles Cloudflare-specific behavior. For example, the Product Feed Shopify adapter handles Cloudflare's HTTP 524 ("A Timeout Occurred") response:
// src/app-product-feed/.../Services/ShopifyV2/ShopifyV2Service.cs
// 524 is Cloudflare's "A Timeout Occurred" — the Product Feed service is still processing
// the batch but Cloudflare closed the connection before it finished. The batch will complete
// server-side, so we return 200 to GCP Workflow to prevent an unnecessary retry of this batch.
if (response.StatusCode == (HttpStatusCode)524)
return new Response<object> { StatusCode = HttpStatusCode.OK };
This demonstrates that:
- All HTTP traffic flows through Cloudflare
- Long-running operations (like product feed sync with 3600s timeout) may exceed Cloudflare's connection timeout
- The application is designed to handle this gracefully — Cloudflare closing the connection does not cancel the server-side processing
Cloudflare's role in the load-balancing chain:
Client --> Cloudflare Edge PoP (nearest) --> Cloudflare Origin LB --> GCP Cloud Run LB --> Service Instance
Cloudflare performs the first level of load balancing by routing each client to the nearest edge PoP. It then forwards the request to the GCP origin, where Cloud Run's built-in load balancer handles the second level of distribution across service instances.
9.2 Cloud Run — Built-in HTTP Load Balancing
Google Cloud Run includes a fully managed HTTP(S) load balancer that automatically distributes incoming requests across all running instances of a service. No separate load balancer resource is needed.
Cloudflare --> Cloud Run HTTPS Endpoint
|
v
Cloud Run Load Balancer (automatic, per-service)
|
+--- Instance 1 (zone a)
+--- Instance 2 (zone b)
+--- Instance 3 (zone c)
+--- ... up to maxScale (10)
How it works:
- Each Cloud Run service gets a dedicated HTTPS endpoint (e.g.,
https://crm-xxxxx.run.app) - Google's Front End (GFE) receives the request from Cloudflare and routes it to the Cloud Run load balancer
- The load balancer distributes requests across all healthy instances using a least-connections algorithm
- Instances are spread across multiple availability zones within europe-west4 (zones a, b, c)
- When
minScale >= 2(e.g., livechat), instances are guaranteed to be in different zones for redundancy
Traffic configuration in Terraform:
# infra/modules/cloudrun/main.tf
traffic {
percent = 100
latest_revision = true # All traffic goes to the latest healthy revision
}
This means:
- 100% of traffic is routed to the latest deployed revision
- Previous revisions remain available but receive 0% traffic unless explicitly rerouted (for rollback)
- Traffic shifts to a new revision only after health checks pass
Per-service load balancing guarantees:
| Service | Min Instances | Min Zones | Load Balancing |
|---|---|---|---|
| crm | 1 | 1+ | Automatic across all instances |
| livechat | 2 | 2 (spread) | Automatic, multi-zone guaranteed |
| cdp | 1 | 1+ | Automatic across all instances |
| cms | 1 | 1+ | Automatic across all instances |
| comms | 1 | 1+ | Automatic across all instances |
| configuration | 1 | 1+ | Automatic across all instances |
| product-feed | 1 | 1+ | Automatic across all instances |
| integrations | 1 | 1+ | Automatic across all instances |
| jsltservice | 1 | 1+ | Automatic across all instances |
| frontend | 1 | 1+ | Automatic across all instances |
| app-ui | 1 | 1+ | Automatic across all instances |
When traffic increases and instances scale from 1 to 10, the load balancer automatically includes the new instances in the rotation — no configuration change required.
9.3 GKE Background Service — Kubernetes LoadBalancer
The background service (bgservice) runs on GKE Autopilot and is exposed via a Kubernetes LoadBalancer service, which provisions a GCP Network Load Balancer in front of the pods.
# infra/environments/production/workload.tf
module "workloads" {
workloads = [
{
name = "bgservice"
replicas = 1
container_port = 8080
service_port = 80
service_type = "LoadBalancer" # Provisions a GCP Network LB
}
]
}
# infra/modules/workload/main.tf
resource "kubernetes_service" "this" {
for_each = { for w in var.workloads : w.name => w }
spec {
selector = {
app = each.value.app_label
}
port {
port = each.value.service_port
target_port = each.value.container_port
}
type = each.value.service_type # "LoadBalancer"
}
}
How it works:
- Kubernetes creates a GCP Network Load Balancer with an external IP
- The load balancer distributes traffic across all pods matching the
app=bgserviceselector - GKE Autopilot schedules pods across multiple zones in the regional cluster
- If a pod fails health checks, the load balancer stops routing traffic to it
9.4 VPC Access Connector — Distributed Egress
Outbound traffic from Cloud Run services to internal resources (Redis, Cloud SQL) and external APIs (via NAT) is load-balanced across the VPC Access Connector instances:
# infra/modules/vpc_network/access_connector.tf
resource "google_vpc_access_connector" "shared-nat" {
min_instances = 2 # Minimum 2 instances for redundancy
max_instances = 10 # Scale up under load
min_throughput = 200 # 200 Mbps baseline
max_throughput = 1000 # 1000 Mbps maximum
}
- Minimum 2 instances ensures egress traffic is always distributed across at least 2 connector instances
- Traffic is balanced automatically — no sticky sessions
- If one connector instance fails, the remaining instance(s) absorb the traffic
9.5 Cloud SQL — Connection Distribution
While Cloud SQL itself is not load-balanced (single writer, with standby for failover), application-level connection pooling distributes query load efficiently:
ConnectionStrings: "...Pooling=true;Maximum Pool Size=1024;"
- Each Cloud Run instance maintains its own connection pool (up to 1024 connections)
- With 10 instances of a service running, up to 10,240 concurrent connections can be distributed across the pool
- Cloud SQL's internal connection handling distributes queries across available compute resources
9.6 Pub/Sub — Event Distribution
Pub/Sub push subscriptions deliver messages to Cloud Run endpoints, which are themselves load-balanced:
Publisher -> Pub/Sub Topic (Events)
|
v
Subscription filter (attribute-based routing)
|
v
Push delivery -> Cloud Run endpoint (load-balanced across instances)
- Each push delivery is an independent HTTP POST to the Cloud Run service URL
- Cloud Run's load balancer distributes these deliveries across all service instances
- Multiple messages are delivered in parallel — Pub/Sub does not wait for one delivery to complete before sending the next
9.7 Cloud NAT — Distributed Egress
# infra/modules/vpc_network/router.tf
resource "google_compute_router_nat" "shared-nat" {
nat_ip_allocate_option = "MANUAL_ONLY"
nat_ips = [google_compute_address.static-marketplace.id]
}
Cloud NAT is a managed, distributed service — it is not a single VM or instance. Google distributes NAT processing across its infrastructure automatically. All outbound connections share the static IP but are processed in parallel with no single point of failure.
Summary
| Component | Scaling Type | Min | Max | Trigger | Load Balancing |
|---|---|---|---|---|---|
| Cloudflare (edge) | Global anycast | 300+ PoPs | Unlimited | Automatic | Global edge LB, DDoS protection |
| Cloud Run services | Horizontal auto-scale | 1-2 instances | 10 instances | Request concurrency | Built-in HTTP LB, multi-zone |
| GKE bgservice | Kubernetes-managed | 1 pod | Autopilot | Pod resources | GCP Network LoadBalancer |
| Cloud SQL | Vertical (tier upgrade) | db-custom-1-3840 | Configurable | Manual | N/A (single writer + failover) |
| Cloud SQL storage | Automatic disk growth | 20 GB SSD | Unlimited | Disk usage | N/A |
| Redis | Vertical (memory resize) | 1 GB | 300 GB | Manual | N/A (single instance) |
| Pub/Sub | Fully automatic | — | Unlimited | Message volume | Built-in (managed service) |
| Cloud Tasks | Rate-limited dispatch | — | 1000 concurrent | Task queue depth | Built-in (managed service) |
| VPC Connector | Horizontal auto-scale | 2 instances | 10 instances | Network throughput | Automatic across instances |
| Cloud NAT | Distributed (managed) | — | — | Automatic | Distributed (no SPOF) |
Last updated: May 2026