Predictive Autoscaling for Stateful Applications on Kubernetes
Reactive autoscaling works for stateless services because a new pod is useful seconds after it starts. A new replica of a database, broker or cache is not. It has to get data first, and while it does, it often adds load to the existing replicas. By the time a CPU-based HPA reacts to a peak, the new replica is ready after the peak has passed.
Predictive scaling fixes the timing: add capacity before the load arrives.
What adding a replica actually does
Be clear what a new replica gives you:
- PostgreSQL or MySQL read replica: more read capacity once the base backup is restored and replication has caught up. Write capacity doesn't change.
- Kafka broker: nothing, until partitions are reassigned to it. Reassignment copies data and loads the existing brokers. Strimzi with Cruise Control can automate the rebalance.
- Cassandra node: streams its token ranges from existing nodes when it bootstraps. Old nodes need
nodetool cleanupafterwards. - Redis Cluster node: empty until hash slots are migrated to it.
- Elasticsearch or OpenSearch node: shards rebalance automatically, which is disk and network load on the whole cluster.
So scaling out is a heavy operation that is often better triggered once, early, than repeatedly in reaction to metrics.
Measure your lead time
Lead time is the time from "scale decision" to "new replica serves traffic at full capacity". Measure it in a staging environment with production-sized data: pod scheduling, volume provisioning, data sync, cache warmup. Lead time is your forecast horizon. If it takes 25 minutes, a forecast for "now" is useless; you need one for 30 minutes from now.
Also check the StatefulSet's podManagementPolicy. The default OrderedReady creates pods one at a time and waits for each to be Ready, so adding three replicas takes three times as long. Parallel creates them together, if the application can handle several members joining at once.
Start with schedules
Most load that can be predicted follows the calendar: business hours, nightly batch jobs, known campaigns. A schedule covers this with no model at all. KEDA's cron scaler sets a replica floor for a time window:
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
name: my-db-readers
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: StatefulSet
name: my-db-readers
minReplicaCount: 3
maxReplicaCount: 8
advanced:
horizontalPodAutoscalerConfig:
behavior:
scaleDown:
stabilizationWindowSeconds: 1800
policies:
- type: Pods
value: 1
periodSeconds: 600
triggers:
- type: cron
metadata:
timezone: Europe/Berlin
start: 30 7 * * 1-5
end: 0 20 * * 1-5
desiredReplicas: "6"
Start the window earlier than the load by your lead time. Scale-down is deliberately slow, one replica every ten minutes after a 30-minute stabilization window. That is acceptable for read replicas, which hold no unique data. Systems where each member owns data need more care, see below.
Add a forecast when schedules aren't enough
When load has seasonality but also trends and irregular days, a time series forecast helps. Prophet is a reasonable starting point. The package is called prophet now, the old fbprophet name is no longer maintained.
import pandas as pd
from prophet import Prophet
from prometheus_client import CollectorRegistry, Gauge, push_to_gateway
LEAD = pd.Timedelta(minutes=30)
# history: columns ds (UTC timestamp, 15-minute steps) and y (requests per second)
df = pd.read_csv("qps_history.csv", parse_dates=["ds"])
model = Prophet(daily_seasonality=True, weekly_seasonality=True)
model.fit(df)
future = model.make_future_dataframe(periods=8, freq="15min")
forecast = model.predict(future)
target = pd.Timestamp.now(tz="UTC").tz_localize(None) + LEAD
value = forecast.loc[forecast["ds"] >= target, "yhat_upper"].iloc[0]
registry = CollectorRegistry()
gauge = Gauge("my_app_qps_forecast", "Forecast QPS at now + lead time",
["horizon"], registry=registry)
gauge.labels(horizon="30m").set(max(value, 0))
push_to_gateway("pushgateway.monitoring.svc:9091", job="qps-forecast", registry=registry)
Run it every few minutes as a CronJob. Two choices matter:
- Use the upper bound (
yhat_upper), not the point forecast. Being under capacity for a stateful system costs much more than one extra replica. - The Pushgateway keeps the last value forever. If the job stops, the forecast freezes. Alert on the age of
push_time_secondsfor this job.
Combine forecast and reality
Never let the forecast scale below what current traffic needs. Add both as triggers to the ScaledObject above, next to the cron trigger. With several triggers the HPA takes the highest replica count, so the forecast can only add capacity:
triggers:
- type: prometheus
metadata:
serverAddress: http://prometheus.monitoring.svc:9090
query: max(my_app_qps_forecast{horizon="30m"})
threshold: "500"
- type: prometheus
metadata:
serverAddress: http://prometheus.monitoring.svc:9090
query: sum(rate(my_app_requests_total[2m]))
threshold: "500"
threshold is the load one replica handles at your target utilization, measured in a load test. KEDA uses it as a per-replica average, so desired replicas are roughly total load divided by threshold.
If the stateful system runs on VMs instead of Kubernetes, EC2 Auto Scaling predictive scaling and Compute Engine predictive autoscaling for managed instance groups give you a built-in forecast.
Readiness must mean "caught up"
A pod that is Running but still syncing must not receive traffic. Make the readiness probe check the thing that matters: replication lag under a threshold, bootstrap finished, cache warm. Set minReadySeconds so a replica that flaps right after joining doesn't count as available.
Scale-down is a different problem
Removing a stateful replica can lose data or availability if done carelessly.
- A StatefulSet removes the highest ordinal first. That pod must hold nothing that isn't replicated elsewhere.
- Data must move off before the pod goes:
nodetool decommissionfor Cassandra, partition reassignment for Kafka, slot migration for Redis. ApreStophook is limited byterminationGracePeriodSecondsand is the wrong place for long data moves. This is a job for an operator. - By default, PVCs are kept when a StatefulSet scales down (
persistentVolumeClaimRetentionPolicy.whenScaled: Retain). That is safe, but the disks keep costing money until someone cleans them up. - A PodDisruptionBudget keeps node drains from taking out a quorum.
For most stateful systems, scale up automatically and scale down deliberately: on a schedule, through the operator, during low traffic, or after a human approves it.
Storage settings
Use a CSI provisioner and let volumes be created in the zone where the pod lands:
apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
name: db-ssd
provisioner: pd.csi.storage.gke.io
parameters:
type: pd-ssd
reclaimPolicy: Retain
volumeBindingMode: WaitForFirstConsumer
allowVolumeExpansion: true
The in-tree kubernetes.io/gce-pd provisioner is deprecated in favor of the CSI driver.
Checklist
- Know what a new replica does for your system and what load it adds.
- Measure lead time, and forecast at least that far ahead.
- Start with cron-based scaling, add a forecast only where schedules miss.
- Scale on the maximum of forecast and current load, using the forecast's upper bound.
- Readiness probes check sync state, not just process health.
- Scale down through the operator or by hand, with data moved first.
