Why Kubernetes Autoscaling Scales Up but Not Down
Scale-up problems are loud: latency goes up, someone gets paged. Scale-down problems are quiet: the cluster stays big after a traffic peak, and you find out from the bill. Autoscaling in Kubernetes works in two layers, and either one can get stuck.
- Pods: the Horizontal Pod Autoscaler (HPA), or KEDA, which creates and drives an HPA.
- Nodes: Cluster Autoscaler or Karpenter.
Nodes can only go away after pods go away and the remaining pods can be moved. So debug from the top: first the HPA, then the nodes.
Layer 1: the HPA doesn't reduce replicas
The HPA computes desiredReplicas = ceil(currentReplicas × currentMetric / targetMetric). It scales down only when that number is lower than the current count. Common reasons it isn't:
Defaults are slower than people expect. Scale-down has a 300-second stabilization window: the HPA uses the highest recommendation from the last five minutes. There is also a 10% tolerance around the target where nothing happens. That is a delay, not a failure, but it often gets mistaken for one.
Utilization is measured against requests. CPU utilization is usage divided by the pod's CPU request. If requests are set too low, utilization looks high at normal load and the HPA never sees a reason to shrink.
One metric stays high. With several metrics, the HPA picks the highest replica count among them. Memory is the usual culprit: JVM heaps, caches and allocators rarely give memory back, so a memory target keeps replicas up long after CPU drops. Scale on CPU, request rate or queue depth, and leave memory to requests and limits.
Missing metrics block scale-down. If some pods have no metrics, the HPA assumes they use 100% of the target when it considers scaling down. A broken metrics pipeline or pods that never report make it conservative.
Something else owns the replica count. If spec.replicas is set in Git and Argo CD or Flux syncs it, every sync resets the Deployment. Remove replicas from the manifest when an HPA manages it.
It is configured not to. Check minReplicas and look for scaleDown.selectPolicy: Disabled copied from somewhere.
External metrics don't return to baseline. A queue that never drains because of retries, or a Prometheus query that sums a counter instead of a rate.
Debug it:
kubectl describe hpa my-app
kubectl get hpa my-app -o yaml # status.conditions and status.currentMetrics
kubectl top pods -l app=my-app
Look at the conditions. AbleToScale, ScalingActive and ScalingLimited usually tell you exactly what is happening, for example "the desired replica count is less than the minimum replica count".
An explicit HPA, so nobody has to remember the defaults:
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: my-app
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: my-app
minReplicas: 2
maxReplicas: 30
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70
behavior:
scaleDown:
stabilizationWindowSeconds: 300
policies:
- type: Percent
value: 20
periodSeconds: 60
Use autoscaling/v2. The v2beta2 API was removed in Kubernetes 1.26.
KEDA specifics
KEDA's cooldownPeriod (default 300 seconds) only applies to scaling from one replica to zero. Between one and N replicas the HPA it creates is in charge, so tune scale-down through advanced.horizontalPodAutoscalerConfig.behavior. minReplicaCount defaults to 0, which is fine for queue workers but not for services that must answer requests.
Layer 2: nodes don't go away
Cluster Autoscaler
A node becomes a scale-down candidate when the sum of CPU and memory requests of its pods is below 50% of allocatable (--scale-down-utilization-threshold). It is removed after it has been unneeded for 10 minutes (--scale-down-unneeded-time), and only if all its pods can be moved. Note that this is requests, not actual usage.
Pods that block removal:
- Pods whose PodDisruptionBudget allows no disruption.
kube-systempods without a PDB (with default flags).- Pods not managed by a controller.
- Pods with local storage such as
emptyDir(with default flags). - Pods that fit nowhere else because of node selectors, affinity or lack of room.
- Pods annotated with
cluster-autoscaler.kubernetes.io/safe-to-evict: "false".
kubectl -n kube-system get configmap cluster-autoscaler-status -o yaml
kubectl -n kube-system logs deploy/cluster-autoscaler | grep -i "scale-down\|cannot be removed"
Karpenter
Karpenter consolidates according to the NodePool's disruption settings. The default consolidationPolicy is WhenEmptyOrUnderutilized. With WhenEmpty, nodes running even one pod are kept. Consolidation is also limited by disruption budgets (default nodes: 10%), by pods annotated karpenter.sh/do-not-disrupt, and by blocking PDBs.
Karpenter publishes events on nodes explaining why they can't be consolidated:
kubectl get nodeclaims
kubectl describe node <node-name> | grep -A10 Events
PDBs that block everything
A PDB with minAvailable equal to the replica count, or maxUnavailable: 0, means no pod can ever be evicted voluntarily. The same happens with minAvailable: 1 on a single-replica Deployment. Such pods pin their node forever. Use maxUnavailable: 1 for most services, and make sure every critical service runs at least two replicas.
Alert on it
Scale-down failures don't page anyone unless you make them. Two signals from kube-state-metrics:
# HPA pinned at maxReplicas
kube_horizontalpodautoscaler_status_current_replicas
== on (namespace, horizontalpodautoscaler)
kube_horizontalpodautoscaler_spec_max_replicas
Fire it when the condition holds for an hour or more. Also track the ratio of requested to allocatable CPU across the cluster. A low ratio that stays low means nodes are not being removed.
Checklist
- HPA conditions explain the current state, check them first.
- Requests reflect real usage, otherwise utilization targets are meaningless.
- No memory-based scaling for runtimes that don't release memory.
- No
replicasin Git for HPA-managed workloads. - PDBs allow at least one disruption.
- No stray
safe-to-evict: "false"ordo-not-disruptannotations. - Alerts exist for HPAs pinned at max and for low cluster request utilization.
