Scaling StatefulSets Without Overloading Volume Attach
A StatefulSet pod with a persistent volume does not start until its disk is created, attached to the node and mounted. Each of those steps calls the storage backend or the cloud API. When many of them happen at once, after a large scale-up, a node drain or a zone failure, they queue up, hit rate limits and time out. Pods sit in ContainerCreating and retries make the queue longer.
This post walks through the path from scale-up to running pod, where the bottlenecks are, and the settings that keep it manageable.
From replica count to running pod
- The StatefulSet controller creates a PVC from each
volumeClaimTemplatesentry, named<template>-<statefulset>-<ordinal>. PVCs are created once per ordinal and reused afterwards; scaling down does not delete them unless you set a retention policy. - With
volumeBindingMode: WaitForFirstConsumer, the scheduler picks a node first, then the CSI provisioner creates the disk in that node's zone (a cloud API call). - The attach/detach controller creates a VolumeAttachment. The CSI external-attacher asks the driver to attach the disk to the node (another cloud API call).
- The kubelet stages and mounts the volume. On first use, the filesystem is created. If the pod sets
fsGroup, the kubelet may change ownership of every file on the volume. - The container starts.
Steps 2 and 3 go through the cloud provider's API, which has rate limits per account and region. Step 4 can take long on large volumes with many files.
When storms happen
The default podManagementPolicy is OrderedReady. Pods are created one at a time, and each must be Running and Ready before the next one starts. A scale-up of a single StatefulSet under OrderedReady does not create a herd by itself. Storms usually come from:
podManagementPolicy: Parallel, where all new pods are created at once.- Many StatefulSets scaling or restarting together, for example after an operator upgrade or a cluster-wide rollout.
- Node drains and failures. Every pod on the node needs a detach from the old node and an attach to the new one.
- Node upgrades that replace many nodes in a short time.
- Per-node attach limits. Each instance type supports only a certain number of attached volumes. The CSI driver reports it in the CSINode object, and the scheduler respects it, so pods stay
Pendingonce nodes are full.
Detach after node failure
If a node becomes unreachable, its pods cannot finish terminating, since the kubelet never confirms it. The StatefulSet controller will not create a replacement until the old pod is confirmed gone, because two pods with the same identity must never run at once. The pod stays Terminating or Unknown, and its volume stays attached to the dead node. If someone force-deletes the pod by hand, the new pod can still hit a Multi-Attach error until the volume is detached from the old node.
If you know the node is really down, mark it out of service. Pods without a matching toleration are force-deleted and their volumes detached right away:
kubectl taint nodes <node-name> node.kubernetes.io/out-of-service=nodeshutdown:NoExecute
Only do this when the node is truly powered off or isolated, and remove the taint after recovery.
Keep scale-ups controlled
Keep OrderedReady unless you need Parallel. If your app needs Parallel for faster start, scale in steps and wait for readiness between them:
for n in 6 7 8 9 10; do
kubectl scale statefulset my-app --replicas="$n"
kubectl rollout status statefulset my-app --timeout=10m
done
rollout status returns when all replicas are ready, so each step waits for real progress instead of a fixed sleep.
Rate-limit autoscaling. If an HPA drives the StatefulSet, limit how fast it adds pods:
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: my-app
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: StatefulSet
name: my-app
minReplicas: 3
maxReplicas: 12
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70
behavior:
scaleUp:
policies:
- type: Pods
value: 1
periodSeconds: 120
For predictable peaks, scale ahead of time on a schedule instead of reacting to load. Stateful apps often need time to rebalance data after a new replica joins, which an autoscaler does not see.
Use minReadySeconds. With OrderedReady, the next pod waits until the previous one has been Ready for that long. This gives the app time to join its cluster before more members arrive.
Make each attach and mount cheaper
fsGroupChangePolicy: OnRootMismatchin the podsecurityContext. The kubelet then skips the recursive ownership change when the volume root already has the right owner and permissions. On large volumes this is often the biggest part of mount time.WaitForFirstConsumeron the StorageClass, so disks are created in the zone where the pod is scheduled and never need to be recreated or moved.- Spread replicas with
topologySpreadConstraintsacross nodes and zones. A single node drain or zone failure then moves one replica, not several. - PodDisruptionBudget with
maxUnavailable: 1, and drain nodes one at a time, so maintenance never triggers a mass reattach. - Check node attach limits before choosing instance types:
kubectl get csinode <node-name> -o jsonpath='{range .spec.drivers[*]}{.name}{"\t"}{.allocatable.count}{"\n"}{end}'
- CSI sidecar tuning. The external provisioner and attacher have flags for worker threads and API client rate limits. Defaults vary by driver and version; look at your driver's Helm chart values before changing them, and raise them only if the cloud API is not the bottleneck.
Monitor it
storage_operation_duration_seconds, a histogram withoperation_namevalues such asvolume_attach(attach/detach controller) andvolume_mount(kubelet). Watch p95 and error counts.- Events with reasons
FailedAttachVolume,FailedMountandProvisioningFailed. kubectl get volumeattachmentsduring an incident shows which attachments are stuck and on which node.- Cloud API throttling errors in the CSI controller logs.
Checklist
OrderedReadyby default; stepwise scaling with readiness checks if you useParallel.- HPA scale-up policies limit pods per period; planned peaks scaled ahead of time.
fsGroupChangePolicy: OnRootMismatchandWaitForFirstConsumereverywhere.- Replicas spread across nodes and zones, PDBs set, drains one node at a time.
- Node attach limits known for every instance type in use.
- Attach and mount latency alerting in place.
- Out-of-service taint procedure documented for confirmed dead nodes.
