Debugging Data Loss with Persistent Volumes in Kubernetes

Debugging Data Loss with Persistent Volumes in Kubernetes

Reading time1 min
#kubernetes#devops#stateful-apps#persistent-storage

Debugging Data Loss with Persistent Volumes in Kubernetes

When a stateful app on Kubernetes comes back empty, the data is often not gone. The pod may be looking at a new volume, the data may never have been on the volume, or the old volume may still exist in a Released state. Sometimes it really is deleted. This post covers how the pieces fit together, the common causes, and how to tell them apart quickly.

The pieces

  • A PersistentVolumeClaim (PVC) is a namespaced request for storage.
  • A PersistentVolume (PV) is the cluster-level object that represents the real disk, bound one to one to a PVC.
  • A StorageClass tells the CSI driver how to create new volumes, and sets two defaults that matter a lot: reclaimPolicy and volumeBindingMode.
  • The backing disk (EBS volume, Persistent Disk, Azure Disk, Ceph RBD image) is where the bytes live.

The key rule: dynamically provisioned PVs inherit the StorageClass reclaimPolicy, and the default is Delete. With Delete, removing the PVC deletes the PV and the cloud disk.

Common causes

The PVC was deleted. With reclaim policy Delete, the disk is gone too. Typical triggers:

  • Deleting the namespace.
  • helm uninstall of a chart that templates the PVC directly.
  • A GitOps tool pruning a PVC that disappeared from Git after a refactor.
  • A StatefulSet with persistentVolumeClaimRetentionPolicy set to Delete (stable since Kubernetes 1.32) being deleted or scaled down.

The pod got a new, empty volume. The old PVC still exists, but the pod uses a different one. Renaming a StatefulSet or its volumeClaimTemplates entry changes PVC names (<template>-<statefulset>-<ordinal>), so new empty PVCs are created. Also happens when a PVC was recreated by hand and bound to a fresh PV.

The data was never on the volume. The volume is mounted at /data, but the app writes to /var/lib/my-app. Everything goes to the container's writable layer or to an emptyDir, and disappears with the pod. A wrong subPath, or an image whose data directory changed between versions, gives the same result.

Two writers. ReadWriteOnce means one node, not one pod. Two pods on the same node can mount the same RWO volume and both write to it. For apps that are not built for shared storage, that corrupts data. ReadWriteOncePod restricts the volume to a single pod.

Full disk or unclean shutdown. The app crashes or the filesystem needs a repair after a node failure. Data is there but unreadable until fixed.

Zone mismatch. Not data loss, but it looks like it: with volumeBindingMode: Immediate, a disk can be created in a zone where the pod cannot run. The pod stays Pending with a volume node affinity conflict. WaitForFirstConsumer creates the disk after the pod is scheduled.

Debugging steps

1. Stop automation. Pause the GitOps sync and CI deploys for the app so nothing deletes or recreates more objects while you look.

2. Check claims and volumes.

kubectl get pvc -n my-app
kubectl get pv -o custom-columns=NAME:.metadata.name,CLAIM:.spec.claimRef.name,NS:.spec.claimRef.namespace,STATUS:.status.phase,POLICY:.spec.persistentVolumeReclaimPolicy,CREATED:.metadata.creationTimestamp
kubectl describe pvc data-my-db-0 -n my-app

A PVC created minutes ago next to a Released PV with the same claim name means the pod is on a new volume and the old one still exists.

3. Find the real disk.

kubectl get pv <pv-name> -o jsonpath='{.spec.csi.driver}{" "}{.spec.csi.volumeHandle}{"\n"}'
kubectl get volumeattachments | grep <pv-name>

The volume handle is the cloud disk ID. Look it up in the cloud console or CLI, including its snapshots.

4. Check where the app writes.

kubectl exec -n my-app my-db-0 -- sh -c 'df -h /var/lib/my-app; ls -la /var/lib/my-app'

If df shows the overlay filesystem instead of the volume device, the data path is not on the volume.

5. Read events and logs. kubectl get events -n my-app --sort-by=.lastTimestamp shows attach, mount and provisioning errors. The CSI controller and node plugin logs show what the driver did with the disk.

Recovering a Released volume

A Released PV with policy Retain still has its data. To bind it to a new PVC, clear the old claim's UID so the PV can be claimed again, then create a PVC that names it:

kubectl patch pv <pv-name> --type json -p '[{"op":"remove","path":"/spec/claimRef/uid"}]'
apiVersion: v1
kind: PersistentVolumeClaim
metadata:
  name: data-my-db-0
  namespace: my-app
spec:
  accessModes: ["ReadWriteOnce"]
  storageClassName: db-retain
  volumeName: <pv-name>
  resources:
    requests:
      storage: 100Gi

Access mode and storage class must match the PV, and the request cannot be larger than its capacity. For a StatefulSet, use the exact PVC name the pod expects. Take a snapshot of the disk before touching anything.

Prevention

Use a separate StorageClass for data you cannot lose:

apiVersion: storage.k8s.io/v1
kind: StorageClass
metadata:
  name: db-retain
provisioner: ebs.csi.aws.com
parameters:
  type: gp3
  encrypted: "true"
reclaimPolicy: Retain
volumeBindingMode: WaitForFirstConsumer
allowVolumeExpansion: true
  • Switch existing PVs to Retain: kubectl patch pv <pv-name> -p '{"spec":{"persistentVolumeReclaimPolicy":"Retain"}}'.
  • Add helm.sh/resource-policy: keep to PVCs templated by Helm, and disable pruning for PVCs in your GitOps tool.
  • Use ReadWriteOncePod for single-writer databases.
  • Take regular snapshots (VolumeSnapshot API or the cloud's own) and application-level backups. Test restores on a schedule.
  • Alert on disk usage with kubelet_volume_stats_used_bytes and kubelet_volume_stats_capacity_bytes, and on I/O latency from the cloud's disk metrics. On AWS, gp3 has a baseline of 3,000 IOPS and 125 MiB/s that you can raise per volume; size it from measured load, not defaults.
  • In code review, treat any change to a StatefulSet name, volumeClaimTemplates or PVC manifest as a data migration.

Checklist

  • Retain policy and WaitForFirstConsumer on storage classes for important data.
  • PVCs protected from Helm uninstall and GitOps pruning.
  • Data path verified to be on the volume.
  • Single-writer apps use ReadWriteOncePod.
  • Snapshots plus tested restores.
  • On incident: pause automation, list PVs with claims and policies, find the disk ID, snapshot before fixing.