Kubernetes Upgrades in Regulated Environments: Process and Evidence
Kubernetes ships a new minor version roughly three times a year, and the project maintains only the three most recent minor releases. Since 1.19, each minor gets about one year of patch support. Managed services have their own support windows, and extended support usually costs extra. In practice, a cluster that is not upgraded several times a year falls out of support.
In a regulated environment (PCI DSS, SOC 2, ISO 27001, HIPAA and similar), every upgrade is also a change that needs approval, testing evidence and a rollback plan. The technical part is rarely the slow part. The slow part is producing the paperwork by hand each time. The goal of this post is a process where the pipeline produces most of that evidence on its own.
Constraints you cannot negotiate
From the Kubernetes version skew policy:
- The API server must not skip minor versions. Going from 1.33 to 1.35 means 1.33 to 1.34, then 1.34 to 1.35.
- In HA control planes, API server instances may differ by at most one minor version during the upgrade.
- The kubelet must not be newer than the API server and may be up to three minor versions older (two for kubelets older than 1.25).
- kubectl is supported within one minor version of the API server.
- Upgrade order: control plane first, then nodes.
Most managed services do not let you downgrade the control plane to a previous minor version. Your rollback plan for a control plane upgrade is therefore "fix forward" or "move workloads to a second cluster", not "click downgrade". Write that down before the change board asks.
Falling behind is also a compliance problem. Most frameworks expect timely patching of known vulnerabilities, and an end-of-life Kubernetes version no longer receives security fixes.
Make the upgrade a standard change
If every upgrade goes to a change advisory board as a new, unique change, you will never keep up. Work with your change management or GRC team to define Kubernetes minor upgrades as a standard (pre-approved) change. That usually requires:
- A documented runbook that does not change between upgrades.
- Defined test and verification steps with recorded results.
- Defined risk level and rollback approach.
- The same path every time: dev, then staging, then production.
Once approved, each upgrade needs only a change record that references the runbook and attaches the evidence.
Pre-upgrade checks
Removed APIs. The most common upgrade failure is a manifest or Helm chart using an API version that the target release removed. Check what clients still call:
kubectl get --raw /metrics | grep apiserver_requested_deprecated_apis
The metric lists group, version, resource and the release in which it will be removed. Scan your manifests and Helm releases too, with pluto or kube-no-trouble (kubent). Managed services also provide upgrade insights (EKS, GKE and AKS each have their own version).
Release notes. Read the "Urgent Upgrade Notes" and the deprecations in the Kubernetes changelog for every minor version you pass through.
Add-on compatibility. CNI, CSI drivers, ingress controller, cert-manager, service mesh, operators and policy engines each publish supported Kubernetes versions. Upgrade add-ons that need it before the control plane.
Disruption budgets. A PodDisruptionBudget that allows zero disruptions blocks node drains. Find them before the maintenance window:
kubectl get pdb -A -o custom-columns=NS:.metadata.namespace,NAME:.metadata.name,ALLOWED:.status.disruptionsAllowed
Upgrade through code
Keep the version in Terraform (or your IaC tool) so the change is a pull request with a reviewable plan:
resource "aws_eks_cluster" "main" {
name = "prod-eu"
version = var.kubernetes_version
role_arn = aws_iam_role.cluster.arn
vpc_config {
subnet_ids = var.private_subnet_ids
}
}
The pull request becomes most of your change record: what changed, who reviewed it, who approved it, and the plan output. Pin node images and add-on versions the same way, so a node pool rollout is also a reviewed change.
Then:
- Upgrade dev, run the test suite, keep it running for an agreed soak period.
- Upgrade staging with the same pull request flow.
- Upgrade the production control plane.
- Roll node pools one at a time with surge capacity, so pods are rescheduled before old nodes go away.
Evidence the pipeline can collect
Store these as CI artifacts or attach them to the change record automatically:
- The approved pull request and plan output.
- The deprecated API scan, showing nothing in use that the target removes.
- Versions before and after:
kubectl version -o json
kubectl get nodes -o custom-columns=NAME:.metadata.name,KUBELET:.status.nodeInfo.kubeletVersion
(kubectl version --short no longer exists; the flag was removed in kubectl 1.28.)
- Test results from the post-upgrade smoke tests.
- Monitoring screenshots or queries covering the change window: error rates, latency, pod restarts.
If auditors ask "how do you know the upgrade worked and nothing else changed", this set answers it without anyone writing a report.
Plan the calendar
- Track end-of-support dates for your provider's versions and schedule upgrades well ahead of them.
- Avoid change freezes: many regulated businesses freeze changes during peak periods or before audits, so the upgrade windows are fewer than the calendar suggests.
- Do not batch several minor versions into one project. Each hop is a separate change with its own testing.
Checklist
- Minor upgrades are an approved standard change with a fixed runbook.
- No skipped minor versions; control plane first, then nodes.
- Deprecated API scan and add-on compatibility check before each hop.
- Version and node images defined in IaC, upgrades go through pull requests.
- Dev, staging, production, with a soak period between.
- Rollback plan written as fix-forward or failover, since downgrades are not supported.
- Evidence collected automatically by the pipeline and attached to the change record.
