Blue-Green Deployments on Kubernetes: Where the Traffic Switch Fails

Blue-Green Deployments on Kubernetes: Where the Traffic Switch Fails

Reading time1 min
#devops#cloud#kubernetes#deployment#blue-green

Blue-Green Deployments on Kubernetes: Where the Traffic Switch Fails

Blue-green means running two complete versions side by side. Blue serves production traffic, green gets the new release. When green is verified, all traffic moves to green at once, and blue stays around as an instant rollback target.

It is different from a canary, which shifts traffic gradually by percentage. Most blue-green outages come from a broken assumption: that green is fully ready, that the switch is atomic, that blue is still there to go back to, or that both versions can share the same data.

The basic switch

On plain Kubernetes you need two Deployments with a version label and one Service. The switch is a selector change:

apiVersion: v1
kind: Service
metadata:
  name: my-app
spec:
  selector:
    app: my-app
    version: blue
  ports:
    - port: 80
      targetPort: 8080
kubectl -n my-app patch service my-app \
  -p '{"spec":{"selector":{"app":"my-app","version":"green"}}}'

An Ingress with two backends on the same path does not split traffic. For percentages you need canary support in your ingress controller, a service mesh, or Gateway API weights.

Where it goes wrong

Green is not at full capacity. Green was started with fewer replicas, or an HPA starts it at minReplicas. Caches are cold, connection pools are empty, JIT-compiled runtimes are slow on first requests. Scale green to blue's current replica count before switching, and send it synthetic traffic through a preview Service first.

Readiness probes pass too early. If the probe returns 200 before the app can actually serve requests, the switch moves traffic to pods that are not ready.

Existing connections stay on blue. A selector change affects new connections only. Keep-alive HTTP, gRPC and WebSocket connections stay on blue pods until they close. If blue is scaled down right after the switch, those clients get errors. Keep blue running for a while, handle SIGTERM by draining connections, and add a short preStop delay so endpoint removal reaches all proxies before the process stops. The sleep action for preStop is stable since Kubernetes 1.34 (beta and on by default since 1.30), so you do not need a sleep binary in the image.

Shared state is not compatible. Both colors talk to the same database, queues and caches. If green runs a migration that blue cannot handle, rollback is impossible. Use expand and contract migrations so the old version keeps working against the new schema.

Background work runs in both colors. Queue consumers and scheduled jobs inside both Deployments will both process work. Old code may receive messages in the new format. Either keep workers out of the blue-green flow or make both versions handle both formats.

Blue is gone too early. The point of blue-green is a fast rollback. Once blue is scaled down, a rollback means starting pods again, which is a normal deploy, not a switch.

DNS-based switching. Flipping DNS between two load balancers depends on client caching, and some clients ignore TTLs. Switch at the load balancer or Service level instead.

Monitoring per version

If dashboards aggregate both colors, you cannot tell which one is failing. Put the version label on the pods and copy it to metrics. With the Prometheus Operator, podTargetLabels: [version] on the ServiceMonitor does that. Then compare error rates side by side:

sum by (version) (rate(http_requests_total{job="my-app", code=~"5.."}[5m]))
/
sum by (version) (rate(http_requests_total{job="my-app"}[5m]))

Watch latency percentiles, restarts and a business metric (logins, orders) the same way. Decide before the switch which numbers trigger a rollback.

With Argo Rollouts

Argo Rollouts automates the Service switch and keeps the old ReplicaSet for a configurable time:

apiVersion: argoproj.io/v1alpha1
kind: Rollout
metadata:
  name: my-app
spec:
  replicas: 4
  selector:
    matchLabels:
      app: my-app
  template:
    metadata:
      labels:
        app: my-app
    spec:
      containers:
        - name: my-app
          image: registry.example.com/my-app:1.2.3
          ports:
            - containerPort: 8080
  strategy:
    blueGreen:
      activeService: my-app
      previewService: my-app-preview
      autoPromotionEnabled: false
      scaleDownDelaySeconds: 600
      prePromotionAnalysis:
        templates:
          - templateName: smoke-test

The controller adds a pod template hash to both Service selectors, so my-app points at the stable version and my-app-preview at the new one. The smoke-test AnalysisTemplate (defined separately) has to pass before promotion. With autoPromotionEnabled: false, someone promotes explicitly with kubectl argo rollouts promote my-app. The default scaleDownDelaySeconds is 30. Raise it to cover the time you need to spot a problem, because rolling back is only instant while the old ReplicaSet is still running.

Weighted shifting with Gateway API

If you want to move a small share of traffic first, that is a canary step. Gateway API supports it natively:

apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
  name: my-app
  namespace: my-app
spec:
  parentRefs:
    - name: my-gateway
  hostnames: ["app.example.com"]
  rules:
    - backendRefs:
        - name: my-app-blue
          port: 80
          weight: 90
        - name: my-app-green
          port: 80
          weight: 10

Note that ingress-nginx, which offered canary annotations, was retired by the Kubernetes project in March 2026. For new setups, use Gateway API or another maintained controller.

Checklist

  • Green runs at full production capacity before the switch.
  • Readiness probes reflect real readiness.
  • Graceful shutdown and a preStop delay protect in-flight requests.
  • Blue stays up long enough to roll back by switching, not redeploying.
  • Database changes are backward compatible with the old version.
  • Workers and scheduled jobs are handled separately.
  • Metrics carry a version label, and rollback thresholds are agreed beforehand.
  • The switch happens at the Service or load balancer, not in DNS.