Safe Deploys on Any Day: Release Gates, Rollbacks and Freeze Windows

Safe Deploys on Any Day: Release Gates, Rollbacks and Freeze Windows

Reading time1 min
#devops#cicd#software#engineering#deployment

Safe Deploys on Any Day: Release Gates, Rollbacks and Freeze Windows

"No deploys on Friday" is a proxy rule. The day itself does nothing. What changes on Friday afternoon is that fewer people are around, problems get noticed later, and the time to fix them gets longer. A big or irreversible change makes all of that worse.

You can attack those causes directly. Then a Friday deploy is just a deploy, and a freeze is something you use on purpose, not out of habit.

What makes a deploy risky

  • Size. A release with fifty changes is harder to verify and harder to bisect than a release with one.
  • Reversibility. Schema migrations, data backfills, messages already published to a queue and changed public API contracts do not roll back with the code.
  • Detection time. If alerts fire only on hard failures, a broken checkout flow can go unnoticed for hours.
  • Response capacity. Someone who knows the change has to be reachable while it settles.

Keep changes small and reversible

Deploy every merge to main instead of batching. Hide unfinished work behind feature flags, so deploying code and enabling a feature become separate steps with separate rollbacks.

For database changes use expand and contract:

  1. Add the new column or table. Old code ignores it.
  2. Deploy code that writes to both old and new, and reads from old.
  3. Backfill.
  4. Switch reads to the new structure.
  5. In a later release, stop writing the old one and drop it.

At every step the previous version of the application still works, so rolling back the code is always safe.

Let the pipeline judge the rollout

A deploy job that exits 0 after kubectl apply proves nothing. Wait for the rollout and undo it if it fails:

#!/usr/bin/env bash
set -euo pipefail
kubectl -n my-app set image deployment/my-app my-app="registry.example.com/my-app:${GIT_SHA}"
if ! kubectl -n my-app rollout status deployment/my-app --timeout=10m; then
  kubectl -n my-app rollout undo deployment/my-app
  exit 1
fi

A Deployment also has progressDeadlineSeconds (600 by default), after which it reports the rollout as failed. With Helm, --atomic in Helm 3 rolls back a failed upgrade. Helm 4 renamed it to --rollback-on-failure.

All of this depends on readiness probes that mean something. A probe that returns 200 before the app can actually serve requests will pass a broken release.

After the rollout, run a smoke test against real endpoints and watch error rate and latency for a fixed period before calling the deploy done. Argo Rollouts and Flagger automate that step: they shift traffic gradually and roll back when metrics cross a threshold.

Gates and serialization

In GitHub Actions, put the deploy job in an environment. Environments support required reviewers, a wait timer and restrictions on which branches can deploy. A concurrency group keeps two deploys from running at once:

on:
  push:
    branches: [main]

jobs:
  deploy:
    runs-on: ubuntu-latest
    environment: production
    concurrency:
      group: deploy-production
      cancel-in-progress: false
    steps:
      - uses: actions/checkout@v7
      - run: ./scripts/deploy.sh
        env:
          GIT_SHA: ${{ github.sha }}

In GitLab, resource_group serializes deploy jobs, and "prevent outdated deployment jobs" stops an older pipeline from overwriting a newer deploy.

Freeze windows in config, not in people's heads

If you need a freeze (a holiday, a big sales event, the weekend for a small team), define it in the tool so it is enforced and visible.

GitLab has deploy freezes in project settings. During a freeze, jobs get the CI_DEPLOY_FREEZE variable:

deploy-production:
  stage: deploy
  environment: production
  resource_group: production
  script: ./scripts/deploy.sh
  rules:
    - if: $CI_COMMIT_BRANCH != $CI_DEFAULT_BRANCH
      when: never
    - if: $CI_DEPLOY_FREEZE
      when: manual
      allow_failure: true
    - when: on_success

During a freeze the job turns into a manual one. An urgent fix can still go out, but someone has to press the button on purpose. In GitHub Actions you can get the same effect with a wait timer, required reviewers or a custom deployment protection rule.

Be ready to restore

  • Practice rollback. A rollback path that is never exercised tends to fail when it is needed.
  • Write down who is on call after a deploy and what they should check.
  • Alert on symptoms users see (error rate, latency, failed checkouts), not only on crashed pods.
  • Record deploys as annotations on your dashboards, so a spike can be matched to a change in seconds.

Measure instead of guessing

The DORA metrics are a good starting point: deployment frequency, lead time for changes, change failure rate and how long it takes to recover from a failed deployment. If you want to know whether Friday is actually riskier for your team, record a timestamp for every deploy and every incident and compare failure rate and restore time by day of week. Then decide on a freeze using your own data.

Checklist

  • Small changes, deployed one merge at a time.
  • Feature flags separate deploy from release.
  • Migrations use expand and contract, so code rollback is always safe.
  • The pipeline waits for the rollout and rolls back on failure.
  • Readiness probes reflect whether the app can actually serve requests.
  • Deploys are serialized and gated by environment rules.
  • Freezes are defined in CI config, with a manual override.
  • Rollback is practiced, deploys are annotated, failure rate and restore time are tracked.