Argo CD vs Flux: Deployment Lag and Rollback Speed

Argo CD vs Flux: Deployment Lag and Rollback Speed

Reading time1 min
#devops#kubernetes#argocd#flux

Argo CD vs Flux: Deployment Lag and Rollback Speed

Argo CD and Flux both pull desired state from Git and apply it to Kubernetes. When people compare them on speed, the result usually says more about the setup than about the tool. Applying a few manifests takes seconds in either one. The time goes into polling, rollouts and health checks.

This post covers what drives deployment lag and rollback time in each tool, what to tune, and how to measure it in your own cluster instead of trusting someone else's benchmark.

Where the time goes

The time from git push to a healthy new version has four parts:

  1. Change detection. How long until the controller notices the new commit.
  2. Render and apply. Running Helm or Kustomize and sending objects to the API server. Usually seconds.
  3. Rollout. Pulling images, starting pods, passing readiness probes. Usually the biggest part, and the same for both tools.
  4. Health assessment. How the tool decides the change worked.

Parts 1 and 4 are where Argo CD and Flux actually differ.

Change detection

Argo CD polls Git based on timeout.reconciliation in the argocd-cm ConfigMap. The default is 120 seconds plus up to 60 seconds of jitter (timeout.reconciliation.jitter), so a commit can wait up to three minutes before anything happens. A webhook from your Git provider to https://argocd.example.com/api/webhook removes most of that wait.

Flux has no global default. Every GitRepository has a required spec.interval, and source-controller fetches at that interval. When a new revision appears, the Kustomizations and HelmReleases that use it are reconciled right away. To skip polling, add a Receiver:

apiVersion: notification.toolkit.fluxcd.io/v1
kind: Receiver
metadata:
  name: github-receiver
  namespace: flux-system
spec:
  type: github
  events: ["push"]
  secretRef:
    name: receiver-token
  resources:
    - apiVersion: source.toolkit.fluxcd.io/v1
      kind: GitRepository
      name: my-app

The generated path is in .status.webhookPath. Expose notification-controller's webhook endpoint through your ingress and register that URL in GitHub with the same token.

With webhooks on both sides, detection takes seconds in either tool. Keep polling as a fallback, because webhook deliveries do get lost.

Health assessment

Argo CD marks an Application Healthy using built-in health checks per resource kind (a Deployment counts as healthy once its rollout is complete) and Lua health checks for custom resources.

Flux Kustomizations only check health when you ask for it: spec.wait: true for everything, or a spec.healthChecks list, bounded by spec.timeout. Without that, Flux reports Ready as soon as the apply succeeds. That looks faster but says nothing about the pods. HelmReleases wait for resources by default, with a default spec.timeout of 5 minutes.

When you compare the two tools, use equivalent health settings, or you are measuring different things.

Measure it yourself

Push a harmless change (an annotation bump is enough) and poll until the new revision is applied and healthy:

#!/usr/bin/env bash
set -euo pipefail
git commit --allow-empty -m "lag test" -q   # or bump an annotation in the manifests
sha=$(git rev-parse HEAD)
start=$(date +%s)
git push -q origin main

app() { kubectl -n argocd get applications.argoproj.io my-app -o jsonpath="$1"; }
until [[ "$(app '{.status.sync.revision}')" == "$sha" &&
         "$(app '{.status.health.status}')" == "Healthy" ]]; do
  sleep 2
done
echo "commit to healthy: $(( $(date +%s) - start ))s"

An empty commit only measures detection and sync. To include the rollout, change something in the pod template. For Flux, poll kubectl -n flux-system get kustomization my-app -o jsonpath='{.status.lastAppliedRevision}' until it contains the SHA and the Ready condition is True.

Run it many times, with and without webhooks, at quiet and busy hours. Keep the worst case, not only the median. During an incident the worst case is what you get.

Rollback

With GitOps, the normal rollback is git revert plus a forced reconcile. Anything you change directly in the cluster gets reverted by self-heal (Argo CD) or drift correction (Flux) on the next pass.

git revert --no-edit HEAD && git push
argocd app sync my-app                              # Argo CD
flux reconcile kustomization my-app --with-source   # Flux

Both tools also have an emergency path that bypasses Git:

  • Argo CD: argocd app history my-app, then argocd app rollback my-app <ID>. Rollback is refused while automated sync is enabled, so you have to disable it first. The app then stays OutOfSync until Git is reverted too. History length is revisionHistoryLimit, 10 by default.
  • Flux: flux suspend kustomization my-app, fix the cluster by hand, revert in Git, then flux resume kustomization my-app.

For Helm there is a real difference. Flux runs actual Helm installs and upgrades, so a HelmRelease can roll itself back when an upgrade fails:

apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
metadata:
  name: my-app
  namespace: my-app
spec:
  interval: 10m
  chart:
    spec:
      chart: my-app
      sourceRef:
        kind: HelmRepository
        name: my-charts
        namespace: flux-system
  upgrade:
    remediation:
      retries: 2
      strategy: rollback

Argo CD renders charts with helm template and applies the manifests itself. There is no Helm release history in the cluster, so helm rollback does not apply.

Automatic rollback on metrics

Neither core tool does canary analysis. Two separate controllers do:

  • Argo Rollouts, a separate Argo project. It replaces the Deployment with a Rollout resource that supports canary and blue-green steps with metric analysis.
  • Flagger, part of the Flux project. It works with regular Deployments and drives traffic shifting through a service mesh or ingress.

Both work with either GitOps tool. Both abort and shift traffic back when analysis fails, which is faster than any human doing git revert.

Differences that matter more than speed

  • Argo CD ships a web UI, SSO, project-level RBAC, ApplicationSets and a hub model that manages many clusters from one instance.
  • Flux is driven by CRDs and the CLI, usually runs in every cluster, performs native Helm releases and decrypts SOPS secrets in kustomize-controller.

Checklist

  • Webhooks configured, polling kept as a fallback.
  • Health checks enabled (wait: true in Flux) so "synced" means "running".
  • Commit-to-healthy measured with a script, many runs, worst case recorded.
  • Rollback is git revert plus forced reconcile, and the team has practiced it.
  • Everyone knows how to pause the controller: disable auto-sync in Argo CD, flux suspend in Flux.
  • Metric-based automatic rollback with Argo Rollouts or Flagger where it is worth the setup.