Argo CD vs Flux: Deployment Lag and Rollback Speed
Argo CD and Flux both pull desired state from Git and apply it to Kubernetes. When people compare them on speed, the result usually says more about the setup than about the tool. Applying a few manifests takes seconds in either one. The time goes into polling, rollouts and health checks.
This post covers what drives deployment lag and rollback time in each tool, what to tune, and how to measure it in your own cluster instead of trusting someone else's benchmark.
Where the time goes
The time from git push to a healthy new version has four parts:
- Change detection. How long until the controller notices the new commit.
- Render and apply. Running Helm or Kustomize and sending objects to the API server. Usually seconds.
- Rollout. Pulling images, starting pods, passing readiness probes. Usually the biggest part, and the same for both tools.
- Health assessment. How the tool decides the change worked.
Parts 1 and 4 are where Argo CD and Flux actually differ.
Change detection
Argo CD polls Git based on timeout.reconciliation in the argocd-cm ConfigMap. The default is 120 seconds plus up to 60 seconds of jitter (timeout.reconciliation.jitter), so a commit can wait up to three minutes before anything happens. A webhook from your Git provider to https://argocd.example.com/api/webhook removes most of that wait.
Flux has no global default. Every GitRepository has a required spec.interval, and source-controller fetches at that interval. When a new revision appears, the Kustomizations and HelmReleases that use it are reconciled right away. To skip polling, add a Receiver:
apiVersion: notification.toolkit.fluxcd.io/v1
kind: Receiver
metadata:
name: github-receiver
namespace: flux-system
spec:
type: github
events: ["push"]
secretRef:
name: receiver-token
resources:
- apiVersion: source.toolkit.fluxcd.io/v1
kind: GitRepository
name: my-app
The generated path is in .status.webhookPath. Expose notification-controller's webhook endpoint through your ingress and register that URL in GitHub with the same token.
With webhooks on both sides, detection takes seconds in either tool. Keep polling as a fallback, because webhook deliveries do get lost.
Health assessment
Argo CD marks an Application Healthy using built-in health checks per resource kind (a Deployment counts as healthy once its rollout is complete) and Lua health checks for custom resources.
Flux Kustomizations only check health when you ask for it: spec.wait: true for everything, or a spec.healthChecks list, bounded by spec.timeout. Without that, Flux reports Ready as soon as the apply succeeds. That looks faster but says nothing about the pods. HelmReleases wait for resources by default, with a default spec.timeout of 5 minutes.
When you compare the two tools, use equivalent health settings, or you are measuring different things.
Measure it yourself
Push a harmless change (an annotation bump is enough) and poll until the new revision is applied and healthy:
#!/usr/bin/env bash
set -euo pipefail
git commit --allow-empty -m "lag test" -q # or bump an annotation in the manifests
sha=$(git rev-parse HEAD)
start=$(date +%s)
git push -q origin main
app() { kubectl -n argocd get applications.argoproj.io my-app -o jsonpath="$1"; }
until [[ "$(app '{.status.sync.revision}')" == "$sha" &&
"$(app '{.status.health.status}')" == "Healthy" ]]; do
sleep 2
done
echo "commit to healthy: $(( $(date +%s) - start ))s"
An empty commit only measures detection and sync. To include the rollout, change something in the pod template. For Flux, poll kubectl -n flux-system get kustomization my-app -o jsonpath='{.status.lastAppliedRevision}' until it contains the SHA and the Ready condition is True.
Run it many times, with and without webhooks, at quiet and busy hours. Keep the worst case, not only the median. During an incident the worst case is what you get.
Rollback
With GitOps, the normal rollback is git revert plus a forced reconcile. Anything you change directly in the cluster gets reverted by self-heal (Argo CD) or drift correction (Flux) on the next pass.
git revert --no-edit HEAD && git push
argocd app sync my-app # Argo CD
flux reconcile kustomization my-app --with-source # Flux
Both tools also have an emergency path that bypasses Git:
- Argo CD:
argocd app history my-app, thenargocd app rollback my-app <ID>. Rollback is refused while automated sync is enabled, so you have to disable it first. The app then staysOutOfSyncuntil Git is reverted too. History length isrevisionHistoryLimit, 10 by default. - Flux:
flux suspend kustomization my-app, fix the cluster by hand, revert in Git, thenflux resume kustomization my-app.
For Helm there is a real difference. Flux runs actual Helm installs and upgrades, so a HelmRelease can roll itself back when an upgrade fails:
apiVersion: helm.toolkit.fluxcd.io/v2
kind: HelmRelease
metadata:
name: my-app
namespace: my-app
spec:
interval: 10m
chart:
spec:
chart: my-app
sourceRef:
kind: HelmRepository
name: my-charts
namespace: flux-system
upgrade:
remediation:
retries: 2
strategy: rollback
Argo CD renders charts with helm template and applies the manifests itself. There is no Helm release history in the cluster, so helm rollback does not apply.
Automatic rollback on metrics
Neither core tool does canary analysis. Two separate controllers do:
- Argo Rollouts, a separate Argo project. It replaces the Deployment with a
Rolloutresource that supports canary and blue-green steps with metric analysis. - Flagger, part of the Flux project. It works with regular Deployments and drives traffic shifting through a service mesh or ingress.
Both work with either GitOps tool. Both abort and shift traffic back when analysis fails, which is faster than any human doing git revert.
Differences that matter more than speed
- Argo CD ships a web UI, SSO, project-level RBAC, ApplicationSets and a hub model that manages many clusters from one instance.
- Flux is driven by CRDs and the CLI, usually runs in every cluster, performs native Helm releases and decrypts SOPS secrets in kustomize-controller.
Checklist
- Webhooks configured, polling kept as a fallback.
- Health checks enabled (
wait: truein Flux) so "synced" means "running". - Commit-to-healthy measured with a script, many runs, worst case recorded.
- Rollback is
git revertplus forced reconcile, and the team has practiced it. - Everyone knows how to pause the controller: disable auto-sync in Argo CD,
flux suspendin Flux. - Metric-based automatic rollback with Argo Rollouts or Flagger where it is worth the setup.
