Feature Flag Drift: How Stale Toggles Cause Bugs and How to Control Them

Feature Flag Drift: How Stale Toggles Cause Bugs and How to Control Them

Reading time1 min
#devops#feature-flags#engineering#software-development

Feature Flag Drift: How Stale Toggles Cause Bugs and How to Control Them

Feature flags separate deploying code from releasing a feature. That is useful, and it has a cost: every flag doubles the number of possible code paths in the area it touches. Ten independent flags give 1024 combinations, and only a few of them are ever tested.

"Toggle drift" is what happens when the flag state you think you have is not the one production actually runs. This post covers how that happens and the habits and checks that prevent it.

Not all flags are the same

Pete Hodgson's article on martinfowler.com splits flags into four categories: release toggles, experiment toggles, ops toggles and permission toggles. They have different lifetimes. Unleash uses similar types with expected lifetimes: 40 days for release and experiment flags, 7 days for operational flags, and no expiry for kill switches and permission flags. The exact numbers matter less than the idea that a release flag is temporary by definition.

How drift happens

Environment drift. A flag is on in staging and off in production, or the targeting rules differ. Tests pass against a configuration production does not have.

Default drift. Every SDK call has a fallback value in code. When the flag service is unreachable, at startup or during an outage, the application quietly uses that fallback. If the fallback is the old behavior and the flag has been on for months, a network problem turns into a functional rollback.

Cross-service drift. Two services check the same flag, but one evaluates it per user and the other per account, or one caches the value for longer. A single request takes the new path in one service and the old path in the other.

Stale flags. A flag has been at 100% for months. Nobody tests the "off" path anymore, but the code is still there. Someone switches it off during cleanup, or by mistake, and old code runs against data written by the new code.

Reused keys. A flag key gets reused for a new feature and inherits old targeting rules, or old clients still running somewhere interpret it the old way.

Orphans. Flags that exist in the flag service but not in code, and code that references flags already deleted from the service. The second kind silently falls back to the default.

Rules that keep flags manageable

  1. Every flag has an owner, a type and an expiry date. Store this in the flag tool if it supports it, or in a registry file in the repository.
  2. Removing a flag is part of the feature. Create the cleanup task when you create the flag.
  3. Defaults in code are the safe value, and the same in every service. Define the flag key and its default once, in a shared module, instead of repeating string literals.
  4. Test both states of every active flag in CI. Once a flag is permanently on, delete the "off" path.
  5. Never reuse a flag key.
  6. Evaluate once per request where possible, at the edge, and pass the result downstream in the request context, so all services make the same decision.
  7. Keep flag changes auditable. Production flag changes deserve the same visibility as deploys: who changed what and when, ideally shown on your dashboards.

A flag registry with a CI check

If your flag tool does not track ownership and expiry, a small file in the repository does:

# flags.yaml
flags:
  - key: new-checkout-flow
    type: release
    owner: team-payments
    expires: "2026-11-30"
  - key: disable-recommendations
    type: kill-switch
    owner: team-platform
    expires: never

A CI job then fails on expired flags and on flags no code uses:

#!/usr/bin/env bash
set -euo pipefail
today=$(date +%F)
status=0
while IFS=$'\t' read -r key expires; do
  if [[ "$expires" != "never" && "$expires" < "$today" ]]; then
    echo "expired: $key ($expires)"
    status=1
  fi
  if ! git grep -q -F "$key" -- src/; then
    echo "unused: $key is in flags.yaml but not referenced in src/"
    status=1
  fi
done < <(yq '.flags[] | [.key, .expires] | @tsv' flags.yaml)
exit "$status"

ISO dates compare correctly as strings, so no date parsing is needed. The script uses mikefarah's yq v4. An expired flag then blocks merges until someone removes it or extends the date on purpose, with a reason in the pull request.

The opposite check (flag keys in code that are missing from the registry) is easier if all keys are defined as constants in one module, as rule 3 suggests.

Tooling that helps

  • Unleash marks flags as potentially stale once they outlive the expected lifetime of their type, and tracks lifecycle stages based on usage metrics.
  • LaunchDarkly has code references (ld-find-code-refs) that scan your repository and show where each flag is used.
  • OpenFeature, a CNCF project, defines a vendor-neutral SDK API. Code written against it does not depend on a specific flag provider, which makes migrations and local testing easier.
  • Any provider with an audit log and webhooks can post flag changes to your chat and dashboards.

Cleaning up an existing mess

  1. Export all flags with their current state per environment and last change date.
  2. Flags fully on or off in every environment for longer than their expected lifetime are candidates for removal.
  3. For each one, hardcode the current behavior in code, deploy, then delete the flag from the service. In that order, so a missing flag cannot fall back to an old default.
  4. Flags with different states across environments get an owner and an explicit decision.
  5. Add the registry check so the list does not grow back.

Checklist

  • Every flag has an owner, type and expiry.
  • Defaults are safe and defined once.
  • Both states of active flags are tested.
  • Flags are evaluated once per request and passed downstream.
  • Expired and unused flags fail CI.
  • Code paths are removed before flags are deleted from the service.
  • Flag changes in production are audited and visible next to deploys.