Terraform State: How It Breaks and How to Recover

Terraform State: How It Breaks and How to Recover

Reading time1 min
#terraform#devops#cloud#infrastructure-as-code

Terraform State: How It Breaks and How to Recover

Terraform state is the mapping between your code and the real resources in the cloud. If it is wrong, terraform plan is wrong, and the next apply will happily "fix" infrastructure that was fine. Almost every scary Terraform incident comes down to state that no longer matches reality.

This post covers the common ways state goes bad, what to set up so it doesn't, and what to do when it already has.

How state gets damaged

Hand-edited state. Someone opens the JSON to fix a resource address or delete an entry. One typo or a dropped resource and the next plan wants to recreate a database. There is almost never a reason to edit state by hand today: moved, import and removed blocks cover renames, adoption and forgetting resources.

Two applies at once. Without locking, two people or two CI jobs can read the same state, change different things and write back. The last write wins and the other change disappears from state while the resource still exists in the cloud.

An apply that dies halfway. CI runner gets killed, laptop goes to sleep, credentials expire. Resources were created, but the new state was never written. If Terraform could not save state to the backend, it writes errored.tfstate in the working directory. On an ephemeral CI runner that file is gone with the runner unless you collect it.

Refactoring without moved. Renaming a resource or moving it into a module changes its address. Terraform sees "old address gone, new address added" and plans destroy plus create. For a bucket or a database that is data loss.

Provider version skew. A newer provider can upgrade the resource schema stored in state. After that, a job pinned to an older provider fails with "Resource instance managed by newer provider version". Not dangerous by itself, but it pushes people toward quick fixes like editing state.

Wrong workspace or backend key. Applying the staging config against the production state key, or the other way around. Plan output gets long, nobody reads it, apply goes through.

Guardrails

Remote backend with locking and versioning.

terraform {
  backend "s3" {
    bucket       = "my-terraform-state"
    key          = "prod/network.tfstate"
    region       = "eu-central-1"
    use_lockfile = true
    encrypt      = true
  }
}

Since Terraform 1.10 the S3 backend can lock with a lock file in the bucket itself (use_lockfile), and the DynamoDB lock table is deprecated. The GCS backend locks out of the box. In both cases turn on object versioning on the bucket: that is your undo button.

Pin Terraform and providers. Set required_version and required_providers with version constraints, and commit .terraform.lock.hcl. Everyone, including CI, then runs the same provider builds.

Apply exactly what was reviewed.

terraform plan -out=tfplan
terraform show tfplan        # what reviewers look at
terraform apply tfplan       # applies that plan, nothing else

If state changed between plan and apply, Terraform refuses to apply a stale plan file.

Protect what must not be destroyed. Use lifecycle { prevent_destroy = true } on databases and buckets with data. It only works while the resource block exists, so also enable deletion protection in the provider (deletion_protection on RDS and Cloud SQL, for example).

Refactor in code, not in state.

moved {
  from = aws_s3_bucket.logs
  to   = module.logging.aws_s3_bucket.this
}

moved (Terraform 1.1+), import blocks (1.5+) and removed blocks (1.7+) go through code review and show up in plan, unlike terraform state mv run from someone's laptop.

Limit who can write state. Only the CI role should be able to write the state object. People who run plan locally need read access to state and write access only to the lock file.

Recovery checklist

  1. Stop all applies. Pause the pipeline, tell the team. Do not run force-unlock until you are sure no apply is still running.
  2. Save what you have.
    terraform state pull > state-$(date +%s).json
    
  3. Find a good version. With bucket versioning:
    # S3
    aws s3api list-object-versions --bucket my-terraform-state --prefix prod/network.tfstate
    aws s3api get-object --bucket my-terraform-state --key prod/network.tfstate \
      --version-id <VERSION_ID> restored.tfstate
    
    # GCS
    gcloud storage ls --all-versions gs://my-terraform-state/prod/default.tfstate
    gcloud storage cp "gs://my-terraform-state/prod/default.tfstate#<GENERATION>" restored.tfstate
    
  4. Push it back carefully. terraform state push restored.tfstate refuses a snapshot with a lower serial or a different lineage. That check is there for a reason. Use -force only when you know exactly why it is needed.
  5. If an apply crashed, look for errored.tfstate and push that instead of an older version. It contains the resources that were actually created.
  6. Reconcile. Run terraform plan. Resources that exist but are missing from state go back in with import blocks. Resources in state that no longer exist will show up as "create": check each one before you apply.
  7. Plan until it is clean. You are done when plan shows no changes, or only the changes you expect.

Short version

  • Remote backend, locking, bucket versioning.
  • Pinned versions and a committed lock file.
  • Apply only saved plans, from CI.
  • prevent_destroy plus provider-level deletion protection on stateful resources.
  • moved, import and removed blocks instead of state surgery.
  • Before any recovery, stop applies and take a copy of the current state.