Cloud

What is infrastructure drift, and why your Terraform state keeps lying to you

Infrastructure drift is the gap between what your Terraform state says and what is actually running. Why it happens, what breaks, and how to catch it before an incident.

DevOpsArk EngineeringEngineering team, DevOpsArkPublished 16 September 20267 min read
TL;DR

Your Terraform state describes the infrastructure you think you have. Drift is the gap between that and what is actually running, and it opens whenever a resource is changed outside the tool meant to manage it: a console edit, an emergency fix, another automation. Drift is close to inevitable, so the goal is detecting it quickly with scheduled plans, logging out-of-pipeline changes and codifying manual fixes.

Short answer

What is infrastructure drift?

Infrastructure drift is the difference between the infrastructure declared in code (a Terraform state file, a CloudFormation stack or another infrastructure-as-code source of truth) and the infrastructure actually running. It happens when a resource is changed outside the managing tool, for example by a manual console edit, an emergency fix during an incident or a separate automation touching the same resource.

Why infrastructure drift happens

Drift happens whenever a change bypasses the infrastructure-as-code tool that is supposed to be the source of truth. The common causes:

  • Manual console changes. Someone opens the AWS, Azure or GCP console, often during an incident and under time pressure, and makes a change that never finds its way back into the Terraform configuration.
  • Emergency fixes. The fastest way through an active incident is rarely "write a PR, get it reviewed, run the pipeline". Engineers reach for the console or a direct CLI command, fully intending to backport the change later, and often do not.
  • Other automation touching the same resources. A separate script, another team's pipeline or a third-party tool modifies a resource Terraform also manages, and Terraform never knows.
  • Provider-side changes. Occasionally a cloud provider changes a default or a resource attribute without any action from the team.

None of these are exotic edge cases. They are ordinary parts of operating infrastructure day to day, which is why drift is closer to inevitable than rare.

Why Terraform state "lies" once drift happens

The Terraform state file is a snapshot of what Terraform believes about your infrastructure, based on the last time it applied a change. It does not continuously verify reality. It trusts its own record until something forces a re-check.

Once a resource has drifted, the state describes infrastructure that no longer exists in that form, and Terraform cannot know until someone runs a plan or apply and it queries the provider. Until then, every plan reasons about an assumption rather than a verified fact. That is how a drifted resource sits unnoticed for weeks and then produces a baffling result the next time someone touches that part of the configuration.

# Read-only drift check: refresh against the provider, change nothing.
# Exit code 2 means the live infrastructure differs from the configuration.
terraform plan -refresh-only -detailed-exitcode -input=false

What breaks when drift goes undetected

  • An apply can revert a manual fix. If an emergency change was never backported into code, the next routine apply silently undoes it and reintroduces the original problem.
  • Incident investigations take longer. An engineer reads the Terraform configuration, sees what should be running, and loses time before realising the real infrastructure has diverged.
  • Security posture quietly degrades. A security group rule loosened temporarily during an incident and never reverted or codified can persist far longer than intended. It is exactly the change an audit trail should record and a drift check should flag.
  • Trust in the source of truth erodes. After being burned a few times, engineers start doubting whether the code reflects reality at all, which undermines the point of infrastructure as code.

How teams detect and manage drift

  • Scheduled drift detection. Run a read-only plan against live infrastructure on a schedule, even when nothing is being applied, so drift surfaces before it causes an incident.
  • Treat manual changes as alert-worthy events, not silent ones. A change made outside the pipeline is exactly what an audit trail should capture. Undocumented changes are hard to trace precisely because they do not leave the record they should.
  • Backport emergency fixes as a standard incident follow-up. Make "codify the manual fix" a required postmortem action item rather than optional cleanup.
  • Restrict console access to production resources managed by infrastructure as code, so manual changes become a deliberate exception instead of the easy default under pressure.
Two views of one problem. A drifted resource and an unlogged manual change are usually the same underlying event seen from two angles. Correlating drift reports with the audit trail tells you not just what differs, but who changed it and when.

Common mistakes

  • Assuming Terraform will notice drift on its own without a scheduled detection step.
  • Never backporting emergency manual fixes, so drift accumulates invisibly.
  • Treating drift detection as a one-time setup task rather than an ongoing, scheduled practice.
  • Not correlating drift with the audit trail, and so investigating the same problem twice.

How DevOpsArk approaches this

DevOpsArk manages clusters, servers and cloud accounts centrally across environments, and reports drift from a declared baseline as a specific difference rather than a vague warning. The audit trail records every DevOps event, including changes made outside the normal pipeline, with full history retained, so a drift report comes with a record of exactly when and where the manual change was made. 360 DITE adds automated infrastructure assessment and reliability analysis, which turns drift detection into an ongoing capability rather than a manual step that is easy to skip.

Key takeaways

  • Drift is the gap between declared infrastructure and what is actually running.
  • It comes from ordinary operations: console edits, emergency fixes and overlapping automation.
  • Terraform state is a record of the last apply, not a live view of reality.
  • Undetected drift can let a routine apply revert a manual fix and restart an incident.
  • Schedule read-only plans, log out-of-pipeline changes and make codifying manual fixes a postmortem action.

Frequently asked questions

Infrastructure as codeTerraformDriftCloud

See this working on your own infrastructure

A 30-minute walkthrough with a platform engineer. Bring the problem this article describes.