An intended private ingress rule compared with an out-of-band public ingress rule before remediation.
Drift

What is Infrastructure Drift and How to Fix It

Infrastructure drift is the gap between IaC and live cloud state. Learn how ops0 detects drift, maps impact, and routes safe remediation.

Reviewed for technical accuracy: October 1, 2026

  • Infrastructure drift is a difference between intended configuration and observed state
  • Refresh-only and normal plans answer different questions and do not change production by themselves
  • Review whether code, the live change, or an owned exception should determine the intended state
  • Discovery is needed for unmanaged resources and gaps in provider coverage
  • Deployment gates protect their own workflow; access controls and scans cover other change paths

Infrastructure drift is when your actual cloud environment no longer matches what your Terraform, CloudFormation, or other infrastructure-as-code says it should be. Someone changed a security group in the AWS console. An emergency SSH fix modified a config file on a server. An auto-scaling event created resources your state file doesn't know about. The gap between declared state and real state is drift, and every team running infrastructure at any scale deals with it.

Drift means the declared configuration, Terraform state, and live system need comparison. A normal plan refreshes managed objects and proposes actions to reach the configuration; it does not automatically change production. A recovery runbook or compliance record may be incomplete when it relies on code that no longer reflects the intended environment.

Why Drift Happens

Drift isn't caused by bad engineers. It's caused by real situations.

Emergency fixes can create drift. Production goes down at 2 AM, an operator changes a security group or a server configuration, and the matching code is not updated. The effect depends on what the IaC manages: a file changed inside a VM is not necessarily visible to Terraform unless a separate configuration system tracks it.

Console changes are another source. An operator changes a cloud setting during troubleshooting and intends to update the repository later. Without a review and reconciliation step, the live resource and the approved configuration diverge.

Multi-team environments make it worse. When multiple teams touch the same infrastructure, one team's change can create drift for another team's code. Nobody's at fault. The tooling doesn't handle shared ownership well.

Auto-scaling and controllers add another layer. Changes they own can be expected behavior. Decide which system owns each field and compare against that boundary; not every new pod or scaled instance is an IaC defect.

How to Detect Drift

For Terraform-managed resources, compare a refreshed view of state with your configuration and inspect the normal plan. A refresh-only plan shows out-of-band changes for managed objects without proposing configuration-driven reconciliation. Unmanaged resources and settings the provider does not read require additional discovery. HashiCorp explains normal and refresh-only planning.

Better tools scan your actual cloud APIs and compare the results against your state files. This catches both changes to managed resources and new resources that were created outside your IaC.

ops0's Resource Graph does this continuously. It compares your Terraform state files against live cloud state across AWS, GCP, and Azure, and shows you exactly what changed, when it changed, and what depends on the changed resource. That last part matters because a drifted security group might affect dozens of instances, and knowing the blast radius before you fix anything prevents making things worse.

Worked Example: An Emergency Security Group Change

Illustrative lab scenario. A Terraform-managed security group declares an approved internal source range for administration. During troubleshooting, an operator broadens the live rule. This example describes a review process, not a measured customer incident or an automatically completed ops0 fix.

In the correct workspace, with the correct account and region, inspect the managed-resource difference:

BASH
terraform plan -refresh-only -out=observed.tfplan
terraform show observed.tfplan
terraform plan -out=reconcile.tfplan
terraform show reconcile.tfplan

The refresh-only plan helps inspect the observed change. The normal plan shows what applying the declared code would do next. Do not apply either plan merely to discover the difference. Saved plans may contain sensitive data and belong in the same protected workflow as state.

Compare the rule with the incident record, access requirements, and attached workloads. If the temporary range should be removed, approve the normal reconciliation plan. If the live change is the new approved design, update the code and review a new plan. Applying a refresh-only plan updates state; it does not update the Terraform configuration or make the setting a durable approved exception.

Capture the resource address, observed difference, discovery time, owner, decision, reviewer, and final verification. Rerun the normal plan and verify the relevant live setting after the approved change. A clean plan for managed objects does not prove that the entire account has no unmanaged resources.

How to Fix Drift

You have three options when you find drift.

Option one: update the code to match reality. The manual change was correct and should be kept. You modify your Terraform to reflect what exists. This is the right choice when the drift was an intentional improvement.

Option two: return the live setting to the approved code. Review a saved normal plan and apply it through the normal approval workflow. Inspect replacements and deletions before proceeding, and confirm that the original emergency change is no longer needed.

Option three: accept an explicitly owned exception. Expected controller-managed settings may need a carefully scoped lifecycle rule or another ownership model. Record the reason and review date. Broadly ignoring changes can hide security-relevant drift.

ops0 connects the drift finding to resource context and a remediation workflow. Review the impact, choose which state should win, and route the change through policy and approval. Inspect the proposed code and infrastructure actions before reconciliation.

How to Prevent Drift

Detection and remediation are reactive. Prevention is better.

Policy gates can block non-compliant changes that pass through the governed deployment path. They do not prevent a separate administrator, unmanaged automation, or another control plane from making a change outside that path. Combine deployment checks with access controls and live-state detection.

RBAC and access controls reduce console access. If fewer people can click buttons in the AWS console, fewer manual changes happen. This isn't about trust. It's about removing the temptation to take shortcuts.

Scheduled scanning can catch drift sooner. Detection time depends on the scan interval, permissions, API availability, and resource coverage. Check the last successful scan and any failures before assuming that a quiet dashboard means no changes occurred.

GitOps workflows require every change to go through version control. No direct applies, no console changes, no SSH fixes without a matching PR. This is the gold standard but it requires discipline and good tooling to enforce.

Drift in Kubernetes

Kubernetes has its own drift problem. Helm charts declare desired state but operators, controllers, and manual kubectl commands modify actual state constantly. Someone runs kubectl edit to fix a broken deployment. A controller scales a replica set. An admission webhook modifies a pod spec.

ops0 tracks 31 Kubernetes resource types for drift, comparing Helm chart state against live cluster state. This is harder than cloud drift because Kubernetes is designed to be dynamic. The challenge is distinguishing intentional changes (auto-scaling, rolling updates) from actual drift (manual edits, configuration mistakes).

The Real Cost of Drift

Drift isn't a technical problem. It's a trust problem. When your code doesn't match reality, you can't trust your disaster recovery plan. You can't trust your compliance reports. You can't trust that a terraform apply won't break production.

Teams that don't manage drift end up afraid to make changes. They stop running terraform plan because the output is too scary. They stop updating infrastructure because nobody knows what will happen. The infrastructure calcifies and becomes impossible to maintain.

Controlling drift means comparing the intended configuration with observed state, assigning ownership, and reconciling through review. Automation helps collect and explain differences; the remediation decision still needs context.

For the operational workflow, see drift prevention. For dependency context before you change a resource, review Resource Graph.

Quick answers

What is infrastructure drift?

Infrastructure drift happens when live cloud resources no longer match the declared state in Terraform, OpenTofu, CloudFormation, or another source of record.

How does ops0 detect infrastructure drift?

ops0 compares declared infrastructure with discovered cloud state and uses Resource Graph context to show what changed and what depends on it.

How should teams fix drift safely?

Teams should inspect blast radius, decide whether the live change or IaC source should win, route remediation through review, and preserve audit evidence.

Sources
From article to workflow

Move from a drift finding to a reviewed change.

Find the resource, confirm the intended configuration, review the plan, and retain the evidence after the approved fix.

Related articles

All articles