Live · Fri, Oct 2, 2026 · 01:01 UTC Block 843,917 Fees 14 sat/vB Fear & Greed 72 · Greed
Newsletter Pro Terminal Sign in
ITop Field News.
Subscribe →
Live · 01:01 UTC Block 843,917 F&G 72
Cloud & infrastructure Cloud & infrastructure desk

Cloud finite state: how to detect and stop configuration drift before it bites

Cloud configuration drift is one of the quietest risks in modern infrastructure: live environments diverge from their declared state, security controls weaken, and nobody notices until something breaks.

Detailed view of network cables plugged into a server rack in a data center.

Photo by Brett Sayles on Pexels

Cloud configuration drift is the gap between what your infrastructure is supposed to look like and what it actually looks like right now. It starts small. A developer opens a port for a weekend test and forgets to close it. A security group gets loosened to fix a deployment. A storage bucket loses a lifecycle rule. None of these feel catastrophic in isolation, but they compound. Over weeks, your live environment quietly diverges from your declared state, and the risk attached to that gap is very real.

For Australian IT teams, the stakes are higher than they might appear. The Australian Signals Directorate consistently flags misconfiguration as one of the top causes of cloud security incidents. Under the Privacy Act and the Notifiable Data Breaches scheme, a configuration gap that exposes personal data isn't just a technical embarrassment; it carries legal obligations and potential penalties.

Why drift is harder to catch than it looks

Most teams assume their infrastructure-as-code gives them a reliable picture of current state. It doesn't. IaC defines the intended state at the moment code is applied. Anything that happens after a terraform apply or an aws cloudformation deploy sits outside that snapshot. Manual console changes are the obvious culprit, but they're not the only one. Auto-scaling events, vendor-applied patches, and even some orchestration tools can alter resource configurations in ways that don't propagate back to the source repo.

This is also why cloud environment drift tends to accelerate over time. The longer a team goes without reconciling live state against declared state, the harder the remediation becomes. What starts as a small delta turns into a significant gap, and closing it requires unpicking changes that may have been in production for months.

What good drift detection actually looks like

Effective drift detection works in continuous time, not in point-in-time audits. Quarterly reviews catch problems weeks after they've already been exploited. The minimum viable approach is a pipeline that compares live resource state against your IaC definitions on every change event, not just on scheduled runs.

Three concrete practices make a real difference:

  • State locking with authoritative backends. Terraform remote state in S3 (with DynamoDB locking) or Azure Blob Storage with lease-based locking ensures a single source of truth. Any live resource that no longer matches a state entry triggers an alert, not just a diff in the next plan.
  • CloudTrail or equivalent audit logging piped to a SIEM. Every API call that modifies a resource should generate a log entry that your SIEM correlates against expected change windows. Out-of-band changes become detectable within minutes, not weeks.
  • Policy-as-code enforcement at the PR stage. Tools like Open Policy Agent (OPA) or AWS Config rules reject non-compliant resource configurations before they reach production, shrinking the surface on which drift can start.

The IAM angle most teams miss

Configuration drift in IAM is particularly dangerous because it's often invisible in standard resource scans. A role that accumulates permissions through manual grants, a service account that gets an admin binding added for a one-off task, a cross-account trust that never gets removed: these are all forms of drift, and they don't show up in a diff of your compute or network resources.

This connects directly to why cloud IAM misconfigurations are so persistent across Australian organisations. Even teams with solid IaC practices for compute and networking often have a manual, ad-hoc approach to IAM changes. The fix is the same: treat IAM definitions as code, review every permission change through a pull request, and run continuous comparison against a known-good baseline.

Tooling worth knowing

No single tool covers the full drift surface. The combination that works best depends on your cloud and your existing stack, but a few tools have earned their place across Australian enterprise environments.

AWS Config continuously records resource configurations and evaluates them against managed or custom rules. It covers most AWS resource types and integrates natively with CloudTrail, making it the practical starting point for AWS-heavy teams. For teams running Terraform across any cloud, terraform plan run in CI (not just locally) provides a live diff against state on every merge request. Driftctl, now merged into Snyk IaC, adds detection of resources that exist in the cloud but have no corresponding IaC definition at all, which is the category most teams forget to look for.

Remediation: the part teams skip

Detection without remediation is just a more detailed alert queue. The harder problem is deciding what to do when drift is found. There are two valid responses: reconcile the code to match the live state (if the change was intentional and correct) or revert the live state to match the code (if the change was unauthorised or mistaken).

The worst outcome is what most teams default to: documenting the drift, adding it to a backlog, and leaving it in place while the backlog grows. Six months later, the original reason for the deviation is forgotten, and the team is afraid to touch it because something might depend on it. Fixing drift at discovery requires a process, not just a tool. Assign ownership, set a time limit for resolution, and treat unresolved drift the same way you treat an open vulnerability.

Making it a team habit, not a compliance exercise

Drift detection programs that live inside a compliance team rarely change developer behaviour. The ones that stick get built into the delivery pipeline so that developers see drift consequences in their own workflow. A failed PR check because the proposed change would introduce a configuration gap is more effective than a quarterly audit report.

The teams that handle this well tend to share one habit: they run terraform plan or an equivalent diff as a required step in every deployment pipeline, not as an optional check. The output isn't ignored. When the plan shows unexpected changes, the pipeline stops. That single practice prevents more drift than any amount of tooling applied after the fact.

→ The Confirmations · Daily newsletter

One email at 06:00 UTC. Six minutes. The only digest written for desks, not for retail.