Infrastructure as code (IaC) is the practice of defining servers, networks, permissions and other infrastructure in version-controlled files that a tool applies until reality matches. Changes go through review like any other code change, environments can be rebuilt the same way, and the repository, not someone's memory, describes what exists.
This guide goes deeper than the definition: how the declarative model works, how Terraform and OpenTofu actually differ, what the state file holds and how to protect it, a complete pull request to apply pipeline in GitHub Actions, repository layout, the failure modes that surface in year two, and when IaC is the wrong tool. Product names, versions and defaults are as documented on 9 October 2026.
What is infrastructure as code?
Infrastructure as code replaces point-and-click console work with definition files. Microsoft's overview puts the mechanism in one line: just as the same source code always generates the same binary, an IaC model generates the same environment every time it deploys. HashiCorp's tutorial describes the language property underneath: Terraform's configuration language is declarative, meaning it describes the desired end state of your infrastructure, in contrast to procedural programming languages that require step-by-step instructions. You write that a network with these subnets and a bucket with these tags should exist; the tool works out the creates, updates and deletes, and Terraform's providers calculate the dependencies between resources automatically.
Two properties follow. The first is idempotence: a deployment command always sets the target environment into the same configuration, regardless of the environment's starting state, which is achieved either by configuring the existing target or by discarding it and recreating a fresh one. The second is a single edit point: to make changes, the team edits the source, not the target. Microsoft explains why this matters: without IaC, each environment becomes a "snowflake," a unique configuration that cannot be reproduced automatically, and IaC evolved to solve that environment drift in release pipelines. The cloud can also provision and tear down environments from IaC definitions, which makes extra test environments practical.
A second idea decides how a change reaches running infrastructure. The AWS Reliability Pillar defines immutable infrastructure as a model where no updates, security patches or configuration changes happen in place on production workloads. When a change is needed, a new set of updated resources is deployed in parallel to the existing ones, the deployment is validated automatically, and if it succeeds, traffic is gradually shifted to the new set. The previous working version is not changed, so you can roll back to it if errors are detected. Among the benefits AWS lists are less configuration drift, because resources are replaced with a known, version-controlled configuration, and remote access such as SSH can be disabled, which reduces the attack vector.

Declarative files describe what should exist; immutability decides how a change arrives. In practice that buys review (every change is a pull request with a diff), repeatability (the same files build the same environment), recovery (rebuilding is an apply, not a memory exercise) and speed (a new test environment is another run of the same pipeline).
Which tools manage infrastructure as code?
The tools split into three jobs: provisioning cloud resources, configuring what runs on them, and continuously reconciling a platform. Each fact in the table comes from the vendor's, project's or foundation's own pages, checked in October 2026.
| Tool | Scope | You write | State and execution | Licence and steward |
|---|---|---|---|---|
| Terraform | Resources on multiple cloud platforms, through providers | Declarative configuration files | A state file you place: local, or a remote backend such as S3 | Business Source License 1.1, licensor IBM |
| OpenTofu | The same providers, with most Terraform code unchanged | Declarative configuration files, with an optional .tofu extension | Backends such as S3, plus optional end-to-end encryption | MPL 2.0, Linux Foundation project, CNCF sandbox |
| Pulumi | Cloud infrastructure, in general-purpose programming languages | Node.js, Python, Go, .NET, Java or YAML (list) | Pulumi Cloud by default, or your own S3, Azure Blob, GCS or PostgreSQL | Apache-2.0 |
| AWS CloudFormation and CDK | AWS resources | Templates; the CDK adds six languages that provision through CloudFormation | Stacks, managed as a single unit by the service | Managed AWS service; the CDK is open source |
| Azure Bicep | Azure resources | A declarative DSL the Bicep CLI converts to a Resource Manager JSON template | No state files to manage: Azure stores all states | Free and open source, supported by Microsoft Support |
| Crossplane | Cloud native software, through a Kubernetes control plane | Compositions (functions in YAML, KCL, Python or Go) and managed resources | Monitored for its whole lifecycle, with drift corrected automatically | Apache-2.0, CNCF graduated |
| Ansible | System configuration and software deployment | YAML playbooks | Playbooks declare the desired state of a system, run over SSH with no agents to install | GPL-3.0, sponsored by Red Hat |
Provisioning and configuration are often paired. IBM's acquisition announcement frames the pairing itself: Terraform and the Red Hat Ansible Automation Platform, with Terraform automating the foundational infrastructure creation across multiple cloud providers, and Ansible automating application configurations and middleware deployments on top of that infrastructure. Ansible's own README also lists cloud provisioning among what it handles, so treat the split as a common pattern, not a rule. Ansible's playbooks declare the desired state of a system and are idempotent in the same sense: when the system is in the state a playbook describes, Ansible does not change anything, even if the playbook runs multiple times.
Crossplane is the different model in the list. It is a control plane framework built on Kubernetes: it configures your software, then monitors it throughout its lifecycle, and if the software drifts from the desired state it corrects the drift automatically. The CNCF accepted it in June 2020 and graduated it in October 2025. The trade-off is that you now run Kubernetes as your control plane, with everything that involves; our OpenShift vs Kubernetes comparison covers that platform decision, including Operators and packaged Argo CD GitOps. The remaining rows differ in how you write the code and where state lives: Pulumi lets you write in general-purpose languages, Bicep's what-if previews changes before deployment, and CloudFormation manages what it deploys as stacks inside AWS.
Terraform or OpenTofu in 2026?
The split is three years old and now mostly about licence, a short list of features, and whether you rely on HashiCorp's HCP Terraform. On 10 August 2023 HashiCorp announced it was changing its source code licence from the Mozilla Public License 2.0 to the Business Source License 1.1 on all future releases, keeping its APIs, SDKs and almost all libraries under MPL 2.0. The licence text, which names IBM as licensor and covers Terraform 1.6.0 and later, grants production use provided you do not offer the work to third parties on a hosted or embedded basis to compete with IBM's paid versions. It spells out that hosting or using Terraform internally within an organization is not a competitive offering, and each version converts to MPL 2.0 four years after it is published. IBM completed its acquisition of HashiCorp on 27 February 2025.
OpenTofu was the community's answer: the Linux Foundation announced it on 20 September 2023 (it was previously named OpenTF), and the project FAQ describes it as a Terraform fork created as an initiative of Gruntwork, Spacelift, Harness, Env0, Scalr and others. It is published under MPL 2.0 and was accepted into the CNCF sandbox on 23 April 2025. OpenTofu will not have its own providers: it works with the current Terraform providers through a separate registry. On state, the FAQ says OpenTofu will work with state files created by Terraform up to 1.5.x, and the 1.7.0 release describes OpenTofu as a drop-in replacement for Terraform 1.5 with easy migration paths from later versions.
Both projects kept moving after the fork. As of 9 October 2026, the GitHub release pages list v1.16.5 as Terraform's latest release (with v1.17.0-rc1 in pre-release) and v1.13.1 as OpenTofu's, and OpenTofu says the 1.13 series will be supported until August 2027. The timeline shows how the feature sets diverged. Each column is its own timeline, so entries in the same row are not equivalents:
| Terraform | OpenTofu |
|---|---|
| 1.5 (June 2023): import blocks, check blocks, config generation for imports | |
1.7 (January 2024): removed block, import for_each, test mocks | 1.7 (April 2024): end-to-end state encryption (AWS KMS, GCP KMS, OpenBao), removed, loopable import |
| 1.10 (November 2024): ephemeral resources, S3 native state locking | 1.8 (July 2024): variables and locals in module sources, backend configuration and state encryption; .tofu extension |
| 1.11 (February 2025): write-only attributes, S3 native locking generally available, DynamoDB locking deprecated | 1.9 (January 2025): provider for_each for multi-region or multi-zone setups, the -exclude flag |
1.14 (November 2025): list resources and terraform query, provider-defined actions | 1.10 (June 2025): OCI registry support, native S3 locking without DynamoDB, deprecation attributes |
1.15 (April 2026): variables and locals in module source and version, deprecated on variables and outputs | 1.11 (December 2025): ephemeral resources and write-only attributes, the enabled meta-argument |
Read the table and two differences stand out. Built-in encryption is the clearest: OpenTofu added end-to-end state encryption in 1.7 and now encrypts state and plan files at rest, for local storage or a backend, while Terraform's own sensitive-data guide says the encryption method depends on your backend (HCP Terraform, for example, encrypts state at rest automatically, and the S3 backend has an encrypt option). The second is evaluation of variables: OpenTofu 1.8's early evaluation let variables configure backend settings and state encryption, its release notes say Terraform did not support those language features, and Terraform 1.15 later adopted the module-source part. In the other direction, Terraform 1.14 added terraform query and provider-defined actions, and HCP Terraform, which HashiCorp's own tutorial offers as an alternative to in-house automation, adds policy enforcement and health assessments (covered below).
The choice follows your constraints. If your organization requires an open-source licence such as MPL 2.0, or state that is encrypted before it reaches any storage service, OpenTofu fits. If you run HCP Terraform or want its health assessments, Terraform is the natural choice. For a greenfield project either can work, because most Terraform code runs unchanged on OpenTofu and the providers are shared. And the decision is not a leap: the migration overview calls the process designed to be safe and reversible, and OpenTofu's migration guide breaks it into six steps.
- Back up your state (for remote state, S3 bucket versioning or a snapshot) and commit all configuration to version control.
- Install OpenTofu and verify with
tofu --version. - Run
tofu initin the project directory, which downloads providers from the OpenTofu registry and prepares the backend. - Run
tofu plan. The expected result is "No changes" or the same plan Terraform would show. If you see unexpected changes, do not apply; investigate or roll back. - Run
tofu applyso OpenTofu updates the state file format where needed, even when no infrastructure changes. - Make a small, non-critical change, such as adding a tag, and run plan and apply again to prove the new tool manages the infrastructure going forward.
Rolling back is the mirror: stop, restore the backups if state changed, then terraform init and terraform plan to verify before continuing with Terraform. One caveat from the migration overview: estates built from several configurations that share data through terraform_remote_state typically need some additional care.

What lives in the state file, and how do you protect it?
State is the record Terraform uses to decide what a plan will do. The state file stores the bindings between real objects and the resource instances declared in your configuration, Terraform expects a one-to-one mapping between them, and before any operation it refreshes the state against the real infrastructure. By default all of this lives in a local terraform.tfstate file (JSON, with a .backup of the previous state), which HashiCorp says not to edit directly.
Why state stays out of Git
The state page warns against storing state in a version control system or other storage that does not support state locking and secure access control, because doing so can result in data loss or exposure of the secrets stored in the file. The sensitive-data page is more blunt: state can contain sensitive values such as initial database passwords or API tokens, and if you work locally, Terraform stores it in a plaintext file that includes any secret values you defined in your configuration.
Marking a variable or output sensitive (available since Terraform 0.15) redacts it from CLI output and the HCP Terraform UI, but the values are still stored in the state and plan files, and terraform output -json or -raw prints them in plain text. The stronger fixes arrived after the fork, in both tools: ephemeral values (Terraform 1.10, OpenTofu 1.11) are not persisted to state or plan files, and write-only attributes (Terraform 1.11, OpenTofu 1.11) let a resource receive a secret without storing it. Write-only arguments typically end with _wo and have corresponding _wo_version arguments, for example password_wo on aws_db_instance.

A backend with locking, versioning and encryption
A remote backend moves state off laptops and lets a team share it. This is a minimal S3 setup with one tagged resource:
terraform {
required_version = ">= 1.11" # S3 native locking became GA in 1.11
required_providers {
aws = {
source = "hashicorp/aws"
version = "~> 6.0"
}
}
backend "s3" {
bucket = "example-iac-state"
key = "prod/terraform.tfstate"
region = "us-east-1"
encrypt = true # server-side encryption of state and lock files
use_lockfile = true # native S3 locking; creates prod/terraform.tfstate.tflock
}
}
provider "aws" {
region = "us-east-1"
default_tags {
tags = {
Environment = "prod"
ManagedBy = "terraform"
}
}
}
resource "aws_s3_bucket" "logs" {
bucket = "example-logs-111122223333" # changing this name replaces the bucket
tags = {
Purpose = "access-logs"
}
lifecycle {
prevent_destroy = true
}
}
Each setting in that block is documented. State locking on the S3 backend is opt-in: use_lockfile defaults to false, and when enabled it needs s3:GetObject, s3:PutObject and s3:DeleteObject on the <key>.tflock object. HashiCorp highly recommends bucket versioning for state recovery after accidental deletions and human error, and kms_key_id (which needs kms:Encrypt, kms:Decrypt and kms:GenerateDataKey on the key) encrypts the state and lock files with a KMS key you choose. The older DynamoDB locking is deprecated and will be removed in a future minor version; both can be configured together while you migrate. OpenTofu's S3 backend takes the gentler line: its native locking uses conditional writes with an If-None-Match header, and both mechanisms are fully supported with no plans to deprecate either.
Credentials need care too. Terraform's S3 backend page recommends environment variables, because values hardcoded in the backend block or passed with -backend-config are included in the .terraform subdirectory and in plan files, and OpenTofu's page notes that -backend-config settings are saved to disk under .terraform.
Locking itself is automatic on every operation that could write state, and if the lock cannot be acquired the command does not continue; with -lock-timeout, a run retries acquiring the lock for a period of time before it returns an error. Occasionally the automatic release fails and a lock stays behind. Terraform has a manual release command for that case:
terraform force-unlock LOCK_ID
The locking page is emphatic: use it only for your own lock, when the automatic release failed, because releasing a lock that someone else is holding could cause multiple writers.

The provider's default_tags block sets tags for every resource that supports tags, with aws_autoscaling_group documented as the exception. The defaults merge with a resource's own tags into tags_all, a resource's own tags can override a default per key, and a default cannot be excluded from a resource. Our FinOps guide picks up the cost side of tagging. Finally, prevent_destroy = true makes Terraform reject any plan that would destroy the bucket. Two caveats from the lifecycle reference: it does not protect a resource whose configuration you delete, and HashiCorp advises using it sparingly. Note also that the bucket argument forces a new resource when it changes, which prevent_destroy turns into a plan error instead of a surprise replacement.
If you chose OpenTofu, encryption of state and plan files is built in and configured in the same files. Variables in the encryption block need OpenTofu 1.8 or later:
variable "state_passphrase" {
type = string
sensitive = true # the passphrase needs at least 16 characters
}
terraform {
encryption {
key_provider "pbkdf2" "state" {
passphrase = var.state_passphrase
}
method "aes_gcm" "state" {
keys = key_provider.pbkdf2.state
}
state {
method = method.aes_gcm.state
enforced = true
}
plan {
method = method.aes_gcm.state
enforced = true
}
}
}
The encryption documentation carries three warnings worth repeating. Once encrypted, state files become unrecoverable without the encryption key, so keep key backups; encryption does not protect against data loss or against a replay attack with an older state or plan file; and do not rename key providers or methods once data is encrypted (use a fallback block to roll keys). For an existing plaintext state, the migration path adds a method "unencrypted" "migrate" fallback first. If you use a key management system (AWS KMS, GCP Cloud KMS, Azure Key Vault or OpenBao), use a separate key for each state file rather than sharing one.
Import, move and remove without destroying
State changes used to be CLI commands run by hand. Since Terraform 1.5 (June 2023) the declarative path is an import block: to names the resource address, and id (or, where the provider supports it, identity; the two are mutually exclusive) names the real object, which must be known during the plan operation. for_each imports whole sets, and plan -generate-config-out writes a starting configuration, a flag the plan reference still marks experimental. Terraform 1.14's terraform query lists existing infrastructure and can optionally generate configuration for importing the results. For an S3 bucket the ID is the bucket name, as the AWS provider documents, so adopting one looks like this:
import {
to = aws_s3_bucket.logs
id = "example-logs-111122223333"
}
Renames and moves are the second operation. By default Terraform interprets an address change as an instruction to destroy the existing resource and create a new one; a moved block (it needs Terraform v1.1 or later) records the rename instead, and the address change does not destroy the resource. Removing a moved block is a breaking change, because any configuration that still refers to the old address will plan to delete the existing object instead of moving it, so HashiCorp strongly recommends keeping the historical moved blocks of a module.
moved {
from = aws_s3_bucket.logs
to = aws_s3_bucket.access_logs
}
The third operation is letting go. The removed block, added in Terraform 1.7 as a configuration-driven replacement for terraform state rm, takes a resource out of state; its required lifecycle block decides whether the real object is destroyed (the default) or kept with destroy = false, which lets you hand a resource to another tool or team.
What does a safe team workflow look like?
The workflow is the point of IaC. HashiCorp's automation tutorial gives the main path as four steps: initialize the working directory (terraform init -input=false), produce a plan (terraform plan -out=tfplan -input=false), have a human operator review that plan, then apply the changes the plan describes (terraform apply -input=false tfplan). Everything else in this section is the machinery that keeps those four steps honest when ten people are involved.

Four points from the same tutorial deserve emphasis. Pull request plans are "throwaway": they exist to aid code review, and after the merge you plan and apply again from the main branch, because other changes may have landed in between (the plan command page makes the same point about speculative plans). The recommended approach is to allow only one plan to be outstanding at a time; the alternative is to connect the approval to the apply so that the exact plan that was approved is the one applied. -auto-approve belongs on non-critical infrastructure: the tutorial says manual review of plans is always recommended when Terraform can make destructive changes, unless downtime is tolerated. And with TF_IN_AUTOMATION set to any non-empty value, Terraform de-emphasizes the specific commands it would otherwise suggest running, which the tutorial says can be confusing and un-actionable when an automation tool wraps Terraform.
Warning
A saved plan file is not a sanitized preview. terraform plan -out stores your full configuration and the values of the planned changes, and any sensitive data in the plan is saved in cleartext even when the terminal output hides it. Treat saved plan files, and any CI artifact containing them, as potentially sensitive: encrypt them, restrict access, and keep retention short.
A pull request to apply pipeline in GitHub Actions
This workflow implements the pattern. On a pull request it checks formatting with terraform fmt -check, initializes, plans, and publishes the plan to the job summary with terraform show. On a push to main it plans again, archives the whole workspace, and applies only after a person approves through a protected environment:
# .github/workflows/iac.yml
name: iac
on:
pull_request:
paths: ["environments/prod/**"]
push:
branches: [main]
paths: ["environments/prod/**"]
permissions:
id-token: write # OIDC token for configure-aws-credentials
contents: read # actions/checkout
concurrency:
group: iac-prod-${{ github.ref }}
cancel-in-progress: false # never cancel a run that may hold the state lock
defaults:
run:
working-directory: environments/prod
env:
TF_IN_AUTOMATION: "true"
TF_VERSION: "1.16.5" # the exact release you run locally
jobs:
plan:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v7
- uses: aws-actions/configure-aws-credentials@v6.3.0
with:
role-to-assume: arn:aws:iam::111122223333:role/iac-plan
aws-region: us-east-1
- uses: hashicorp/setup-terraform@v4
with:
terraform_version: ${{ env.TF_VERSION }}
terraform_wrapper: false
- run: terraform fmt -check -recursive "$GITHUB_WORKSPACE"
- run: terraform init -input=false
- run: terraform plan -out=tfplan -input=false
- name: Publish the plan for review
run: terraform show -no-color tfplan >> "$GITHUB_STEP_SUMMARY"
- name: Archive the initialized workspace for apply
if: github.event_name == 'push'
run: tar --exclude=./iac.tar -C "$GITHUB_WORKSPACE" -cf "$GITHUB_WORKSPACE/iac.tar" .
- uses: actions/upload-artifact@v7
if: github.event_name == 'push'
with:
path: iac.tar
archive: false # upload the tar as one file, so file modes survive
retention-days: 1 # the plan inside is sensitive
apply:
if: github.event_name == 'push'
needs: plan
runs-on: ubuntu-latest
environment: production # required reviewers approve the apply here
steps:
- uses: actions/checkout@v7 # the working directory must exist before run steps
- uses: actions/download-artifact@v8
with:
name: iac.tar # the file name, because it was uploaded with archive: false
- run: tar -xf "$GITHUB_WORKSPACE/iac.tar" -C "$GITHUB_WORKSPACE"
- uses: aws-actions/configure-aws-credentials@v6.3.0
with:
role-to-assume: arn:aws:iam::111122223333:role/iac-apply
aws-region: us-east-1
- uses: hashicorp/setup-terraform@v4
with:
terraform_version: ${{ env.TF_VERSION }}
terraform_wrapper: false
- run: terraform apply -input=false tfplan
A few details in that workflow come straight from the documentation. permissions grants an OIDC token (id-token: write) and read-only contents, per the workflow syntax reference. The concurrency group key is built from the github context, one of only three contexts (github, inputs and vars) a concurrency expression can use. By default only one pending run waits per group and a newer one replaces it, and cancelling runs that are already in progress requires cancel-in-progress: true, so the workflow sets it to false to keep a run that may hold the state lock from being cancelled midway.
The configure-aws-credentials action assumes an IAM role with the OIDC token, and its README carries the warning that decides your trust policy: without any condition, any GitHub user or repository could potentially assume the role. Scope the sub claim, which includes environment:<ENVIRONMENT_NAME> for jobs that run in an environment. Use two roles in the same AWS account: a plan role that can read your resources and take the state lock, and an apply role that can also change them. The automation tutorial notes that Terraform cannot detect whether the credentials used to plan and to apply reach the same resources.
The apply job's environment is where the human gate lives: up to six required reviewers, of whom one approval is enough; an option to prevent self-reviews, so the person who initiates the deployment cannot approve it; and deployment branch rules that can restrict deploys to protected branches. Plan limits apply: on the Free, Pro and Team plans, required reviewers are only available for public repositories, and by default administrators can bypass protection rules.
The archive step exists because a saved plan can contain absolute paths, and Terraform assumes the same operating system, CPU architecture and identical provider plugins at apply time. The automation tutorial says to archive the entire working directory, including .terraform, and extract it at the same absolute path. It also says an error is produced if Terraform or any plugin is upgraded between creating and applying a plan, which is why both jobs install the same pinned terraform_version. Both jobs use $GITHUB_WORKSPACE, which is the default location of your repository on the runner when you use the checkout action.
upload-artifact v7 zips uploads and does not maintain file permissions, so the workflow tars first and uploads the tar with archive: false, and download-artifact v8 then fetches the single file by its file name. checkout runs in the apply job so the defaults.run.working-directory directory exists before the first run step, as the syntax reference's tip requires. The action versions are the ones documented on 9 October 2026: checkout@v7, upload-artifact@v7, download-artifact@v8, setup-terraform@v4 and configure-aws-credentials@v6.3.0. Setting terraform_wrapper: false skips the wrapper script that otherwise exposes Terraform's STDOUT, STDERR and exit code as step outputs.
Policy checks on the plan
Policy as code adds an automatic check on the plan, so a reviewer does not have to catch every forbidden change by eye. OPA's Terraform tutorial writes policies in Rego against the JSON plan (terraform show -json tfplan.binary > tfplan.json) and checks changes before Terraform makes them, with the caveat that unknown values, dynamic blocks and function calls that have not been evaluated yet limit what a plan-time policy can see. Conftest relies on Rego too, runs with conftest test, and looks for policies in a policy directory by default.
For ready-made rules, Checkov ships over 1,000 built-in policies and scans Terraform, Terraform plan files, CloudFormation, Bicep and OpenTofu, and Trivy's trivy config covers Terraform and CloudFormation misconfiguration scanning (tfsec is now part of Trivy). On HCP Terraform, policy enforcement lets you validate plans with one of three frameworks: Terraform policy (beta, a native HCL-based framework), Sentinel or OPA. Pipeline checks like these have a supply-chain counterpart in a software bill of materials, which our SBOM guide covers, and our CI/CD pipeline guide goes into the wider stage design.
Drift detection
Code is only truthful if someone checks reality against it. A scheduled workflow running terraform plan -refresh-only (available only in Terraform v0.15.4 and later) creates a plan whose goal is only to update state to match changes made to remote objects outside of Terraform. In HashiCorp's refresh-only tutorial the output lists the changes detected outside Terraform and notes that a refresh-only plan takes no action to undo them, which makes the drift visible without proposing infrastructure changes. HCP Terraform health assessments, on the Standard and Premium editions, run periodically: drift detection needs Terraform 0.15.4 or later, continuous validation of your check blocks needs 1.3.0 or later, and resolving drift means either overwriting it (apply the code to revert reality) or updating the configuration to keep the change.
CloudFormation has drift detection built in, but it only determines drift for property values that are explicitly set, it does not detect drift on nested stacks (you run it on the nested stack directly), and resources that do not support it are marked NOT_CHECKED. Crossplane, per its documentation, corrects drift automatically.

Pinning providers, modules and tools
Version pinning is the last guardrail, and it is layered. Each root module should declare its providers in required_providers so that Terraform can install and use them. The ~> operator allows only the right-most version component to increment (~> 6.0 allows 6.x but not 7.0), and HashiCorp's guidance is minimum-only constraints in reusable modules and a ~> constraint with both bounds in root modules. The dependency lock file, .terraform.lock.hcl, records the selected provider versions and belongs in version control, but it tracks providers only: Terraform always selects the newest available module version that meets your constraints unless you pin modules exactly, and terraform init -upgrade re-selects provider versions deliberately.
Provider majors need their own ritual. The AWS provider's version 6 upgrade guide is a good model: first upgrade to the latest 5.x, confirm a plan with no errors, no unexpected changes and no related deprecation warnings, then allow ~> 6.0 and run terraform init -upgrade. Expect the cadence: the AWS provider's changelog lists 68 minor releases, 6.1.0 to 6.68.0, between 6.0.0 (18 June 2025) and 6.68.0 (7 October 2026), roughly one a week.
How do you lay out the repository?
Layout is a blast-radius decision: a configuration and its state are what one apply can change. One layout that works puts reusable modules beside per-environment roots, one state each:
infrastructure/
modules/
app-service/ # reviewed once, versioned, reused
environments/
dev/
backend.tf # dev state key, dev role
main.tf
prod/
backend.tf # prod state key, prod role
main.tf
One state per environment means a bad apply in development cannot touch production's record of reality, and each environment can have its own backend credentials. HashiCorp gives related advice for oversized configurations: the plan page says -target is for exceptional circumstances because routine use can lead to undetected configuration drift, and prefers breaking large configurations into several smaller ones that can each be independently applied. A landing zone is a natural first candidate for this layout; our cloud migration guide describes that foundation (accounts, identity, network, guardrails, logging) as something to define as code before the first workload lands.

What about Terraform workspaces, one configuration with several states? The workspaces page says they are not appropriate for system decomposition or for deployments that need separate credentials and access controls, which is what environments often are. (The automation tutorial recommends, where possible, one backend configuration for all environments with terraform workspace to switch between them, and notes that environments in entirely separate accounts need different credentials or endpoints for the backend itself. Read together, the two pages point to workspaces for copies that share a backend and credentials, and to a separate backend configuration, as in the layout above, for anything with its own account.)
Two smaller rules. Keep secrets out of the repository entirely: .gitignore the *.tfstate* files and .terraform/, pass credentials as environment variables, and rotate any secret that was ever committed (our secure coding checklist covers the pattern). And pin third-party modules to exact versions, because the lock file will not do it for you.
What goes wrong in practice?
Day-two trouble with IaC tends to have three roots: someone changed reality outside the code, the state record was not protected, or versions moved under your feet. The table lists symptoms with their likely causes and fixes:
| Symptom | Likely cause | Fix |
|---|---|---|
| A plan shows changes nobody proposed | Someone edited a resource in the console or with a script | Run terraform plan -refresh-only to see the difference, then decide: apply the code to overwrite it, or update the code to keep it. Restrict console write access so it does not recur |
| A run fails because the state is locked | Another run holds the lock, or a run that died did not release it | Wait for the other run; -lock-timeout makes a run retry for a while. If nothing is running and automatic release failed, release your own lock by its ID with the manual command shown earlier, never someone else's |
| After a provider upgrade, the plan shows unexpected changes | Behaviour changed in the new major version | Follow the provider's upgrade guide: the latest previous major first, a clean plan, then the new major (the AWS v6 guide's sequence). Pin with ~> so majors never arrive unplanned |
| A secret shows up in the repository or a build artifact | State or plan files were committed or archived; both can hold sensitive values in cleartext | Move state to an encrypted, locked backend, keep artifact retention short, rotate the leaked value, and use ephemeral or write-only values for new secrets |
| Applies happen that nobody remembers approving | -auto-approve on critical infrastructure, or applies from laptops | Apply only from the pipeline behind a protected environment with required reviewers; keep auto-approval for non-critical infrastructure |
| The drift alarm fires every Monday | Another tool or person manages the same resources | One writer per resource: bring the management into code, or stop managing it from Terraform with a removed block that sets destroy = false |
Each row has a shared moral: the fix flows through the repository, never around it. The moment you correct production by hand and move on, the code stops being the truth and the next apply becomes a negotiation.
When is infrastructure as code not worth it?
Skip it when there is nothing to repeat. A single static site on a managed platform, a proof of concept you will delete next week, a one-person project where the console genuinely is faster: the overhead of backends, locks, reviews and provider upgrades buys nothing when there is no second environment and no second person. Be honest about the running costs: a state backend to secure and monitor, module and provider versions to keep current (recall the AWS provider's weekly minors), plans that need a competent reader on every change, and a drift routine that somebody must own. None of that is free, and all of it outlives the enthusiasm of the first month.
The boundary sits lower than it first appears, though. Even a small serverless deployment benefits: our serverless guide lists defining the infrastructure as code, so triggers, permissions and limits are written down and reviewed rather than clicked together, as one of four habits that keep the way out open. Recovery is the other argument: rebuilding an environment from code turns recovery into an apply instead of a search for whoever remembers how it was built, which is why recovery targets and rebuild discipline belong together, as in RTO vs RPO. The same discipline shows in recovery drills: a financial firm's disaster recovery drill was automated end to end, with a named person approving every step. An estate migration is another good fit: if you are weighing VMware alternatives, define the target platform as code from the first host rather than rebuilding hand-built habits on a new hypervisor.
Adopt infrastructure as code in one environment this month
- Pick one environment of one service that hurts: the staging estate nobody can rebuild, the network only one person understands. Non-production, real enough to matter.
- Stand up the backend first: a versioned, encrypted state bucket with locking enabled, separate from anything else, one state for this environment. This is the piece that protects every later step.
- Import what exists. Write import blocks for the current resources and iterate until
terraform planreports no changes. Now the code is the truth, and every later change is a diff. - Pin everything:
~>constraints in the root module, the dependency lock file committed, exact versions for third-party modules, an exact tool version in CI. - Wire the pipeline: plan on every pull request with the plan published for review, apply on merge through a protected environment with one required reviewer who is not the author.
- Add the guardrails: Checkov or OPA policies on the plan, a scheduled refresh-only plan for drift, and the AWS provider's
default_tagsso every resource carries the same tags from day one. - Write the runbook: who reviews plans, how a stuck lock is handled, how drift is resolved, how a leaked secret is rotated. Thirty lines is enough; the test is that someone else can follow it.
Computese builds landing zones this way: accounts, networks, identity, policies and clusters defined in Terraform and Git, reproducible in every environment, with policy as code, central logging, and tagging, budgets and guardrails from the first day. The signal we look for is the one this post keeps circling: environments drift because they are built by hand. If that sounds like your estate, cloud transformation is the service to start with.


