A pull request changes a Terraform module, the checks turn green, and the release still waits for the same two infrastructure engineers.
They are not debating indentation.
They are trying to determine whether the change will replace a database, broaden production access, or leave the on-call team with an untested recovery path.
Terraform code review needs to answer those questions without turning every infrastructure change into a committee meeting.
The useful distinction is between checks a pipeline can repeat and decisions that require someone who understands the service.
This guide is for platform engineers, DevOps leads, and engineering managers who already use Terraform but need a more reliable review process.
It offers a risk-based checklist, a worked example, and a practical way to assign ownership.
It is not a promise that a clean plan guarantees a safe deployment.
Review the Operational Change, Not Just the Configuration
Treat a Terraform pull request as a proposed change to a running system.
A small source diff can have a large operational effect. A long diff can be a routine regeneration of low-risk configuration.
Start by asking the author to explain three things in plain language:
- What behavior should change for the application or its operators?
- Which environments and dependencies might be affected?
- What evidence will show that the change worked?
“Update module version” describes an edit, not an outcome.
“Increase worker capacity while preserving queue access and the existing recovery path” gives reviewers something meaningful to evaluate.
Microsoft’s Terraform review guidance covers foundations such as provider versions, repository organization, state handling, variables, and tests.
Keep those foundations, but add an operational question:
Can the team explain the consequences of this specific change?
The code’s author does not change that standard.
Human-written, copied, and AI-assisted configuration should pass the same review.
Faster authorship is useful only when reviewers receive enough context to make an informed decision.
Give Reviewers an Evidence Package They Can Actually Use
A reviewer should not have to reconstruct the target environment from job logs or ask the author which plan belongs to the latest commit.
Create a compact review package alongside each change.
Include:
- Source revision
- Target account or subscription
- Environment
- Terraform workspace or state identifier
- Relevant input set
- Dependency versions
- Plan run identifier
- Responsible service owner
Reference sensitive inputs securely rather than copying them into the pull request.
Add a short impact statement covering:
- Expected changes
- Unexpected changes
- Operational checks
- Recovery responsibility
“No expected downtime” is an assumption until the author explains why the affected service can tolerate the transition.
Keep the package proportional.
A non-production tagging change might need only a few sentences.
A shared-network change may require a dependency map and evidence from affected application owners.
The goal is to remove uncertainty, not maximize documentation.
Tie Review to the Plan That Will Actually Run
Terraform’s plan documentation distinguishes a speculative preview from a saved plan intended for execution.
A preview from an earlier revision is useful feedback, but it is not approval of whatever a later deployment job happens to produce.
For a workflow that applies saved plans, bind approval to the actual saved artifact and its source revision.
If your platform generates a new final plan after merge, review that final plan before execution.
Make the relationship between the approved plan, source revision, and deployment visible in the deployment record.
A Terraform Code Review Checklist Organized by Risk
Use the following checks in order.
This sequence is a proposed operating practice, not a Terraform requirement.
The objective is to:
- Confirm the target.
- Examine destructive consequences.
- Inspect access changes.
- Resolve important unknowns.
- Explain drift.
- Confirm recovery and observable acceptance.
1. Confirm the Environment and Scope
Before interpreting resource changes, confirm:
- Account
- Region
- Backend
- Workspace
- Variables
- Module path
A correct modification against the wrong production state is still a dangerous change.
Ask whether the work affects one service or shared infrastructure.
A route table, shared identity role, or central secret can reach consumers that are absent from the pull request.
Require authors to identify those consumers.
If ownership is unclear, stop and assign it rather than assuming another team will notice the change.
Separate Unrelated Changes When Practical
Combining a module upgrade, network refactor, and capacity increase makes it harder to identify which change created the risk.
Split work by independently deployable intent rather than an arbitrary line-count target.
The purpose is not to make pull requests smaller for their own sake.
It is to make the operational consequences easier to understand.
2. Investigate Deletions and Replacements First
Do not approve a plan from the summary count alone.
Identify every deletion and replacement, then explain why each one is expected.
Terraform’s JSON format represents replacement through action combinations that include both creation and deletion.
That makes deletion detection useful for automated triage, but resource count is not a severity score.
A replacement of a disposable worker and a replacement of a stateful service deserve different decisions.
For each consequential resource, ask:
- Is data stored here, or does another component depend on its identity?
- Can old and new instances coexist?
- Will names, endpoints, credentials, or network relationships change?
- Who has verified the transition against application behavior?
Treat lifecycle settings as mechanisms, not guarantees.
The Terraform lifecycle reference explains constraints around creating replacements before destruction and notes that removing a resource’s configuration can still cause destruction even when that configuration previously used prevent_destroy.
The practical response is to review the complete transition, including capacity and naming constraints, rather than accepting a protective-looking setting as proof.
3. Inspect Access and Network Changes Separately
An in-place update can be more consequential than a replacement.
Pay particular attention to changes involving:
- Identity policies
- Trust relationships
- Ingress and egress rules
- Public exposure
- Access to secrets
- Access to data stores
Ask what new action becomes possible and which identity can perform it.
“Needed for deployment” is not enough detail to assess a permission expansion.
For network changes, identify:
- The intended source
- The intended destination
- Protocol
- Service dependency
Compare the proposed configuration with that intent, including rules being removed as well as rules being added.
Define When a Specialist Review Is Required
Use a named security or platform reviewer when the change crosses a trust boundary.
Do not route every routine infrastructure change to that person.
Instead, define the conditions that require specialist judgment.
If a temporary exception is necessary, record:
- Owner
- Reason
- Expiration
- Removal work
A permission with no end condition is not operationally temporary simply because the pull request describes it as a workaround.
4. Treat Unknown Values as Questions to Resolve
Some Terraform values are available only during execution.
That is normal.
The important question is whether the unknown value affects something that could change the risk classification.
A reviewer should distinguish between:
- A harmless unknown identifier
- An unresolved permission
- An uncertain dependency
- A value that could affect production behavior
Terraform’s JSON representation includes information about values that are not yet known.
An automated check should not silently interpret missing information as a safe value.
Decide How Automation Handles Uncertainty
Define what your policy engine should do when it cannot evaluate an important condition.
Possible responses include:
- Block the change
- Escalate it for human review
- Permit a narrowly defined class of changes
Make that behavior explicit.
For human review, document:
- The unresolved question
- The evidence needed to answer it
- Whether approval remains conditional
If the final value could materially change the risk classification, approval should remain conditional until that uncertainty is addressed.
5. Explain Drift Instead of Approving Around It
When a Terraform plan contains changes the author did not expect, investigate before expanding the approval.
Unexpected differences may come from:
- An emergency console modification
- A changed module input
- Another team’s deployment
- Configuration that has diverged from the declared state
The reviewer needs to understand whether Terraform is expected to preserve the current behavior or replace it with the declared configuration.
Do not remove inconvenient plan output simply to make a pull request appear smaller.
Where possible, separate the intended change from reconciliation work.
If an emergency sequence cannot be separated, document it clearly.
Clarify Who Owns the Operational Truth
If two teams independently believe they manage the same resource, the problem is not simply review speed.
It is ownership.
Resolve that ambiguity before treating the plan as a normal infrastructure change.
6. Check Recovery and Observable Acceptance
A technically successful Terraform apply does not necessarily mean the application is healthy.
Ask what happens if execution completes but the service stops working as expected.
Before approval, define:
- Which signals will be checked
- Who will watch them
- How long they will be observed
- What condition triggers intervention
Choose signals that match the affected service rather than relying on a universal dashboard checklist.
For example:
- A queue change may require checking processing progress and backlog behavior.
- A network change may require validating the specific connections the application depends on.
- A database-related change may require a recovery approach appropriate to its data and provider behavior.
Terraform does not automatically roll back a partially completed apply.
Reverting a Git commit is therefore not a complete recovery plan.
Rehearse Consequential Recovery Steps
Test important recovery procedures in an appropriate non-production environment when practical.
When a full rehearsal is not possible, document:
- What was tested
- What remains uncertain
- Who is accepting the residual risk
Recovery should be part of the review decision, not something the team invents after the deployment fails.
Use Automation to Remove Repetitive Work, Not Accountability
A Terraform CI/CD pipeline should reduce repetitive review work and leave humans with a smaller set of meaningful decisions.
Automate tasks such as:
- Formatting checks
- Configuration validation
- Applicable linting
- Security rules
- Plan generation
Keep the result of each gate visible.
A single green aggregate badge is not useful if an underlying security check failed open or never executed.
Terraform validate checks configuration consistency, but it does not validate remote services such as provider APIs.
Passing validation should therefore not be presented as evidence that production behavior will be safe.
Use Plan-Based Policies Carefully
Automated plan checks can help identify issues such as:
- Prohibited exposure
- Deletion of protected resources
- Missing ownership information
Start with rules the team can explain and maintain.
A policy that produces constant irrelevant failures eventually becomes an exception-processing system rather than a useful control.
For every blocking policy, document:
- Owner
- Intent
- Known limitations
- Test cases
Include examples that must pass and examples that must fail.
When a policy changes, review that modification with the same care as the infrastructure it governs.
A small set of trusted rules is more useful than a large collection nobody understands.
Review exceptions periodically.
If the same legitimate exception appears repeatedly, use that evidence to improve the rule or the infrastructure design.
Protect the Review Pipeline Itself
Plan generation is not permission to run arbitrary contributed code with production credentials.
Separate untrusted pull-request checks from trusted infrastructure jobs.
Control:
- Token permissions
- Cloud permissions
- Which revisions can reach privileged runners
- What approvals are required before production access is granted
GitHub’s secure-use guidance describes risks associated with untrusted workflow inputs and recommends limiting credentials and protecting workflow execution.
Apply those principles to the CI platform your team actually uses.
Treat Terraform Plans as Sensitive Artifacts
Saved plans and rendered plan output may contain sensitive information.
The Terraform show documentation warns that JSON output can expose sensitive values in plain text.
Restrict:
- Artifact access
- Artifact retention
- Where plan output can be published
- Which external tools receive that data
Do not publish complete plan JSON in a public comment.
Do not paste it into an unapproved AI tool.
And do not assume that a masked terminal view means the underlying artifact is safe.
Worked Example: A Routine Capacity Change That Needs Two Decisions
Consider a fictional team increasing background-worker capacity before a customer launch.
The pull request upgrades a shared worker module and changes its size setting.
The author expects only additional processing capacity.
But the proposed Terraform plan also:
- Replaces a worker resource
- Broadens a role’s access to a storage location
The exact behavior depends on the provider and module.
This example illustrates review reasoning, not a universal Terraform output.
| Observation | Review Question | Required Decision |
|---|---|---|
| Worker replacement | Can processing continue while the old worker is retired? | Service owner approves the transition and drain procedure. |
| Broader storage access | Is the new permission necessary for this workload? | Platform or security owner accepts a narrowed permission or rejects the expansion. |
| Shared module upgrade | Do other module consumers change in this deployment? | Author demonstrates the deployment scope and separates unrelated work. |
| Changed capacity | What indicates useful capacity rather than idle cost? | Team defines throughput, backlog, and cost observations. |
The reviewer should not approve the change simply because the desired size appears in the configuration.
The review should also not stall with a vague:
“Needs more testing.”
A useful review response identifies the missing evidence.
In this example, that means:
- Demonstrating the worker transition in staging
- Narrowing the permission to the intended resource
- Attaching the final plan for the revised change
The author then separates the unexpected permission change from the capacity work.
The updated review record explains:
- What will happen to in-flight jobs
- Who will monitor the queue
- What condition will stop the rollout
If the module cannot support the required transition, that becomes an explicit design decision before launch.
Increasing instance size does not solve an unsafe replacement sequence.
Close the Loop After Deployment
After execution, record:
- Whether the expected behavior occurred
- Whether unplanned intervention was required
- Which part of the review could have identified the issue earlier
That feedback turns a one-time approval into an opportunity to improve the review process.
Keep Approval Tied to the Change That Actually Runs
Review ownership should continue from pull request through production.
A source-code approval and a deployment approval answer different questions.
Source-Code Approval
This answers:
Is the proposed configuration acceptable?
Deployment Approval
This answers:
Should the final change run now, against this environment, under these operational conditions?
For low-risk changes, a team may choose to combine those decisions.
For consequential changes, make both visible.
Otherwise, a pull request may be approved on Tuesday and applied on Friday after the surrounding system has changed.
Protect the Link Between Plan, Approval, and Deployment
Maintain an auditable relationship between:
- Final plan
- Source revision
- Approval
- Deployment
Use protected artifacts and restrict who can:
- Replace the approved plan
- Modify the workflow that consumes it
- Change the revision reaching production
If a saved plan is applied, remember that Terraform executes that plan without another interactive confirmation.
The approval control therefore needs to exist in the surrounding workflow.
It is not supplied by another Terraform prompt.
State Locking Is Not Review Approval
Review coordination and state protection solve different problems.
Terraform state locking, where supported by the backend, helps protect state-changing operations from concurrent writes.
It does not:
- Reserve an environment for a reviewer
- Confirm that application dependencies remain unchanged
- Prove that the approved deployment conditions still exist
Do not treat state locking as a substitute for operational approval.
Define When Approval Expires
Approval should not remain valid indefinitely.
A reassessment may be required when:
- The source revision changes
- The final plan changes
- Material drift appears
- The deployment window is missed
The response should be proportional to the risk.
The important point is to avoid treating an old approval as a transferable permission slip.
Assign Reviewers by the Decisions They Own
A process that depends on one senior infrastructure engineer will eventually collide with:
- Vacations
- Incidents
- Competing delivery work
- Review backlogs
The solution is not to let everyone approve everything.
Instead, build a small ownership map.
Author
The author:
- Explains intent
- Prepares the evidence
- Identifies expected impact
Service Owner
The service owner evaluates:
- Application impact
- Acceptance criteria
- Operational behavior
Platform Reviewer
The platform reviewer evaluates:
- Infrastructure interactions
- Shared components
- Platform-specific consequences
Security Reviewer
The security reviewer handles defined changes that cross trust boundaries.
This role should not be required for every routine infrastructure modification.
Deployment Owner
The deployment owner confirms:
- Execution conditions
- Timing
- Follow-through after deployment
In a small organization, one person may hold several of these roles.
Document that reality instead of creating artificial separation that the team cannot sustain.
Expand Review Responsibility Through Demonstrated Judgment
For common low-risk changes, document examples that show the expected review standard.
Pair less experienced reviewers with experienced owners when appropriate.
Then expand review responsibility based on demonstrated judgment rather than job title alone.
The objective is to reduce unnecessary dependency on one person without lowering the review standard.
Measure Where Reviews Actually Wait
Do not treat every delay as a reviewer-capacity problem.
Track whether infrastructure changes are waiting for:
- Missing evidence from the author
- A qualified reviewer
- A planned release window
Those queues require different solutions.
Useful measures can include:
- Review waiting time by risk class
- Requests returned for missing context
- Unplanned recovery work
Do not reward lower review times if they are achieved by skipping meaningful controls.
Decide Whether the Gap Needs a Project or Permanent Ownership
Some infrastructure bottlenecks are implementation problems.
For example, the team may lack:
- Consistent plan artifacts
- Policy tests
- Approval gates
- Documented handoff
- Protected deployment controls
Other bottlenecks are ongoing ownership problems.
For example:
- Nobody has time to maintain modules
- Nobody owns infrastructure decisions
- Review responsibility is permanently understaffed
Do not add capacity until that distinction is clear.
When a Bounded DevOps Project Makes Sense
A bounded improvement project can focus on a defined workflow outcome.
That may include:
- Defining the target review workflow
- Implementing agreed controls
- Testing those controls on representative changes
- Handing over maintenance to the internal team
Useful acceptance criteria can include:
- A reproducible review package
- Tested blocking rules
- Protected approval paths
- An internal maintainer who can explain the system
This makes the improvement measurable.
The project has a defined implementation outcome rather than becoming open-ended infrastructure ownership.
When Continuing Ownership Is the Real Need
If the underlying problem is sustained responsibility, assess whether an embedded specialist or a permanent hire is a better fit.
TechAID’s guide to choosing a nearshore SRE engagement model explores the broader decision between different sourcing models for ongoing reliability and infrastructure ownership.
For a defined delivery gap, a nearshore project outsourcing model can be used to structure a scoped DevOps or CI/CD implementation with clear acceptance criteria and handoff.
Keep technical authority and acceptance responsibilities explicit.
An external partner can implement agreed improvements.
Your organization still needs to decide:
- Which risks it accepts
- Who approves production changes
- Who owns the workflow after handoff
Build a Review Process That Produces Explainable Decisions
The goal of Terraform code review is not simply to make the pipeline green.
A strong review process should allow someone to explain:
- What will change
- Why the change is acceptable
- How execution is controlled
- What the team will observe afterward
- What happens if reality differs from the plan
That is a stronger standard than a green checkmark and a more useful foundation for increasing infrastructure delivery capacity.
If your team has a defined gap in Terraform review, approval controls, or CI/CD implementation, talk to TechAID about a scoped project.
