Troubleshooting Terraform¶
Overview¶
Diagnose common Terraform failures with a fixed order: auth → init/providers → validate/plan → state/lock → drift → graph/performance — using import, state rm, and moved only as deliberate recovery tools.
Most “Terraform is broken” tickets are credentials, provider version skew, state lock contention, or drift from console edits. Separate configuration failures from state failures from provider API failures before changing random flags.
This is a core tutorial in Module 20 · Troubleshooting of the REBASH Academy Terraform for Cloud & DevOps Engineers series — written for Cloud, DevOps, Platform, and SRE engineers.
Prerequisites¶
Learning Objectives¶
By the end of this tutorial, you will be able to:
- Classify provider, auth, state, and graph errors
- Clear or wait on state locks safely
- Detect and reconcile drift
- Use
import,state rm, andmovedas recovery — not routine hacks - Spot dependency cycles and slow plans
Architecture¶
This topic’s control points and relationships are shown below.
Theory¶
What it is¶
Troubleshooting Terraform is a disciplined split between layers:
| Symptom | First checks |
|---|---|
| Auth / 401 / 403 | Credentials, OIDC, IAM, subscription/project |
| Provider download / schema | init, version pins, registry access |
| Validate / reference errors | Typos, wrong module outputs, type mismatches |
| State lock | Who holds the lock; CI overlap; crash mid-apply |
| Drift | Console changes; refresh/plan unexpected updates |
| Cycle errors | Mutual depends_on or references |
| Slow plans | Huge roots, remote data sources, chatty APIs |
| State corruption / object missing | Backend restore; state subcommands carefully |
Recovery tools: terraform import adopts an existing object; terraform state rm drops an address without destroying; moved renames addresses across refactors. All require plan review and backups (state pull) first.
Why it matters¶
Mean time to recovery depends on not confusing a bad variable with a bad backend. Platform on-call needs a fixed order juniors can follow under pressure. The same playbook feeds CI: catch validate/provider issues before apply, and treat force-unlock and state surgery as last resorts with audit notes.
How it works¶
Use this order every time:
- Auth — can the runner assume the role / use the subscription? Clock skew? Wrong account?
- Init —
terraform init(upgrade only when intentional); confirm provider versions. - Validate / plan — read the first error; fix config before touching state.
- Lock — if lock error, identify the holder; wait or
force-unlockonly when sure the other process is dead. - Drift — unexpected updates/deletes → decide re-apply desired state vs update config to match reality.
- Graph — cycle messages → break mutual dependencies; prefer data flow over blanket
depends_on. - State recovery — restore backend version; then consider
import/state rm/movedwith a saved plan. - Performance — split roots, cache providers, reduce refresh scope, avoid unbounded
for_eachover huge sets.
Separate questions: Does config parse? Does state match reality? Did the API reject the call?
Key concepts and comparisons¶
| Failure class | Looks like | Not fixed by |
|---|---|---|
| Auth | 403 / invalid token | Reformatting HCL |
| Config | Unknown resource / invalid ref | Force-unlock |
| Lock | state locked by … | Random state rm |
| Drift | perpetual plan changes | Ignoring and re-apply blind |
| Cycle | cycle: A, B | More depends_on both ways |
| API / provider | retryable 5xx or clear API error | Editing unrelated outputs |
| Recovery tool | Effect |
|---|---|
import | Bind real ID → address |
state rm | Forget address (resource may still exist) |
moved | Rename address without replace |
| Backend version restore | Roll back corrupted state object |
Common pitfalls¶
- Jumping to
force-unlockwhile another apply is healthy. state rmthen apply creating duplicates of critical databases.- Fixing production with console edits Terraform will fight forever.
- Assuming
-refresh=false“fixes” drift — it only hides it temporarily. - Expanding one root until plans take thirty minutes instead of splitting state.
Hands-on Lab¶
Create a workspace for this tutorial.
Focus: Reproduce and diagnose a common apply failure, then fix it
Step 1 – Create a configuration with a deliberate error¶
cat > main.tf <<'EOF'
terraform {
required_providers {
local = { source = "hashicorp/local", version = "~> 2.5" }
}
}
resource "local_file" "broken" {
filename = "${path.module}/out.txt"
content = local.missing
}
EOF
terraform init
terraform validate 2>&1 || true
Step 2 – Fix, apply, and use diagnostics habitually¶
cat > main.tf <<'EOF'
terraform {
required_providers {
local = { source = "hashicorp/local", version = "~> 2.5" }
}
}
locals {
message = "fixed"
}
resource "local_file" "broken" {
filename = "${path.module}/out.txt"
content = local.message
}
EOF
terraform validate
TF_LOG=INFO terraform apply -auto-approve 2>&1 | tail -n 20
cat out.txt
Final step – Cleanup note¶
terraform destroy -auto-approve
# Workspace kept for notes; remove with: rm -rf "$(pwd)" when finished
Validation¶
- Lab commands run under
~/rebash-terraform/module-20/ - You can explain each Theory section in your own words
- You used modern tooling where it applies to this topic
- You can describe one production failure mode for this topic
Code Walkthrough¶
Production practice for Troubleshooting Terraform always combines:
- Inspect before you change (status, plan, logs, dry-run)
- Prefer reversible, documented changes (Git, IaC, drop-ins, version pins)
- Capture evidence (command output, pipeline logs) for handovers
- Prefer current tools and APIs over legacy shortcuts
- Least privilege — escalate credentials only when required
Keep runbooks short enough to follow under pressure. Automate checks; keep humans for judgement.
Security Considerations¶
- Treat credentials and tokens for terraform as privileged — never commit them
- Prefer short-lived auth (OIDC, roles, SSO) over long-lived keys
- Validate blast radius before apply/deploy/delete operations
- Restrict who can approve production changes
- Collect audit logs; limit who can read sensitive traces
Common Mistakes¶
Jumping to force-unlock while another apply is healthy.
Validate assumptions against the Theory section and official docs before changing production.
state rm then apply creating duplicates of critical databases.
Lab shortcuts (open security groups, admin roles, skip approvals) must not ship unchanged.
Changing production without a rollback path
Always know how to revert (previous artefact, prior release, state rollback, DNS failback).
Best Practices¶
- Encode Troubleshooting Terraform changes as code and review them in pull requests
- Pin versions (images, modules, actions, provider plugins)
- Separate environments with clear promotion gates
- Alert on symptoms with runbooks attached
- Destroy lab resources; tag everything with owner and expiry where possible
Troubleshooting¶
| Symptom | Likely cause | Fix |
|---|---|---|
| Auth / permission denied | Wrong identity, policy, or scope | Check caller identity, roles, and least-privilege policies |
| Timeout / no route | Network, DNS, security group, or endpoint | Trace path, DNS, and allow-lists before retrying |
| Drift / unexpected plan | Manual change or wrong state/workspace | Reconcile desired vs actual; avoid click-ops on managed resources |
| Pipeline/job red | Flaky step, cache, or missing secret | Read failing step logs; bisect recent workflow/config changes |
| Cost spike | Idle load balancer, NAT, oversized compute | Inventory billable resources; stop/delete labs promptly |
Summary¶
You can test, secure, automate, multi-cloud, Kubernetes-bootstrap, and operate production Terraform — and troubleshoot failures by layer.
Interview Questions¶
- What are common causes of provider authentication errors?
- How do you interpret a resource that must be replaced?
- What does state inconsistency look like after a partial failure?
- How can TF_LOG help, and what must you avoid when sharing logs?
- When is
terraform refresh/ plan with refresh useful versus dangerous?
Sample answer — question 2
Replacement means Terraform cannot update in place—often from force-new arguments. Expect delete/create and plan for downtime or recreation side effects.
Sample answer — question 4
Debug logs may include secrets and tokens. Redact before sharing, and prefer targeted provider logs over dumping full environment variables.
Related Tutorials¶
- Course overview
- Course overview · Production Terraform Patterns · Format, Validate, and Terraform Test