Infrastructure as Code

#6 7 pages

Infrastructure as Code (IaC)

Managing infrastructure through versioned, reviewable code — the same discipline applied to application code.

0/0 checks

Why IaC

graph LR
    classDef blue fill:#3498db,stroke:#2980b9,color:#fff
    classDef green fill:#2ecc71,stroke:#27ae60,color:#fff
    classDef red fill:#e74c3c,stroke:#c0392b,color:#fff
    classDef orange fill:#e67e22,stroke:#d35400,color:#fff
    classDef purple fill:#9b59b6,stroke:#8e44ad,color:#fff
    classDef teal fill:#1abc9c,stroke:#16a085,color:#fff
    classDef dark fill:#2c3e50,stroke:#1a252f,color:#fff
    classDef yellow fill:#f39c12,stroke:#d68910,color:#000
    classDef k8s fill:#326ce5,stroke:#254ea8,color:#fff
    classDef aws fill:#ff9900,stroke:#cc7a00,color:#000
    subgraph WithoutIaC["Without IaC (manual)"]
        CONSOLE["Click-ops in AWS Console<br/>each change is a one-off, undocumented action"]:::blue --> SNOWFLAKE["Snowflake servers<br/>unique, undocumented, nobody remembers why"]:::dark
        SNOWFLAKE --> DRIFT["Config drift<br/>prod != staging != dev"]:::red
        DRIFT --> FEAR["Fear of change<br/>nobody knows what breaks"]:::orange
    end

    subgraph WithIaC["With IaC"]
        CODE["Infra config in .tf / .yaml files<br/>checked into version control"]:::blue --> PR["Code review + PR approval<br/>same discipline as app code"]:::teal
        PR --> PIPELINE["CI pipeline<br/>plan, validate, apply"]:::blue
        PIPELINE --> VERSIONED["Versioned, auditable, reproducible<br/>git blame tells you who and why"]:::teal
        VERSIONED --> IDEMPOTENT["Idempotent<br/>apply twice = same result"]:::green
    end

    DRIFT -.->|"the exact problem IaC removes"| CODE

Someone claims the main benefit of IaC over click-ops is that it's faster to spin up infrastructure. Based on this section, what's the more important benefit?


Declarative vs Imperative

graph TD
    classDef blue fill:#3498db,stroke:#2980b9,color:#fff
    classDef green fill:#2ecc71,stroke:#27ae60,color:#fff
    classDef red fill:#e74c3c,stroke:#c0392b,color:#fff
    classDef orange fill:#e67e22,stroke:#d35400,color:#fff
    classDef purple fill:#9b59b6,stroke:#8e44ad,color:#fff
    classDef teal fill:#1abc9c,stroke:#16a085,color:#fff
    classDef dark fill:#2c3e50,stroke:#1a252f,color:#fff
    classDef yellow fill:#f39c12,stroke:#d68910,color:#000
    classDef k8s fill:#326ce5,stroke:#254ea8,color:#fff
    classDef aws fill:#ff9900,stroke:#cc7a00,color:#000
    subgraph Declarative["Declarative: Terraform, CloudFormation, Pulumi"]
        D1["Describe desired end state<br/>e.g. 3 EC2 instances, one SG"]:::blue --> D2["Tool figures out HOW<br/>diffs desired vs actual"]:::teal
        D2 --> D3["Tool tracks state internally<br/>knows what already exists"]:::purple
    end

    subgraph Imperative["Imperative: Ansible, shell scripts, AWS CLI"]
        I1["Describe the steps in order<br/>e.g. create instance, then attach EIP"]:::orange --> I2["Order matters<br/>idempotency is your problem"]:::orange
        I2 --> I3["Running twice may break things<br/>e.g. duplicate resources"]:::red
    end
Declarative Imperative
You specify What you want How to do it
Idempotency Built-in Your responsibility
Diff/preview Native (terraform plan, cdk diff) Manual
Best for Static infra (VPCs, DBs, LBs) Config management, bootstrapping

You run the same imperative shell script twice against the same environment. Why might that break something, when running `terraform apply` twice never does?


Tools Comparison

graph TD
    classDef blue fill:#3498db,stroke:#2980b9,color:#fff
    classDef green fill:#2ecc71,stroke:#27ae60,color:#fff
    classDef red fill:#e74c3c,stroke:#c0392b,color:#fff
    classDef orange fill:#e67e22,stroke:#d35400,color:#fff
    classDef purple fill:#9b59b6,stroke:#8e44ad,color:#fff
    classDef teal fill:#1abc9c,stroke:#16a085,color:#fff
    classDef dark fill:#2c3e50,stroke:#1a252f,color:#fff
    classDef yellow fill:#f39c12,stroke:#d68910,color:#000
    classDef k8s fill:#326ce5,stroke:#254ea8,color:#fff
    classDef aws fill:#ff9900,stroke:#cc7a00,color:#000
    subgraph TF["Terraform / OpenTofu"]
        TF1["HCL language<br/>1000+ providers"]:::teal
        TF2["State file<br/>S3 + DynamoDB lock"]:::yellow
        TF3["Multi-cloud<br/>large community"]:::teal
    end
    subgraph CFN["AWS CloudFormation"]
        CFN1["YAML/JSON<br/>AWS-only"]:::aws
        CFN2["State managed by AWS natively<br/>no backend to configure"]:::orange
        CFN3["Change sets, StackSets<br/>deep AWS integration"]:::aws
    end
    subgraph CDK["AWS CDK"]
        CDK1["TS/Python/Go/Java<br/>synthesizes to CFN templates"]:::aws
        CDK2["Programmatic<br/>loops, conditionals, OOP abstractions"]:::purple
    end
    subgraph PUL["Pulumi"]
        PUL1["Any language<br/>Go, TypeScript, Python"]:::purple
        PUL2["Multi-cloud<br/>programmatic with declarative semantics"]:::purple
    end
Terraform CloudFormation CDK Pulumi
Language HCL YAML/JSON TS/Python/Go/Java Any language
State S3+DynamoDB (recommended) AWS-managed AWS-managed (via CFN) Pulumi Cloud / self-hosted
Cloud Multi-cloud AWS only AWS only Multi-cloud
Preview terraform plan Change sets cdk diff pulumi preview
Best for Multi-cloud, OSS ecosystem AWS-native, managed service preferred Devs preferring real code over YAML Teams wanting full language features

The two most-compared tools in practice are Terraform and CloudFormation — one is multi-cloud with a state file you own, the other is AWS-native with state you never have to think about:

You own the state file and the backend. Provision an S3 bucket plus a DynamoDB lock table yourself (or use Terraform Cloud) before day one — that's what stores current world-state and arbitrates concurrent applies. In exchange you get HCL, 1000+ providers, and a config that works the same way across every cloud, plus terraform plan as a real diff before every apply.
AWS owns state for you. No backend to design, no lock table to provision, no separate mechanism to get right — CloudFormation manages it natively. The tradeoff is lock-in: it only understands AWS resources, written in YAML/JSON, and its preview is a Change Set rather than a raw diff. What it buys you is the deepest native AWS integration, including StackSets for rolling the same stack out across many AWS accounts at once.

A team is deciding between Terraform and CloudFormation and wants to skip ever provisioning a state backend themselves. Which tool gives them that, and what do they give up for it?


State Management

State tracks what the tool thinks currently exists. Without it, the tool can't compute what to create, update, or destroy.

sequenceDiagram
    participant Dev as terraform apply
    participant State as State File (S3)
    participant Real as Real AWS

    Dev->>State: Read current state
    Dev->>Real: Refresh, query real infra
    Note over Dev: Diff, desired (code) vs actual (real)
    rect rgb(40, 60, 45)
    Note over Dev: Execution plan, + create, ~ update, - destroy
    Dev->>Real: Execute API calls
    Dev->>State: Write updated state
    end
1. Read current state. The tool loads what it last recorded about the world — not necessarily what's actually there right now.
2. Refresh. It queries real AWS directly, because the recorded state can be stale — someone may have changed something outside the tool since the last run.
3. Diff. Desired (what the code says) is compared against actual (what refresh just found), not just against the old state file.
4. Build the execution plan. Every difference becomes a line: + create, ~ update, - destroy.
5. Execute and record. API calls run against real AWS, and only once they succeed does the tool write the new state back — so state always reflects the last known-good apply.

Remote state with locking:

terraform {
  backend "s3" {
    bucket         = "my-tf-state"
    key            = "prod/terraform.tfstate"
    region         = "us-east-1"
    encrypt        = true
    dynamodb_table = "tf-state-lock"  # prevents concurrent applies corrupting state
  }
}

Why the lock matters: an S3 bucket alone stores state, but it doesn't stop two people from reading, planning, and writing to it at the same time. The DynamoDB table is what actually serializes concurrent applies — one gets the lock and proceeds, the other waits instead of racing to write conflicting state:

sequenceDiagram
    participant A as terraform apply (Engineer A)
    participant Lock as DynamoDB Lock Table
    participant B as terraform apply (Engineer B)
    participant S as State File (S3)

    A->>Lock: Acquire lock
    Lock-->>A: Lock granted
    B->>Lock: Acquire lock
    Lock-->>B: Lock already held, wait
    A->>S: Read, plan, apply, write updated state
    A->>Lock: Release lock
    Lock-->>B: Lock now available
    B->>S: Read updated state, then plan, apply, write
    B->>Lock: Release lock
1. Engineer A applies first. A's terraform apply grabs the DynamoDB lock before doing anything else.
2. Engineer B is blocked, not racing. B's apply, started seconds later, tries for the same lock and simply waits — it never gets a chance to read a state file that A is mid-write on.
3. A finishes and releases. A's plan executes against real AWS, the new state is written to S3, and only then is the lock released.
4. B proceeds against fresh state. B now reads the state file A just updated — including A's changes — before computing its own plan, so B is never planning against stale data.

The S3 backend is configured, but `dynamodb_table` is left out. Two engineers happen to run `terraform apply` on the same stack at the same moment. What actually goes wrong?


Drift

Drift = actual infra differs from what IaC thinks it is. Caused by manual console/CLI changes outside the IaC workflow.

graph LR
    classDef blue fill:#3498db,stroke:#2980b9,color:#fff
    classDef green fill:#2ecc71,stroke:#27ae60,color:#fff
    classDef red fill:#e74c3c,stroke:#c0392b,color:#fff
    classDef orange fill:#e67e22,stroke:#d35400,color:#fff
    classDef purple fill:#9b59b6,stroke:#8e44ad,color:#fff
    classDef teal fill:#1abc9c,stroke:#16a085,color:#fff
    classDef dark fill:#2c3e50,stroke:#1a252f,color:#fff
    classDef yellow fill:#f39c12,stroke:#d68910,color:#000
    classDef k8s fill:#326ce5,stroke:#254ea8,color:#fff
    classDef aws fill:#ff9900,stroke:#cc7a00,color:#000
    subgraph Origin["Where drift comes from"]
        STATE["IaC state thinks<br/>SG allows port 443 only"]:::green -. "drift" .-> REAL["Real AWS has<br/>engineer added port 22 via console"]:::teal
    end

    subgraph Response["Detect, fix, prevent"]
        DETECT["Detect<br/>terraform plan or CFN drift detection<br/>shows an unexpected change"]:::red
        FIX["Fix<br/>import the manual change into state,<br/>OR let IaC correct it on next apply"]:::orange
        PREVENT["Prevent<br/>SCP denies console write access in prod<br/>all changes go through the IaC pipeline"]:::blue
        DETECT --> FIX --> PREVENT
    end

    REAL -.->|"surfaces as"| DETECT
1. Manual change happens. An engineer opens port 22 via the AWS console — outside the IaC workflow entirely, no PR, no plan.
2. Drift surfaces. The next terraform plan (or a CloudFormation drift-detection run) shows an unexpected change that no commit produced.
3. Decide. Either import the manual change into state so IaC adopts it as the new source of truth, or take no action and let the next apply correct it.
4. Prevent. An SCP denies manual console-write access in production, so this class of drift can't happen again — every change has to go through the pipeline.

Prevention is the goal: SCPs (Service Control Policies) that deny all manual writes in production accounts. No human IAM write permissions in prod — changes must go through the pipeline.

terraform plan shows a drifted security group rule (someone opened port 22 by hand). Nobody imports it and nobody edits the .tf code. What happens on the next terraform apply?


Environment Isolation

graph TD
    classDef blue fill:#3498db,stroke:#2980b9,color:#fff
    classDef green fill:#2ecc71,stroke:#27ae60,color:#fff
    classDef red fill:#e74c3c,stroke:#c0392b,color:#fff
    classDef orange fill:#e67e22,stroke:#d35400,color:#fff
    classDef purple fill:#9b59b6,stroke:#8e44ad,color:#fff
    classDef teal fill:#1abc9c,stroke:#16a085,color:#fff
    classDef dark fill:#2c3e50,stroke:#1a252f,color:#fff
    classDef yellow fill:#f39c12,stroke:#d68910,color:#000
    classDef k8s fill:#326ce5,stroke:#254ea8,color:#fff
    classDef aws fill:#ff9900,stroke:#cc7a00,color:#000
    subgraph Dir["Option 1: Directory per env (recommended)"]
        MOD["modules/<br/>shared code, reused by every env"]:::blue --> DEV["envs/dev/<br/>own state file"]:::teal
        MOD --> STG["envs/staging/<br/>own state file"]:::teal
        MOD --> PRD["envs/prod/<br/>own state file, own AWS account"]:::green
    end

    subgraph WS["Option 2: Workspaces"]
        SAME["Same config + backend<br/>different state per workspace<br/>terraform workspace new staging"]:::purple
        LIM["Limitation<br/>same codebase<br/>risky for large env differences"]:::orange
        SAME --> LIM
    end

    subgraph TG["Option 3: Terragrunt"]
        DRY["DRY wrapper<br/>generates backend config per env"]:::teal
        DEPS["Dependency graph<br/>applies modules in the right order"]:::teal
        BEST["Best for<br/>5+ envs, many modules"]:::yellow
        DRY --> DEPS --> BEST
    end
Option State isolation DRY Separate AWS accounts Best for
Directories Full Partial Yes Small-medium teams
Workspaces Yes (per workspace) Full No Similar envs, simple configs
Terragrunt Full Full Yes Large teams, many environments

Recommendation: Directories + shared modules. Terragrunt when you hit 5+ environments.

Workspaces give each environment its own state, so why does this section call them "risky for large env differences" compared to directories?


Secrets in IaC

graph LR
    classDef blue fill:#3498db,stroke:#2980b9,color:#fff
    classDef green fill:#2ecc71,stroke:#27ae60,color:#fff
    classDef red fill:#e74c3c,stroke:#c0392b,color:#fff
    classDef orange fill:#e67e22,stroke:#d35400,color:#fff
    classDef purple fill:#9b59b6,stroke:#8e44ad,color:#fff
    classDef teal fill:#1abc9c,stroke:#16a085,color:#fff
    classDef dark fill:#2c3e50,stroke:#1a252f,color:#fff
    classDef yellow fill:#f39c12,stroke:#d68910,color:#000
    classDef k8s fill:#326ce5,stroke:#254ea8,color:#fff
    classDef aws fill:#ff9900,stroke:#cc7a00,color:#000
    BAD["Hardcoded in .tf or tfvars<br/>committed to git history forever"]:::red -->|"never do this"| GOOD
    subgraph GOOD["Secure patterns"]
        ENV["TF_VAR_x env var<br/>set by CI from vault"]:::blue
        SENS["sensitive = true<br/>hides value from plan/apply output"]:::teal
        DATASRC["data source<br/>pulls from SSM/Vault at apply time"]:::yellow
        SOPS["SOPS + KMS<br/>encrypted tfvars, safe to commit"]:::purple
    end
variable "db_password" {
  type      = string
  sensitive = true   # hidden from CLI output and logs
}

# Pull from AWS SSM at apply time — never stored in config
data "aws_ssm_parameter" "db_password" {
  name            = "/prod/db/password"
  with_decryption = true
}
# .gitignore
*.tfvars
terraform.tfstate
terraform.tfstate.backup
.terraform/

Of the four secure patterns shown, which one is the only one that actually lets you commit a tfvars file to git — the exact thing the "never do this" arrow warns against for plain tfvars?


count vs for_each

graph TD
    classDef blue fill:#3498db,stroke:#2980b9,color:#fff
    classDef green fill:#2ecc71,stroke:#27ae60,color:#fff
    classDef red fill:#e74c3c,stroke:#c0392b,color:#fff
    classDef orange fill:#e67e22,stroke:#d35400,color:#fff
    classDef purple fill:#9b59b6,stroke:#8e44ad,color:#fff
    classDef teal fill:#1abc9c,stroke:#16a085,color:#fff
    classDef dark fill:#2c3e50,stroke:#1a252f,color:#fff
    classDef yellow fill:#f39c12,stroke:#d68910,color:#000
    classDef k8s fill:#326ce5,stroke:#254ea8,color:#fff
    classDef aws fill:#ff9900,stroke:#cc7a00,color:#000
    subgraph CountDanger["count: index-based — dangerous for removals"]
        C1["users = [alice, bob, carol]"]:::blue --> C2["index 0 = alice<br/>index 1 = bob<br/>index 2 = carol"]:::blue
        C2 --> C3["Remove bob<br/>index 1 shifts to carol<br/>destroy + recreate carol AND alice"]:::red
    end

    subgraph ForEachSafe["for_each: key-based — safe for removals"]
        F1["users = {alice, bob, carol}"]:::teal --> F2["key alice<br/>key bob<br/>key carol"]:::teal
        F2 --> F3["Remove bob<br/>only bob's resource destroyed<br/>alice and carol untouched"]:::green
    end

Rule: Always for_each for dynamic resources. Only count for simple enable/disable: count = var.enable_monitoring ? 1 : 0.

A count-based list has [alice, bob, carol] at indexes 0, 1, 2. You remove bob. Which resources does Terraform actually destroy and recreate — just bob's?


Lifecycle Rules

resource "aws_instance" "web" {
  lifecycle {
    create_before_destroy = true   # new resource created BEFORE old destroyed (zero downtime)
    prevent_destroy       = true   # blocks destroy — protects prod DBs, S3 buckets
    ignore_changes        = [tags] # ignore attrs managed externally (AWS auto-tags)
    replace_triggered_by  = [aws_security_group.web.id]  # force replace when dependency changes
  }
}
Rule Use for
create_before_destroy EC2, ECS services — must have no downtime during replacement
prevent_destroy Production databases, S3 buckets — protect from accidental destroy
ignore_changes Tags/attrs managed by AWS or external tools
replace_triggered_by Force replacement when a dependency changes but Terraform wouldn't detect it

A production RDS resource has `prevent_destroy = true`. Someone runs `terraform destroy` against that stack. What happens?


Ansible

Ansible handles configuration management — what goes inside the infrastructure that Terraform provisions.

File Topics Level
ansible/README.md Architecture, how Ansible works, SSH internals, Mermaid diagrams, ansible.cfg SDE-1
ansible/core-concepts.md Inventory, playbooks, modules, tasks, handlers, variables, facts, Jinja2 templates SDE-1
ansible/cloud-integration.md AWS SSM + SSH, GCP OS Login + IAP, dynamic inventory, cloud modules SDE-1/2
ansible/advanced.md Roles, collections, Vault, AWX/Tower, performance tuning, Molecule testing SDE-2

Terraform vs Ansible in one line: Terraform creates the VM. Ansible configures what's inside it.

Read order: ansible/README.md → core-concepts → cloud-integration → advanced

Pages in this section