2.2.1.2. Change Management Processes for IaC Platforms
2.2.1.2. Change Management Processes for IaC Platforms
Every IaC change carries risk. CloudFormation change sets let you preview exactly what will be created, modified, or replaced before committing.
Change sets are mandatory for production stacks. They show the action (Add/Modify/Remove), logical resource ID, physical resource ID, and whether the change requires replacement. Always review change sets — a property change that triggers replacement (e.g., changing an RDS Engine) destroys and recreates the resource.
Update behaviors vary per resource property. The CloudFormation documentation marks each property as:
- No Interruption: In-place update, no downtime
- Some Interruption: Brief disruption (e.g., EC2 instance restart)
- Replacement: Resource destroyed and recreated with new physical ID
Stack policies protect resources from accidental updates:
{
"Statement": [{
"Effect": "Deny",
"Action": "Update:Replace",
"Principal": "*",
"Resource": "LogicalResourceId/ProductionDatabase"
}]
}
Drift detection identifies resources modified outside of CloudFormation (manual console changes, CLI commands). Run drift detection periodically and before updates to catch unexpected state.
Resource import brings resources created outside CloudFormation (console, CLI) under stack management without recreating them: describe them in the template, give each one a DeletionPolicy (any value is accepted; Retain is the safe choice), and run an IMPORT change set that supplies each resource's identifier (e.g. an instance ID). Deploying a template that merely describes the resources creates new ones, and tagging a resource does not associate it with a stack.
UPDATE_ROLLBACK_FAILED means the rollback itself couldn't restore a resource — typically one changed or deleted outside CloudFormation. Fix the cause where you can (recreate a named resource with its original name and properties, remove a blocking dependency) and run continue-update-rollback. If the resource can't be restored (AWS-assigned physical IDs can't be reproduced), add --resources-to-skip <LogicalId>: CloudFormation marks it UPDATE_COMPLETE, finishes at UPDATE_ROLLBACK_COMPLETE, and every other resource stays managed. Reconcile the skipped resource with the template before the next update. A StackSet stack instance is an ordinary stack, so a failed instance is repaired the same way in its own account, then the StackSet operation is re-run.
Shift-left template validation in the build stage: aws cloudformation validate-template checks only template syntax and structure; cfn-lint checks templates against the resource specification (property names, types, allowed values) and best practices; cfn-nag and CloudFormation Guard flag security anti-patterns such as 0.0.0.0/0 ingress or unencrypted storage. None of them can prove that referenced resources exist (a real subnet ID, an available AMI) — those failures only appear at deploy time. For CDK, test the synthesized template: snapshot tests fail on any change (review the diff, then update the snapshot), while fine-grained assertions such as Template.hasResourceProperties pin one requirement (e.g., bucket encryption) and survive unrelated edits.
Exam Trap: CloudFormation rollback on update failure reverts the stack to its previous state — but if a resource was replaced, the original resource is already deleted. DeletionPolicy only applies when a resource is removed from the stack; the old physical resource left behind by a replacement is governed by UpdateReplacePolicy (Retain or Snapshot) — or block the replacement outright with a stack policy. Always set DeletionPolicy: Snapshot on RDS instances and Retain on S3 buckets in production.