3.1.1.2. Mitigating Single Points of Failure
3.1.1.2. Mitigating Single Points of Failure
💡 First Principle: A resilient architecture is achieved by systematically identifying and eliminating any single component whose failure would cause the entire system or a critical function to fail.
Scenario: An architect is reviewing an existing application that runs on a single "EC2 instance" and uses an "Amazon RDS" database deployed in a single "Availability Zone (AZ)". The architect identifies both as single points of failure.
A robust architecture actively seeks out and mitigates Single Points of Failure ("SPOFs") at every layer.
- Compute Layer:
"SPOF": Single"EC2 instance", single"ECS task", single"Lambda"function deployment.- Mitigation: Deploy across
"Multi-AZs"using"Auto Scaling Groups"("EC2"/"ECS"), or use managed services like"Lambda"/"Fargate"that inherently distribute resources. Use"Placement Groups"for specific isolation needs.
- Data Layer:
"SPOF": Single-"AZ"database (e.g.,"RDS"), unreplicated data.- Mitigation:
"RDS Multi-AZ","Aurora Global Database","DynamoDB Global Tables","S3 Cross-Region Replication","EBS Snapshots". Ensure data is always replicated and backed up.
- Networking Layer:
"SPOF": Single"Internet Gateway", single"NAT Gateway"in an"AZ", reliance on a single"Direct Connect"link.- Mitigation:
"IGWs"are inherently redundant. Deploy"NAT Gateways"in each"AZ"where private subnets need outbound internet access. For"Direct Connect", implement multiple links over diverse paths and multiple locations, or use"VPN"as backup.
- Application Layer:
"SPOF": Hardcoded endpoints, shared libraries on a single server, tightly coupled services.- Mitigation: Use load balancers (
"ALB","NLB") for traffic distribution, implement service discovery (e.g.,"Cloud Map"), design for loose coupling with message queues ("SQS") or event buses ("EventBridge"), use immutable infrastructure.
Visual: Mitigating Single Points of Failure (SPOFs)
Protecting a Component from Its Dependencies:
- Retries with exponential backoff and jitter: For transient errors and throttling (HTTP 429,
"DynamoDB"throttling), retry with growing, randomized delays and a bounded attempt count so clients do not retry in lockstep. Retries alone do not protect against a dependency that stays down; they add load to it. - Circuit breaker: After a failure-rate threshold the caller stops calling the dependency for a cool-down, fails fast or returns a fallback, then probes with a trial request. This keeps threads and connections from blocking on a slow or unavailable third-party or downstream service. Do not confuse it with Saga (distributed transactions), Fan-out (one-to-many delivery) or Strangler Fig (incremental migration).
- Decouple with a queue: Put
"SQS"between producer and consumer so work buffers when the downstream is slow, throttled or down (for example"API Gateway"->"SQS"->"Lambda"->"DynamoDB"), preventing data loss and keeping a non-critical call (email, analytics) off the user's request path. Longer timeouts, bigger instances or self-hosting the dependency do not remove the coupling. - Graceful degradation and fault isolation: Let a non-critical feature (recommendations) fail or fall back to cached/default content on its own so the core path (search, checkout) stays fast and available. Limiting how far one failure can spread is blast-radius reduction.
- AZ headroom (static stability): Capacity sized exactly to peak across two
"AZs"has no headroom after losing one. Pre-provision the spare capacity instead of relying on launching it during the failure: for example three"AZs"sized so any two carry peak (N+1) costs about 1.5x peak, versus 2x for doubling both of two"AZs"."EC2 Auto Recovery"only addresses a failed host in the same"AZ". - Stateful single instance: If session state can be externalized (
"ElastiCache","DynamoDB"), run an"Auto Scaling Group"across"AZs"behind an"ALB". If it must keep a local block device, an Auto Scaling group of one launches a new instance but not the old volume's data; use continuous block-level replication with automated failover ("AWS Elastic Disaster Recovery", which supports cross-"AZ"recovery for"EC2") to a standby in another"AZ", or"EFS"only if the application can use a file system.
⚠️ Common Pitfall: Overlooking dependencies on services in a single "AZ". For example, if all your "EC2 instances" rely on a single "NAT Gateway" in one "AZ" for internet access, the failure of that "AZ" will cut off internet connectivity for all instances, even those in other healthy "AZs".
Key Trade-Offs:
- Redundancy vs. Cost: Eliminating
"SPOFs"often involves duplicating resources or infrastructure (e.g.,"Multi-AZ"deployments), which increases the overall cost of the solution.
Reflection Question: How would you mitigate the single points of failure identified in this existing application (single "EC2 instance" and single-"AZ RDS" database) using AWS services like "Auto Scaling Groups" and "RDS Multi-AZ", to enhance the application's overall resilience without dramatically increasing cost?