An extra 30% off every course until Sunday, October 11.Choose your certification →

Copyright (c) 2026 MindMesh Academy. All rights reserved. This content is proprietary and may not be reproduced or distributed without permission.

3.1.1.1. Multi-AZ and Multi-Region Deployments (Compute, Data Layer)

3.1.1.1. Multi-AZ and Multi-Region Deployments (Compute, Data Layer)

Multi-AZ protects against single datacenter failures. Multi-Region protects against entire region outages and reduces latency for global users. The exam expects you to know which services support each pattern and the trade-offs involved.

Compute layer HA patterns:
PatternMechanismFailover Time
EC2 + ASG across AZsASG replaces failed instances in healthy AZsMinutes
ECS/EKS across AZsScheduler places tasks in available AZsSeconds
LambdaAutomatically multi-AZ within a regionSeamless
EC2 Multi-RegionSeparate ASGs per region + Route 53 failoverMinutes
Data layer HA patterns:
ServiceMulti-AZMulti-Region
RDSSynchronous standby, automatic failover (1-2 min)Read replicas, manual promotion
AuroraUp to 15 read replicas across AZs, auto failover (30s)Global Database, <1s replication lag
DynamoDBAutomatic across 3 AZsGlobal Tables (active-active, seconds lag)
ElastiCache RedisMulti-AZ with auto-failoverGlobal Datastore (cross-region read replicas)
S3Automatic across ≥3 AZsCross-Region Replication (CRR)

Exam Trap: RDS Multi-AZ is not a read scaling solution — the standby is for failover only and cannot serve read traffic. For read scaling, use read replicas. Aurora is different: its read replicas do serve read traffic AND participate in failover. If the exam asks about both read scaling and HA, Aurora read replicas are the answer.

Designing Multi-AZ capacity (N+1 AZ sizing): size each AZ so that the AZs left after one fails can still carry peak load plus any required headroom. Example: peak needs 4 instances across 3 AZs → 2 per AZ leaves exactly 4 after an AZ loss (no spare); if one spare is also required, you need 3 per AZ (6 survive). Pair this with load balancer health checks that route around a failed AZ and stateless instances (state in ElastiCache, DynamoDB or S3). A cluster placement group packs instances into a single AZ — the opposite of this goal.

High availability vs. fault tolerance: HA minimizes downtime through redundancy and automatic failover, accepting a brief interruption (e.g., an RDS Multi-AZ failover). Fault tolerance eliminates the interruption: redundant components are already serving, so a failure is absorbed with no user-visible outage. Fault tolerance costs more.

See how it connects
Alvin Varughese
Written byAlvin Varughese
Founder•20 professional certifications