30% off every course until Sunday, October 11. Our biggest update yet, and we'd like you to try it. Applied automatically at checkout.

Choose your certification
Copyright (c) 2026 MindMesh Academy. All rights reserved. This content is proprietary and may not be reproduced or distributed without permission.

1.2.3. πŸ’‘ First Principle: Reliability Pillar

First Principle: Designing systems to recover from infrastructure or service outages, dynamically acquire computing resources to meet demand, and mitigate disruptions (e.g., misconfigurations or transient network issues) ensures continuous availability and functionality.

Scenario: To ensure a critical application remains operational even if an entire Availability Zone ("AZ") experiences an outage, an architect designs the application to deploy its components across multiple "AZs". For the database, they use "Amazon RDS Multi-AZ" for synchronous replication and "Amazon S3 Cross-Region Replication" for data backup to a different region.

The Reliability pillar of the AWS Well-Architected Framework is about ensuring your workload performs its intended function correctly and consistently when it's expected to. For a Solutions Architect, this involves designing systems that are fault-tolerant, highly available, and capable of self-healing.

Key Design Considerations:
  • Foundations: Designing for high availability within a Region ("Multi-AZ") and across Regions ("Multi-Region") using redundant resources.
  • Change Management: Implementing automated, repeatable changes with minimal impact, and planning for rollbacks.
  • Failure Management: Designing for graceful recovery from failures, anticipating problems, and implementing self-healing mechanisms. Scale horizontally across multiple AZs (an Auto Scaling group behind a load balancer replaces failed instances automatically) rather than relying on one bigger instance or one AZ, either of which remains a single point of failure. Keep compute stateless by holding session state in an external store ("DynamoDB", "ElastiCache") so any instance can be replaced. Route messages that repeatedly fail processing to a dead-letter queue, and retry transient errors a limited number of times with exponential backoff and jitter; immediate, unbounded retries create retry storms.
  • Capacity Management: Dynamically scaling resources to meet fluctuating demand, preventing overload.
    • Service quotas: Most quotas are per account and per Region, and a request beyond a quota fails immediately. Some quotas grow gradually with usage, and opt-in Service Quotas Automatic Management can notify at 80%/95% and file increase requests, but approval is neither guaranteed nor instant, so check limits in "Service Quotas" and get increases approved before a known launch. For continuous monitoring, alarm in "CloudWatch" on usage against the quota (Service Quotas publishes AWS/Usage metrics; some services have their own, such as Lambda ConcurrentExecutions); Trusted Advisor's service-limit checks are periodic, not a real-time alarm. EC2 On-Demand quotas are counted in vCPUs per instance category (all Standard families, A, C, D, H, I, M, R, T and Z, share one quota), and moving a workload to another Region to dodge a limit is a workaround, not quota management.
    • Quota vs. capacity: A higher quota lifts only the account limit; it does not guarantee physical capacity. An "On-Demand Capacity Reservation" holds capacity for a specific instance type in a specific AZ, is billed at On-Demand rates whether or not instances run, and counts against the On-Demand quota. Regional Reserved Instances and Savings Plans are billing discounts only (no capacity, no extra quota headroom); zonal Reserved Instances do reserve capacity in one AZ but require a 1- or 3-year term. Spot capacity is never guaranteed.
Practical Implementation: Creating a Multi-AZ Auto Scaling Group via AWS CLI
# This command creates an Auto Scaling group that launches EC2 instances
# across two different Availability Zones (us-east-1a and us-east-1b)
# ensuring high availability.
aws autoscaling create-auto-scaling-group \
  --auto-scaling-group-name my-reliable-asg \
  --launch-template LaunchTemplateName=my-launch-template \
  --min-size 2 \
  --max-size 4 \
  --vpc-zone-identifier "subnet-0a1b2c3d,subnet-0e4f5g6h"
Visual: Reliability Pillar - Multi-AZ & Cross-Region Deployment

⚠️ Common Pitfall: Confusing high availability ("HA") with disaster recovery ("DR"). A "Multi-AZ" deployment provides HA within a region, but a "Multi-Region" strategy is required for DR against a regional failure.

Key Trade-Offs:
  • Reliability vs. Cost: Higher levels of reliability (e.g., "Multi-Region" active-active) require more infrastructure and data replication, which significantly increases cost and complexity.

Reflection Question: How do these combined architectural patterns (e.g., "RDS Multi-AZ", "S3 Cross-Region Replication") enhance application resilience against various types of failures, from component failures to full regional disasters, and what are the associated trade-offs in cost and complexity?

See how it connects
Alvin Varughese
Written byAlvin Varughese
Founderβ€’20 professional certifications