30% off every course until Sunday, October 11. Our biggest update yet, and we'd like you to try it. Applied automatically at checkout.

Choose your certification
Copyright (c) 2026 MindMesh Academy. All rights reserved. This content is proprietary and may not be reproduced or distributed without permission.

2.1.1.1. Designing for Scalability and Elasticity (Auto Scaling, Load Balancing)

2.1.1.1. Designing for Scalability and Elasticity (Auto Scaling, Load Balancing)

💡 First Principle: Architectures must dynamically adapt to fluctuating demand by automatically scaling resources proportionally, ensuring consistent performance during peaks and optimal cost-efficiency during lulls.

Scenario: A popular online gaming application experiences massive, unpredictable spikes in traffic during new game releases. The architect needs to ensure the application can scale rapidly to handle these surges without manual intervention and maintain performance.

Scalability and elasticity are critical for cloud-native applications. Scalability refers to a system's ability to handle increasing load, while elasticity is the ability to automatically grow or shrink resources based on demand.

  • "Amazon EC2 Auto Scaling": A service that dynamically adjusts "EC2 instance" capacity based on demand, using policies and health checks. Ensures performance during peaks and cost savings during lulls.
    • Key "Auto Scaling" Components:
      • Launch Templates/Configurations: Define how new instances are launched ("AMI", instance type, "security groups").
      • Scaling Policies: Define how to scale (e.g., "Target Tracking", "Simple/Step Scaling", "Scheduled Scaling").
      • Health Checks: Determine if an instance is healthy and should remain in service.
  • "Elastic Load Balancing (ELB)": A service that automatically distributes incoming application traffic across multiple targets, such as "EC2 instances", containers, and "Lambda functions", in multiple "Availability Zones". Continuously monitors target health and routes traffic only to healthy instances, ensuring high availability and fault tolerance. Supports "Application Load Balancer (ALB)", "Network Load Balancer (NLB)", and "Gateway Load Balancer (GLB)" for different protocol and routing needs.
    • Choosing the load balancer: "ALB" works at Layer 7 (HTTP/HTTPS, host- and path-based routing, Lambda targets). "NLB" works at Layer 4 (TCP/UDP/TLS), handles millions of requests per second at very low latency, and gives each AZ a static IP (or Elastic IP) that clients can allow-list. "GLB" is for deploying and scaling third-party virtual appliances such as firewalls transparently. The Classic Load Balancer is the legacy option.
  • Stateless tiers: Instances can only be added, removed or replaced freely if no user state lives on them. Externalize session state to "Amazon ElastiCache" or "Amazon DynamoDB": memory or local disk is lost with the instance, and sticky sessions (ALB session affinity) pin a user to one instance, so that user's session is still lost when the instance fails or scales in. Spread the "Auto Scaling group" across multiple AZs behind the load balancer so lost capacity is replaced automatically, and absorb spikes with managed services that scale on demand and decouple tiers (e.g., "SQS", "DynamoDB").
  • Rolling out a new AMI: A new launch template version (however it is deployed) affects only instances launched afterwards; running instances keep the old AMI until they are replaced. An "Instance Refresh" replaces them in place, with MinHealthyPercentage controlling how much capacity stays in service. "AWS Systems Manager Automation" runbooks can run across multiple accounts and Regions by assuming a role in each, which suits fleet-wide operational changes.
Practical Implementation: Creating a Target Tracking Scaling Policy
{
    "AutoScalingGroupName": "my-gaming-app-asg",
    "PolicyName": "cpu-utilization-scaling-policy",
    "PolicyType": "TargetTrackingScaling",
    "TargetTrackingConfiguration": {
        "PredefinedMetricSpecification": {
            "PredefinedMetricType": "ASGAverageCPUUtilization"
        },
        "TargetValue": 50.0
    }
}
Visual: Auto Scaling and ELB for Scalability and Elasticity

⚠️ Common Pitfall: Setting a cooldown period that is too short. This can lead to "flapping," where the "Auto Scaling group" rapidly scales in and out, causing instability and potentially higher costs.

Key Trade-Offs:
  • Aggressiveness vs. Stability: A very aggressive scaling policy (low target utilization, short cooldown) responds quickly to spikes but can be unstable. A more conservative policy is stable but may lag behind sudden traffic surges.

Reflection Question: How does combining "Amazon EC2 Auto Scaling" with "Elastic Load Balancing" address both the rapid scaling and high availability requirements for this gaming application, particularly when dealing with unpredictable traffic surges and ensuring fault tolerance across instances?

See how it connects
Alvin Varughese
Written byAlvin Varughese
Founder•20 professional certifications