30% off every course until Sunday, October 11. Our biggest update yet, and we'd like you to try it. Applied automatically at checkout.

Choose your certification
Copyright (c) 2026 MindMesh Academy. All rights reserved. This content is proprietary and may not be reproduced or distributed without permission.

3.2.1.1. Compute Cost Optimization (Right-sizing, RI, Savings Plans, Spot Instances)

3.2.1.1. Compute Cost Optimization (Right-sizing, RI, Savings Plans, Spot Instances)

💡 First Principle: Maximizing compute cost efficiency requires dynamically matching capacity to actual workload demand and leveraging flexible pricing models to pay only for what is needed.

Scenario: A company runs a batch processing application that can be interrupted and restarted without data loss. They also have a core web application with a stable, predictable baseline load that runs 24/7. The architect needs to minimize compute costs for both workloads.

Compute often represents the largest portion of cloud spend. Effective optimization is crucial.

  • Right-Sizing: Continuously evaluating compute instance sizes to ensure they are appropriately matched to workload requirements. Avoiding over-provisioning (paying for unused capacity) or under-provisioning (leading to performance issues).
    • Practical Relevance: Use "Amazon CloudWatch metrics" (CPU, memory, network) and "AWS Compute Optimizer recommendations".
  • "Reserved Instances (RIs)": A purchasing option that provides a discounted hourly rate for specific "EC2 instances" in exchange for a 1- or 3-year commitment.
    • Practical Relevance: Ideal for steady-state, predictable workloads. Can be "Convertible" (flexible instance family/OS) or "Standard" (fixed).
  • "Savings Plans": A flexible pricing model that offers significant discounts on compute usage in exchange for a 1- or 3-year commitment to a consistent amount of compute usage, measured in $/hour.
    • Practical Relevance: Provides more flexibility than "RIs", automatically applying discounts across instance families, regions, and even services. Ideal for reducing overall compute spend.
  • "Spot Instances": A purchasing option that leverages unused "EC2" capacity. Offer up to 90% discount compared to "On-Demand" prices.
    • Practical Relevance: Ideal for fault-tolerant, flexible, and interruption-tolerant workloads (e.g., batch jobs, testing, stateless microservices). Instances can be interrupted with a 2-minute warning.
  • Managed Services/Serverless: (Covered in 3.2.1.4) Often more cost-efficient for highly variable workloads due to pay-per-use models.
Visual: Compute Cost Optimization Strategies
Purchasing and Right-Sizing Details:
  • Commitment types: "Compute Savings Plans" (up to about 66%) apply to any instance family, size, Region, OS or tenancy plus "Fargate" and "Lambda"; "EC2 Instance Savings Plans" (up to about 72%) commit to one family in one Region, with size, OS and tenancy flexibility inside it, so resizing within the family keeps the discount and it is the highest-discount flexible choice for a stable single-family workload. "Standard RIs" (up to about 72%) fix the instance attributes; "Convertible RIs" (up to about 66%) can be exchanged for another family but do not cover "Fargate"/"Lambda". Longer term and more upfront payment mean a deeper discount (3-year All Upfront is the maximum). Commit only to the steady baseline (a commitment sized for peak pays for idle hours, while running the baseline On-Demand forgoes its discount); cover mixed families, "Fargate", or a planned family change (such as moving to Graviton) with "Compute Savings Plans", and leave bursts on "On-Demand"/"Spot".
  • Spot in an Auto Scaling group: A mixed instances policy lets one group use several instance types and purchase options: an "On-Demand" base capacity for the baseline that must never be starved, "Spot" above it for fault-tolerant work, diversified across many instance types and "AZs" with the price-capacity-optimized or capacity-optimized allocation strategy to limit interruptions (lowest-price concentrates capacity in the cheapest pools and raises interruption risk), and handle the two-minute Spot interruption notice so in-flight work can drain. Running the baseline on Spot, or buying RIs sized for peak, defeats the purpose.
  • Capacity and licensing: "On-Demand Capacity Reservations" reserve capacity but give no discount on their own. "Dedicated Hosts" give a whole physical server with visibility of sockets, cores and host ID for per-socket/per-core BYOL licensing; "Dedicated Instances" isolate hardware but do not give that visibility.
  • Cut idle time first: Environments used ~25% of the time (dev/test, business hours) are cheapest when stopped outside hours, using "Instance Scheduler on AWS" or "EventBridge" schedules driving automation; stopped "EC2" instances stop incurring compute charges (their "EBS" volumes still bill), and "RDS" instances can be stopped too (up to 7 days at a time). That beats an RI for a resource that is off most of the time. "Trusted Advisor" flags idle or low-utilization resources and unassociated "Elastic IP" addresses.
  • Right-sizing evidence: "Compute Optimizer" uses machine learning over historical utilization to recommend a better type/size for "EC2", Auto Scaling groups, "EBS", "Lambda" and others; memory-based findings need the "CloudWatch" agent publishing memory. The "Trusted Advisor" low-utilization check is a coarse CPU/network screen, not a full recommendation. Right-size before buying commitments, and apply it to "RDS" instance classes too when CPU and memory are far below capacity.

⚠️ Common Pitfall: Buying "Reserved Instances" for workloads with unpredictable or fluctuating usage patterns. This leads to paying for capacity that isn't used. "Savings Plans" or "On-Demand" with "Auto Scaling" are better for such workloads.

Key Trade-Offs:
  • Flexibility vs. Discount: "Standard RIs" offer the highest discount but are the least flexible (locked to an instance family). "Savings Plans" offer slightly lower discounts but are much more flexible, applying across instance families, regions, and even services.

Reflection Question: How would you apply different "EC2" purchasing options (e.g., "Spot Instances", "Reserved Instances", "Savings Plans") to optimize costs for both a fault-tolerant, interruptible batch workload and a stable, predictable 24/7 web application? Explain the rationale for each choice based on workload characteristics.

See how it connects
Alvin Varughese
Written byAlvin Varughese
Founder•20 professional certifications