30% off every course until Sunday, October 11. Our biggest update yet, and we'd like you to try it. Applied automatically at checkout.

Choose your certification
Copyright (c) 2026 MindMesh Academy. All rights reserved. This content is proprietary and may not be reproduced or distributed without permission.

3.2.1.2. Storage Cost Optimization (Lifecycle Policies, Data Tiering)

3.2.1.2. Storage Cost Optimization (Lifecycle Policies, Data Tiering)

💡 First Principle: Maximizing compute cost efficiency requires dynamically matching capacity to actual workload demand and leveraging flexible pricing models to pay only for what is needed.

Scenario: A media streaming service stores video files in "Amazon S3". Recently uploaded videos are frequently accessed by users. After 60 days, their access frequency drops significantly but they still need to be available within minutes. After 1 year, videos are rarely accessed but must be retained for 10 years for archival purposes.

Storage costs can accumulate rapidly, especially for large datasets or long retention periods.

  • "S3 Storage Classes":
    • "S3 Standard": General-purpose, frequently accessed data (hot).
    • "S3 Intelligent-Tiering": Automatically moves objects between frequent, infrequent, and archive access tiers based on changing access patterns, without performance impact. Ideal for unpredictable workloads.
    • "S3 Standard-Infrequent Access (IA)": Data accessed less frequently but requiring rapid access when needed (warm). Higher retrieval cost, lower storage cost.
    • "S3 One Zone-IA": Same as "Standard-IA" but stored in a single "AZ" (less durable to "AZ" failure). Lowest cost for infrequent access if data can be lost in an "AZ" event.
    • "S3 Glacier": Archival data, very low cost, delayed retrieval.
    • "S3 Glacier Deep Archive": Lowest cost archival, retrieval in hours (coldest).
  • "S3 Lifecycle Policies": Automate the transition of objects between these "S3 storage classes" or their expiration.
    • Practical Relevance: Essential for managing data retention and cost by moving aging data to cheaper tiers and deleting outdated data.
  • "EBS Volume Types": Selecting the most cost-effective "EBS" volume type (e.g., gp3 often more cost-effective than gp2) and right-sizing the provisioned "IOPS"/throughput.
  • "EFS Performance Modes": Choose between "General Purpose" and "Max I/O" or "Bursting Throughput" and "Provisioned Throughput" based on actual needs to avoid overpaying for performance.
Visual: S3 Storage Cost Optimization Flow
Lifecycle, Visibility and EBS Details:
  • Lifecycle actions and limits: A lifecycle rule can transition objects between classes and expire (delete) them; expiration is the simplest reliable way to delete data after N days. Minimum storage durations apply (30 days for "Standard-IA"/"One Zone-IA", 90 for "Glacier Flexible Retrieval", 180 for "Glacier Deep Archive"), and objects must be 30 days old before moving from "Standard" to "Standard-IA". "Deep Archive" retrieval takes 12 hours (standard) to 48 hours (bulk). For data with a known cold-after-90-days pattern and long retention, a lifecycle transition to "Deep Archive" is cheaper than "Intelligent-Tiering", which carries a per-object monitoring fee (objects under 128 KB are not tiered) but no retrieval charges and suits unknown access patterns.
  • Finding the savings: "S3 Storage Lens" gives organization-wide dashboards (via "Organizations") across accounts, buckets and, with advanced metrics, prefixes, with usage, activity (request counts by type) and recommendations; use it to find cold or costly buckets. "Macie" finds sensitive data, not cold data; "S3 Inventory" is only an object listing; server access logs plus "Athena" is a custom build.
  • "EBS" choices: "gp3" costs less per GB than "gp2" with a 3,000 IOPS / 125 MiB/s baseline independent of size, so it is the low-cost general-purpose and boot choice; "io1"/"io2" only pay off for sustained high IOPS, and "st1"/"sc1" HDD volumes cannot be boot volumes. If metrics show provisioned IOPS far above use, change "io1" to "gp3" (for "RDS", modify the storage type). Find unattached ("available") volumes with "Trusted Advisor"'s underutilized-volume check, snapshot, then delete.
  • "Amazon Data Lifecycle Manager": Schedules "EBS" snapshots and EBS-backed AMIs (as often as hourly) and deletes them by retention rule. It does not delete volumes; removing temporary volumes after a set time needs custom automation such as an "EventBridge" schedule invoking "Lambda".

⚠️ Common Pitfall: Using "S3 Lifecycle Policies" for data with unpredictable access patterns. If an object is moved to an infrequent access tier and then suddenly becomes popular again, the retrieval costs can negate the storage savings. "S3 Intelligent-Tiering" is the better choice for unpredictable access.

Key Trade-Offs:
  • Storage Cost vs. Retrieval Fee: Infrequent access and archive tiers have very low storage costs but charge a per-GB fee for data retrieval, which can be expensive if access patterns are misjudged.

Reflection Question: How would you design a storage cost optimization strategy for a media streaming service using "Amazon S3" storage classes and "S3 Lifecycle Policies" to manage video files with varying access patterns (frequently, infrequently, rarely) and retention requirements (60 days, 1 year, 10 years), minimizing overall storage costs while meeting availability needs?

See how it connects
Alvin Varughese
Written byAlvin Varughese
Founder•20 professional certifications