3.1.2.1. Appropriate Metrics for Scaling Services
3.1.2.1. Appropriate Metrics for Scaling Services
Scaling on the wrong metric wastes money (premature scale-out) or causes outages (delayed scale-out). The best metric directly reflects user experience.
Metric selection by workload type:
| Workload | Best Scaling Metric | Why |
|---|---|---|
| Web servers | ALBRequestCountPerTarget | Directly measures user load |
| API servers | Target response time (p95 latency) | Detects degradation before failures |
| Queue processors | ApproximateNumberOfMessagesVisible (SQS) | Matches processing capacity to queue depth |
| Batch processing | CPU utilization | Compute-bound workloads |
| In-memory cache | Memory utilization | Prevents evictions |
| GPU workloads | GPU utilization (custom metric) | Matches GPU demand |
Target tracking vs. step scaling:
- Target tracking: Set a target value (e.g., CPU = 50%). ASG automatically adjusts capacity to maintain the target. Simplest and recommended for most cases. It scales in more gradually than it scales out and holds off a scale-in that would push the metric back over the target, so it avoids the add-then-remove flapping of a fixed simple-scaling step around one threshold.
- Step scaling: Define thresholds with specific capacity adjustments (e.g., CPU > 70% → add 2, CPU > 90% → add 4). More control, more complexity.
- Scheduled scaling: Set capacity based on known traffic patterns (e.g., scale up at 8 AM, down at 8 PM). Use alongside dynamic scaling.
Custom metrics published to CloudWatch enable scaling on application-specific indicators:
# Publish custom metric: active WebSocket connections
aws cloudwatch put-metric-data \
--namespace "MyApp" \
--metric-name "ActiveConnections" \
--value 1250 \
--dimensions InstanceId=i-1234567890abcdef0
Exam Trap: CPU utilization is a poor scaling metric for I/O-bound workloads (APIs waiting on database queries). CPU stays low while requests queue up and latency spikes. Use ALBRequestCountPerTarget or custom latency metrics instead. If the exam describes "low CPU but high latency," the answer is changing the scaling metric, not adjusting the CPU threshold.
Memory-bound fleets: EC2 publishes no memory metric (detailed monitoring only raises CPU/network/disk metrics to 1-minute frequency), so install the CloudWatch Agent to publish memory utilization and attach a target tracking policy to that custom metric. For queue consumers, AWS recommends a backlog per instance metric — ApproximateNumberOfMessagesVisible ÷ running instances — rather than the incoming message rate.