3.2.1. Amazon CloudWatch for Application Monitoring (Metrics & Logs)
First Principle: Amazon CloudWatch provides a centralized service for collecting and analyzing application metrics and logs, enabling developers to monitor application health, performance, and identify operational issues.
Amazon CloudWatch is the primary monitoring and observability service for AWS. For developers, it's essential for understanding how their applications are performing in real-time.
- Metrics: (Time-series data points that represent a measurement of a particular aspect of a resource or application.) Collects standard metrics from AWS services (e.g., Lambda invocations, DynamoDB read/write capacity). Developers can also publish custom metrics from their applications (e.g., successful transactions, API call latency). Lambda publishes metrics such as
Invocations,Errors,Duration,ThrottlesandConcurrentExecutionsby default, but not memory used: that appears only in each invocation'sREPORTlog line (or through Lambda Insights), so publish it as a custom metric or extract it with a metric filter. - Logs: (Centralizes logs from various sources, such as Lambda functions, EC2 instances, and custom applications.) Allows for real-time monitoring and searching. Developers use CloudWatch Logs Insights for ad-hoc querying.
- Alarms: (Monitors metrics and automatically triggers actions when a defined threshold is breached.) Notifies developers of critical issues (e.g., high error rates).
- Dashboards: Create customizable visualizations of metrics and alarms for operational oversight. One dashboard can combine metrics from different services (Lambda errors, API Gateway latency, DynamoDB throttles), and a line, stacked-area or bar graph widget can split a metric by a dimension (e.g., Lambda
ErrorsbyFunctionName). - Publishing custom metrics: Call
PutMetricDatafrom the SDK or CLI (namespace, metric name, dimensions, value), or use the Embedded Metric Format in logs (see 3.2.1.2) to avoid the API call. - Reading key metrics: Lambda
Throttlesrising whileConcurrentExecutionssits at the concurrency limit means new invocations are being rejected (synchronous callers get 429 errors; asynchronous events are retried), so users see latency or failures. On an ALB,HTTPCode_Target_5XX_Countmeans the targets returned 5xx (HTTPCode_ELB_5XX_Countis the load balancer itself), andTargetConnectionErrorCountmeans the ALB could not connect to its targets. Sustained highCPUUtilizationon the RDS primary points at the database as the bottleneck.
Scenario: Your application, deployed on AWS Lambda and using DynamoDB, starts experiencing increased latency. You need to quickly identify if the issue is with the Lambda function itself, the DynamoDB table, or an integration point.
⚠️ Exam Trap: CloudWatch custom metrics are NOT free. Each custom metric costs money, and high-resolution metrics (1-second intervals) cost more. If a question asks about cost-effective monitoring, standard-resolution (1-minute) custom metrics are preferred.