An extra 30% off every course until Sunday, October 11.Choose your certification →

Copyright (c) 2026 MindMesh Academy. All rights reserved. This content is proprietary and may not be reproduced or distributed without permission.

3.2.1.1. Monitoring Applications & Infrastructure Overview

3.2.1.1. Monitoring, Logging & Observability Overview

Monitoring tells you something is wrong. Observability tells you why. The distinction matters: you can monitor a system you don't understand, but you can't diagnose it without observability.

Three pillars of observability:
  • Metrics: Numeric measurements over time (CPU utilization, request count, error rate). CloudWatch Metrics is the primary service.
  • Logs: Discrete events with context (application errors, access logs, audit trails). CloudWatch Logs, S3, OpenSearch.
  • Traces: End-to-end request paths across distributed services. AWS X-Ray captures traces showing latency at each service hop.

How they work together: A CloudWatch alarm fires on high error rate (metric). You query CloudWatch Logs Insights to find the failing requests (logs). You trace a failing request through X-Ray to discover a downstream service timeout (trace). Without all three, you're guessing.

AWS observability stack:
PillarServicePurpose
MetricsCloudWatch MetricsCollect, store, and alarm on time-series data
LogsCloudWatch LogsCentralized log ingestion, search, and retention
TracesAWS X-RayDistributed tracing across microservices
DashboardsCloudWatch DashboardsUnified visualization of metrics and logs
SyntheticsCloudWatch SyntheticsCanary scripts that probe endpoints on schedule
RUMCloudWatch RUMReal user monitoring from browsers
ML anomaly detectionCloudWatch anomaly detectionTrains on up to two weeks of a metric's history (trends plus hourly, daily, and weekly patterns) to draw a band of expected values; anomaly detection alarms fire above or below the band with no static threshold, and their state changes go to EventBridge for automated remediation
ProfilingAmazon CodeGuru ProfilerContinuous, low-overhead runtime profiling of live Java/JVM and Python apps — flame graphs of the most expensive methods plus CPU-cost recommendations (CodeGuru Reviewer is static review of source code, not profiling)

Exam Trap: CloudWatch collects metrics from most AWS services automatically (EC2 CPU, RDS connections, ALB request count). But memory utilization and disk usage for EC2 are not collected by default — you must install the CloudWatch Agent to publish these as custom metrics. This is one of the most frequently tested facts on the exam.

See how it connects
Alvin Varughese
Written byAlvin Varughese
Founder•20 professional certifications