3.2.2. AWS X-Ray for Distributed Tracing
First Principle: AWS X-Ray provides end-to-end visibility into requests as they traverse distributed applications, enabling developers to precisely identify performance bottlenecks and errors across multiple services.
In this X-Ray trace, it's immediately clear the external API call (2500ms) is the bottleneck — not DynamoDB or S3. Without tracing, you'd be guessing.
In modern, distributed applications (especially those using microservices or serverless architectures), a single user request might traverse many different services and components. AWS X-Ray helps developers understand how their application and its underlying services are performing.
- Distributed Tracing: Collects data about requests that your application serves and traces them as they flow through various components.
- Service Map: Visualizes the components of your application and their connections, showing latency and error rates for each. This helps identify unhealthy services or bottlenecks.
- Segment Timelines: Provides a detailed breakdown of what each service or component is doing within a trace, showing execution time for each step.
- Integration: Easily integrate with AWS Lambda, Amazon API Gateway, Amazon EC2, Amazon ECS, AWS Elastic Beanstalk, and AWS SDKs (for custom instrumentation).
- Segments & Subsegments: A segment is the data one service or resource records for a request. Subsegments break it down: downstream AWS and HTTP calls (recorded automatically once the X-Ray SDK patches or wraps the AWS SDK and HTTP clients), or a custom subsegment you open around a block of code (e.g., payment validation) to time it.
- Reading the map and timeline: Nodes are coloured by health: green (successful), yellow (4xx client errors), red (5xx faults), purple (429 throttling). The client node is where traced requests enter (end users or external callers). The trace timeline is a waterfall of segments and subsegments that shows which downstream call used the time.
- Enabling it: On Lambda and API Gateway, turn on active tracing (on the function and on the API stage). On EC2 or ECS, run the X-Ray daemon (or the ADOT collector) to relay the SDK's trace data. AWS now recommends OpenTelemetry (ADOT) for new instrumentation; the X-Ray SDKs and daemon entered maintenance mode in February 2026.
- Annotation & Metadata: Developers can add custom annotations and metadata to traces for more context during debugging. Annotations are indexed key-value pairs (e.g.,
customerTier = premium) that you can search with filter expressions. Metadata stores extra objects on a segment but is not indexed, so you cannot filter traces by it. - Sampling: X-Ray does not record every request. The default rule records the first request each second plus 5% of any additional requests; custom sampling rules change the rate.
Scenario: You're developing a microservices application where a single user request passes through API Gateway, then a Lambda function, and finally interacts with a DynamoDB table. Users report intermittent delays, and you need to pinpoint where the latency is occurring.
CloudWatch vs. X-Ray: When to Use Which
| Question | Use CloudWatch | Use X-Ray |
|---|---|---|
| "What happened?" | ✅ Logs show errors, output | ❌ |
| "How is it performing?" | ✅ Metrics show duration, errors, throttles | ❌ |
| "Alert me when..." | ✅ Alarms on metrics/log patterns | ❌ |
| "Where is the bottleneck?" | ❌ | ✅ Service map + trace analysis |
| "Which downstream call failed?" | ❌ | ✅ Segments and subsegments |
| "How do services interact?" | ❌ | ✅ Service map visualization |
⚠️ Exam Trap: X-Ray requires both the X-Ray SDK in your code AND the X-Ray daemon running on the host (or enabled on the Lambda service). Missing either one means no traces. On Lambda, you enable "active tracing" — you don't run a daemon.