An extra 30% off every course until Sunday, October 11.Choose your certification →

Copyright (c) 2026 MindMesh Academy. All rights reserved. This content is proprietary and may not be reproduced or distributed without permission.

6.3.1. CloudWatch Metrics and Alarms for GenAI

💡 First Principle: GenAI monitoring requires a custom metrics layer on top of standard AWS infrastructure metrics because most quality signals — retrieval relevance, response accuracy, user satisfaction, hallucination rate — are not native CloudWatch metrics and must be published from your application.

Critical metrics for GenAI monitoring:
Metric CategoryMetric NameAlarm ThresholdSource
Availability5xxErrorRate> 1% over 5 minCloudWatch (native)
LatencyP99ResponseTime> 15sCustom metric from Lambda
CostDailyTokenCost> budget thresholdCustom from token counts
QualityGroundingScore< 0.7 averageCustom from Guardrails trace
SafetyGuardrailTriggerRate> 5% of requestsCustom from Guardrails trace
RetrievalAverageRetrievalScore< 0.6Custom from Knowledge Bases
ThroughputThrottledRequestRate> 2%Custom from retry logic
AccuracyUserCorrectionRate> 10%Custom from user feedback
Publishing custom quality metrics:
def publish_response_quality_metrics(response_data):
    metrics = [
        {
            'MetricName': 'GroundingScore',
            'Value': response_data['grounding_score'],
            'Unit': 'None',
            'Dimensions': [
                {'Name': 'KnowledgeBaseId', 'Value': response_data['kb_id']},
                {'Name': 'ModelId', 'Value': response_data['model_id']}
            ]
        },
        {
            'MetricName': 'RetrievalTopScore',
            'Value': response_data['top_retrieval_score'],
            'Unit': 'None'
        },
        {
            'MetricName': 'ResponseTokenCount',
            'Value': response_data['output_tokens'],
            'Unit': 'Count'
        },
        {
            'MetricName': 'TotalLatencyMs',
            'Value': response_data['total_latency_ms'],
            'Unit': 'Milliseconds'
        }
    ]
    cloudwatch.put_metric_data(Namespace='GenAI/Quality', MetricData=metrics)
Composite alarms for multi-condition alerting:
# Composite alarm: alert when BOTH latency is high AND error rate is elevated
cloudwatch.put_composite_alarm(
    AlarmName='GenAI-Critical-Degradation',
    AlarmDescription='Both latency and error rate elevated — likely capacity issue',
    AlarmRule='ALARM("GenAI-HighLatency") AND ALARM("GenAI-ElevatedErrorRate")',
    AlarmActions=['arn:aws:sns:...:GenAI-PagerDuty-Critical'],
    OKActions=['arn:aws:sns:...:GenAI-Recovery-Notification']
)

⚠️ Exam Trap: CloudWatch alarms on Bedrock native metrics (like InvocationLatency) measure the FM API call time, not your application's end-to-end response time. If your retrieval pipeline is slow, this metric won't capture it. Always instrument end-to-end latency in your application layer separately from Bedrock API latency.

Monitoring agents: AgentCore Observability

Agent failures often hide between the metrics above: every model call succeeds, yet the agent loops, calls the wrong tool, or stalls in a slow one. AgentCore Observability records each step as OpenTelemetry spans and stores the metrics, spans, and logs in CloudWatch.

  • Built-in metrics cover runtime, memory, gateway, built-in tools, and identity resources, including session count, latency, duration, token usage, and error rates. The CloudWatch GenAI Observability page adds trace visualizations and error breakdowns for runtime agents.
  • Setup: enable CloudWatch Transaction Search once per account; without it, spans do not appear. For full trace data and custom metrics, add the AWS Distro for OpenTelemetry (ADOT) SDK (aws-opentelemetry-distro) and start the agent with opentelemetry-instrument. Memory and gateway resources need tracing and log delivery turned on; they are not configured automatically.
  • Agents outside Runtime (on EKS, ECS, or Lambda) are supported through the ADOT SDK or the AWS Lambda Layer for OpenTelemetry, with OTEL environment variables such as AGENT_OBSERVABILITY_ENABLED=true and a named log group. The ADOT Collector is not supported for agent observability.
  • Correlation: pass the session ID (the X-Amzn-Bedrock-AgentCore-Runtime-Session-Id header, or session.id in OTEL baggage) so every span of a conversation groups together.
  • Quality, not only health: AgentCore Evaluations scores sessions, traces, and spans with LLM-as-a-judge evaluators (built-in ones such as helpfulness, or your own), online against live traffic or on demand.

⚠️ Exam Trap: Amazon Bedrock model invocation logging records prompts and responses for each model call. It does not show the agent's reasoning path or tool calls. Scenarios about "which step of the agent failed" point to agent traces.

Finding what drives a latency tail: a percentile alarm shows that P99 is bad, not why. CloudWatch Contributor Insights rules over structured JSON logs, such as application request logs or model invocation logs, rank the top contributors to a pattern by user, model ID or input-token size. That shows whether the tail comes from long prompts, one tenant or one query type.

Reflection Question: At 2pm on a Tuesday, users start reporting that the chatbot "keeps making things up." Your CloudWatch dashboard shows: Bedrock InvocationLatency = normal, Lambda error rate = 0%, 5xx rate = 0%. What category of metric is missing from your monitoring setup, and what specifically would have caught this issue?

See how it connects
Alvin Varughese
Written byAlvin Varughese
Founder•20 professional certifications