The ANS-C01 exam retires on December 31, 2026
You can still take and pass the exam until then — plan your exam date accordingly.
4.2.2. Troubleshooting VPN/Direct Connect
Troubleshooting VPN and Direct Connect (DX) involves systematically verifying configuration, routing, and tunnel/VIF status on both AWS and on-premises sides to restore hybrid cloud connectivity.
Scenario: Your on-premises data center has lost connectivity to your AWS VPC over a Site-to-Site VPN connection. You've confirmed your internet connection is up.
Connectivity issues in hybrid cloud environments can be complex due to interactions between on-premises devices and AWS.
Key Troubleshooting Steps for VPN/Direct Connect:
- Check AWS Side:
- Site-to-Site VPN:
- Verify the status of the two VPN tunnels in the AWS Management Console (VPN tunnel state CloudWatch metric). Both should be UP.
- Check Virtual Private Gateway (VPG) or Transit Gateway (TGW) VPN attachment route tables for correct routes to on-premises.
- Verify VPC route tables have routes to the VPG or TGW.
- If the tunnels and BGP are UP but on-premises prefixes are missing from the VPC route table, check that route propagation from the VGW is enabled on that route table. With a static VPN, each new on-premises subnet also needs a new static route on the VPN connection itself.
- AWS Direct Connect:
- Verify the status of the DX connection and Virtual Interfaces (VIFs) in the Direct Connect console. Both should be UP.
- Check BGP session status on DX VIFs.
- Verify route tables (VPC, Direct Connect Gateway, TGW) have correct routes for on-premises prefixes.
- Site-to-Site VPN:
- Check On-premises Side:
- Verify the status of your customer gateway device (VPN router/firewall) or Direct Connect router.
- Check VPN tunnel status or DX physical connection.
- Verify on-premises route tables have routes to AWS VPC CIDRs.
- Ensure firewall rules on-premises are not blocking traffic.
- Check DNS Resolution: Verify DNS settings are correct on both sides for resolving hostnames.
- Use VPC Flow Logs: Analyze logs from the VPC side to see if traffic is reaching the gateway or being rejected.
- Separate Control Plane from Data Plane: Tunnel
UPand BGP routes received prove the control plane; then prove the data plane withping/traceroutefrom an EC2 instance to an on-premises private IP (and the reverse).traceroutefrom on-premises also shows which path — Direct Connect or the backup VPN — packets actually take. Reachability Analyzer analyzes only AWS-side configuration: a path can end at a virtual private gateway or transit gateway, but it cannot see into the on-premises network. - Inspect BGP Routes on the VIF: A virtual interface's Accepted routes and Advertised routes tabs in the Direct Connect console (or
aws directconnect list-virtual-interface-routes) show each prefix with its AS path and communities — the place to confirm that an AS_PATH prepend or community change was received. VPC and TGW route tables show the resulting routes but not BGP attributes. - Work Up the Layers on Direct Connect: Physical link up but the VIF never comes up and the router cannot ARP the AWS peer → Layer 2, most often a VLAN ID mismatch. Layer 2 fine but BGP down → peer IP, ASN or MD5 key mismatch, or more than 100 prefixes advertised on a private or transit VIF (the session goes idle). BGP flapping → hold-timer expiry (keepalives lost or processed too late, e.g., a CPU-starved router) or an MTU mismatch that drops large UPDATE packets.
- Read the Negotiation Logs: For IKE/IPsec negotiation and rekey problems, read the customer gateway device's logs and, on the AWS side, Site-to-Site VPN logs published to CloudWatch Logs (IKE phase states, IPsec/DPD messages, BGP status). CloudTrail shows only API calls, and the console's tunnel details show status, not negotiation steps. Alert on outages with a CloudWatch alarm on the
TunnelStatemetric (1 = UP, 0 = DOWN). - Test Failover Deliberately: Shut down the primary path's BGP session on the on-premises router, or run the Direct Connect Resiliency Toolkit failover test (or the AWS FIS
aws:directconnect:virtual-interface-disconnectaction), which puts the chosen VIF's BGP sessions down for a set time and then restores them. Deleting gateways or changing VLAN tags is destructive, not a test.
Practical Implementation: Checking VPN Tunnel Status (CLI)
# 1. Describe VPN connections to get tunnel details
aws ec2 describe-vpn-connections --vpn-connection-ids vpn-0abcdef1234567890 --query "VpnConnections[0].VgwTelemetry"
# Expected output will show TunnelState (UP/DOWN), LastStatusChange, StatusMessage
# Example:
# [
# {
# "OutsideIpAddress": "198.51.100.1",
# "Status": "UP",
# "LastStatusChange": "2023-10-27T10:00:00.000Z",
# "StatusMessage": "Tunnel is up.",
# "AcceptedRouteCount": 10
# },
# {
# "OutsideIpAddress": "203.0.113.1",
# "Status": "DOWN",
# "LastStatusChange": "2023-10-27T09:50:00.000Z",
# "StatusMessage": "IKE negotiation failed.",
# "AcceptedRouteCount": 0
# }
# ]
⚠️ Common Pitfall: Focusing only on the AWS side. Hybrid connectivity issues often stem from misconfigurations or failures on the on-premises network devices.
Key Trade-Offs:
- Manual Inspection vs. Automated Monitoring: While manual checks are necessary, automated monitoring (e.g., CloudWatch alarms on VPN TunnelState) provides proactive alerts.
Reflection Question: How does systematically verifying configuration, routing, and tunnel/VIF status on both the AWS side (e.g., VPN tunnel state in CloudWatch) and the on-premises side fundamentally help you troubleshoot VPN and Direct Connect (DX) issues and restore hybrid cloud connectivity?