30% off every course until Sunday, October 11. Our biggest update yet, and we'd like you to try it. Applied automatically at checkout.

Choose your certification
Copyright (c) 2026 MindMesh Academy. All rights reserved. This content is proprietary and may not be reproduced or distributed without permission.

3.1.1.4. Chaos Engineering and Resiliency Testing (Fault Injection Simulator)

3.1.1.4. Chaos Engineering and Resiliency Testing (Fault Injection Simulator)

💡 First Principle: Proactively and deliberately injecting failures into a system in a controlled environment is the only way to build true confidence in its ability to withstand real-world disruptions.

Scenario: A financial trading platform's architect needs to ensure the application's high availability mechanisms are robust and truly work as designed under stress. They want to systematically test how the system reacts to unexpected failures, such as the termination of key compute instances or network isolation.

Chaos Engineering is the discipline of experimenting on a system in production to build confidence in that system's capability to withstand turbulent conditions. It's about breaking things intentionally to learn and improve.

  • Purpose:
    • Verify assumptions about system resilience.
    • Uncover hidden architectural weaknesses.
    • Build muscle memory for incident response.
    • Validate automated recovery mechanisms.
  • "AWS Fault Injection Service (AWS FIS)" (formerly Fault Injection Simulator): A fully managed service for running chaos engineering experiments on AWS.
    • Practical Relevance: Allows you to create and run experiments that inject faults (e.g., terminate "EC2 instances", disrupt network connectivity, induce CPU/memory stress) into your AWS workloads in a controlled, safe manner. It helps validate your "HA", "DR", and auto-healing strategies.
    • Experiment Templates: Define targets (e.g., specific "EC2 instances"), actions (e.g., aws:ec2:stop-instances), and stop conditions (e.g., "CloudWatch Alarm" threshold).
  • Game Days: Structured, simulated production incidents where teams practice incident response and validate system resilience in a realistic environment.
    • Practical Relevance: Beyond automated tooling, "Game Days" build team confidence, expose procedural gaps, and improve communication under pressure.
Visual: Chaos Engineering with AWS FIS
FIS Actions and Non-Disruptive DR Tests:
  • Pick the action for the failure you want: aws:ec2:terminate-instances/stop-instances (target by tag or "AZ" to emulate losing an "AZ"'s instances); aws:ssm:send-command and aws:ssm:start-automation-execution run an "SSM" document on instances (CPU stress, latency or packet loss, a host-firewall script); aws:network:disrupt-connectivity targets subnets (it clones the subnet's network ACL to deny traffic, with a scope such as all or availability-zone); aws:ecs:task-network-latency and the other ECS task actions inject faults into tasks (Fargate needs the SSM agent sidecar); aws:rds:failover-db-cluster and aws:rds:reboot-db-instances exercise database failover; aws:fis:inject-api-internal-error makes API calls fail rather than slow. A custom script or a manual console termination lacks FIS's IAM scoping, stop conditions and logging.
  • Test the DR plan without touching production: Build or scale up the DR environment in isolation (infrastructure as code, or scale a pilot light up from its AMIs), restore or promote data into it, and exercise it through a separate test DNS name pointing at the DR load balancer. Do not repoint production DNS, fail production over in business hours, or shut down the production primary. Time the run to validate the "RTO"/"RPO", then tear down on a schedule. Reading the runbook proves nothing, and replication alone is not a tested recovery. To test writes against a replica-based DR database, promote the DR-Region read replica: that exercises the real failover path and gives a writable copy isolated from production, but it ends replication, so recreate the replica afterward. Restoring a copied snapshot tests a restore, not the replica failover.

⚠️ Common Pitfall: Running chaos experiments without clear hypotheses and safeguards. An uncontrolled experiment can easily cause a real production outage. Always define what you expect to happen and have automatic stop conditions in place.

Key Trade-Offs:
  • Testing Realism vs. Production Safety: The goal is to test as realistically as possible without causing an actual user-impacting outage. This requires starting with a small blast radius and gradually increasing the scope of experiments.

Reflection Question: How would you use "AWS Fault Injection Service (AWS FIS)" to conduct controlled chaos engineering experiments on a financial trading platform to validate its high availability mechanisms, and what safeguards (e.g., stop conditions, blast radius management) would you implement to ensure these experiments don't cause unintended outages in production?

See how it connects
Alvin Varughese
Written byAlvin Varughese
Founder•20 professional certifications