Power Your Software Testing with AI Agents and Cloud
The Native AI-Agentic Cloud Platform to Supercharge Quality Engineering. Test Intelligently and Ship Faster.
- TestMu AI (Formerly LambdaTest)
- /
- Blog
- /
- Chaos Engineering in Software Testing: A Practical Guide
Chaos Engineering in Software Testing: A Practical Guide
Learn how chaos engineering works in software testing, how QA teams design chaos experiments, and how a blast radius keeps production failures contained.
Last Updated on:
Chaos engineering in software testing injects random failures into a live production system to expose weaknesses before customers hit them. Each experiment sets a steady state baseline, injects a fault such as pod termination or network latency with a tool like LitmusChaos or AWS Fault Injection Service, and contains it inside a defined blast radius. This guide covers what chaos engineering is, its benefits, how QA and chaos testing mix and match, testing with a blast radius, how to chaos test an AI powered feature, and productive chaos.
Key Takeaways
- Chaos engineering tests a distributed system on a live production server with random, unexpected failure conditions to find weaknesses before customers are affected.
- A chaos experiment starts by defining a steady state baseline for the application and the server, then checks whether that steady state still holds while failures are injected.
- LitmusChaos, Chaos Mesh and AWS Fault Injection Service inject faults such as pod termination, network latency and CPU stress, and save each chaos experiment as a reusable template that can run from a pipeline.
- A blast radius confines a chaos test to one zone of related functions, so an injected failure stays contained and the production server can be restored quickly.
- Chaos tests run against a live production server, so QA testers design the failure scenarios while a DevOps or IT engineer stands by to restore the server.
- Microsoft Dev Proxy chaos tests an AI powered feature by returning a 429 once a prompt or completion token budget is spent, and its LanguageModelFailurePlugin simulates 15 named failure types including Hallucination.
What is Chaos Engineering?
Chaos engineering means testing a distributed computer system using random and unexpected failure conditions to identify weaknesses present in the system. Random and unexpected actions, failures, and conditions equal chaos.
Chaos engineering is a software development methodology that enables testing creativity and expanded test coverage to discover and plan for system errors. Not the average system error, but catastrophic errors that take down the network and cause customer access interruptions for any length of time.
Netflix established the practice while transferring its entire infrastructure to AWS. Netflix developed two principles to test to prevent or minimize the impact of the move on customers.
Chaos engineering principles include:
- Systems never have a single point of failure.
- A single point of failure refers to the possibility a failure in the system leads to customer interruption or significant access downtime.
- Systems always have at least one single point of failure.
- Software development teams must create effective tests and monitor the system to ensure there is never a single point of failure.
Chaos engineering proactively identifies errors to prevent production server outages from impacting customers. Chaos engineering is not random, or undisciplined testing. Chaos engineering relies on the ability to monitor the production server and execute real-life test simulations to determine how the application responds to failures in integrated or connected services and systems.
Chaos engineering includes performing the following functions on the production server:
- Define a steady-state or baseline to measure the application and server against.
- Determine if the defined steady-state holds during experimental testing.
- Test with minimal impact on users by defining and implementing tests within a blast radius.
- Defining a blast radius means chaos tests are focused on a particular area and the resources are available to immediately respond to failures.
- Introduce the planned chaos events in order, contained by the defined blast radius.
- Introduce scenarios to mimic real-world failure scenarios. Failure scenarios examples include:
- Monitor testing and repeat test scenarios being as creative with failure scenarios as possible.
Netflix built its own tool during that migration. Chaos Monkey randomly terminates virtual machine instances and containers running inside a production environment, and Netflix released the source on GitHub in 2012 under the Apache 2.0 license. That version carries one dependency worth checking before you plan around it. Its README states you must be managing your apps with Spinnaker to use Chaos Monkey to terminate instances.
Teams now run these experiments with dedicated tools rather than hand written scripts. LitmusChaos and Chaos Mesh are CNCF incubating projects that inject faults into Kubernetes clusters, including pod termination, network latency, packet loss, CPU and memory stress, and disk faults. AWS Fault Injection Service does the same job inside AWS and stops an experiment on its own when a CloudWatch alarm you nominate fires. Each tool saves an experiment as a reusable template, so a chaos test can sit in version control and run from a pipeline like any other test.
Key Takeaway: Chaos engineering injects random failure conditions into a live distributed system, measured against a defined steady state and contained by a blast radius, to expose catastrophic weaknesses that ordinary testing never triggers.
Benefits of Chaos Engineering & Chaos Testing
Chaos engineering benefits an organization by identifying server and application vulnerabilities, integration failures, and system crashes before the customer experience is impacted. The production system continues to perform as expected with each new release regardless of the nature of the changes or updates.
Other benefits of chaos engineering include:
- Faster issue identification and correction not captured by other QA testing efforts.
- Fewer unplanned outages and downtime.
- Provides ongoing system monitoring on the production server.
- Increases test depth and coverage with controlled testing in production.
Key Takeaway: Chaos engineering finds server vulnerabilities, integration failures and crashes before customers see them, which reduces unplanned outages and adds test coverage that other QA testing does not reach.
QA and Chaos Testing - Mix & Match
Chaos engineering appears similar to stress, load, and performance testing. However, the primary purpose is chaos or the randomness of the testing. For example, in chaos engineering, the system's optimal or baseline state is set. Then, testers consider potential weaknesses and the effects of those on the customer experience and create a test scenario for each. Each test is then executed with assistance from DevOps and with resources available to repair the production server when tests successfully find problems.
In other types of performance testing, the application performance is tested when running on a test or development server. Often functional application tests are transformed into performance tests based on the user workflow. In a typical performance, stress, or load test, testers execute based on known factors against an expected result, rather than crash or cause production server failures.
Chaos engineering also must involve IT or DevOps to manage issues on the production server. If failures are caused by testing in a blast radius, resources must be ready to reinstate the production server as needed.
It's common for a DevOps engineer to execute chaos engineering testing. However, there's no reason QA testers cannot also design and execute chaos engineering testing. Coordination and cooperation between QA testing and DevOps during testing are key. QA testers have the skills to break software including hardware and backend connections, but they may not have the skills to restore the production server to normal operations rapidly. Leverage the QA tester's ability and desire to break software to the business's advantage with chaos engineering.
Mix and match QA testing resources with DevOps to ensure optimal chaos test development, execution, and support when testing in production. Add chaos test scenarios to scheduled regression testing even on a test server. Determine what all can be tested first on the test servers and then move into production. Adding chaos tests improves the depth and test coverage of QA testing while providing business value.
Key Takeaway: Chaos testing differs from stress, load and performance testing because chaos tests run on the production server against unknown failures, so QA testers design the scenarios while DevOps stands ready to restore the server.
Testing with a Blast Radius
Using a blast radius enables production level testing without negatively impacting the production server or taking it down completely. Designate distinct blast radius zones for similar functions. Next, group test scenarios into their related blasting zones. Executing tests by blast radius ensures failure control and reduces the possibility of unexpectedly and completely crashing the production server.
During chaos engineering testing, expect disruption. Coordinating efforts between IT, DevOps and QA testing is critical to minimize adverse effects on the production server and the customer experience. Ensure redundancy measures are in place to keep the server operational when chaos engineering testing causes issues.
One basic blast radius worth considering is the timing of test execution. Execute tests at non-peak periods to minimize performance impact on customers.
Resilience testing is now a legal requirement in some markets, so record what you do. The EU Digital Operational Resilience Act has applied to banks, insurers, investment firms and their critical ICT providers since 17 January 2025, and it requires both basic and advanced testing of ICT systems, including threat led penetration testing. Teams in regulated sectors should log the scope of each chaos experiment, the blast radius they set, the failures they injected and the result, because those records serve as audit evidence.
Key Takeaway: Grouping chaos tests into blast radius zones of related functions and running them at non-peak hours keeps an injected failure contained, and teams covered by the EU Digital Operational Resilience Act should log each experiment as audit evidence.
How Do You Chaos Test an AI Powered Feature?
You chaos test an AI powered feature by injecting faults into the model API it calls, then checking that the app degrades in a controlled way instead of failing in front of the user. A language model dependency breaks in ways a CPU or network fault does not reproduce. The API can throttle the request, or it can return an answer that is confident, well formed and wrong. A steady state defined only on server metrics misses both cases.
Microsoft Dev Proxy injects these faults at the network layer, so the application code does not change. Its LanguageModelRateLimitingPlugin sets a prompt token limit and a completion token limit for a fixed window, and returns a standard 429 once either budget is used up. The documented example allows 1,000 prompt tokens and 500 completion tokens in a 60 second window. You can swap that for a custom error body with a retry-after header that counts down to the reset. It works with any OpenAI compatible API, including a local model served through Ollama.
The content side has its own plugin. LanguageModelFailurePlugin simulates 15 named failure types, among them Hallucination, PlausibleIncorrect, IncorrectFormatStyle, FailureFollowInstructions and ContradictoryInformation. It injects a system role message into the outgoing request, so the model produces the bad answer itself and your parsing, validation and fallback paths receive a realistic input rather than a canned string.
Run these as ordinary chaos experiments. State the hypothesis first, for example that a throttled model call shows the last cached result and a retry control. Keep the blast radius to one feature and one slice of traffic, and stop the run when the alert on that path fires.
Key Takeaway: Chaos testing an AI powered feature means injecting model API faults such as a 429 token limit response or a deliberately hallucinated answer, then confirming the feature degrades in a controlled way instead of failing in front of the user.
Productive Chaos
Chaos engineering creates real-world hardware, distributed software, and application failures in distributed systems. Chaos provides deeper testing into the vulnerabilities present in complex, integrated computer systems and the hardware they use. The purpose of chaos engineering is to ensure production server integrity.
Chaos engineering improves customer experience by reducing the number of failures or system crashes possible or present in production. Chaos engineering testing is executed by DevOps or QA testing teams on production servers with resources ready and able to keep production running in case of issues. The key to success is coordination and cooperation between DevOps and QA testing teams. Chaos works better by leveraging operational, test development, and defect-finding skills. Eliminate downtime on production and disruptions to the customer experience by executing chaos testing frequently.
Austin Siewert
Co-Founder, Steadfast Systems
Discovered @TestMu AI yesterday. Best browser testing tool I've found for my use case. Great pricing model for the limited testing I do 👏
2M+ Devs and QAs rely on TestMu AI
Deliver immersive digital experiences with Next-Generation Mobile Apps and Cross Browser Testing Cloud
Author
Amy Reichert is a software quality assurance professional with 25+ years of experience in manual testing for web and mobile applications across healthcare, enterprise, and SaaS domains. She specializes in test case design, exploratory testing, regression, integration, and API testing using Postman, with strong experience in QA process leadership and test strategy. Amy holds ISTQB CTFL and CTAL-TA certifications and has authored multiple articles on software testing practices and QA careers, combining hands-on testing expertise with technical writing.
Chaos Engineering FAQs
Did you find this page helpful?
More Related Blogs
TestMu AI forEnterprise
Get access to solutions built on Enterprise
grade security, privacy, & compliance
- Advanced access controls
- Advanced data retention rules
- Advanced Local Testing
- Premium Support options
- Early access to beta features
- Private Slack Channel
- Unlimited Manual Accessibility DevTools Tests



