Running Chaos Experiments Without Breaking Production
Chaos engineering is running controlled experiments to check how a system behaves when something fails. What separates it from randomly breaking things is the setup around the fault: a hypothesis, a measured steady state, a limited blast radius, and a way to stop immediately. Without those, a chaos experiment is just an incident you scheduled yourself.
The method
The Principles of Chaos Engineering (principlesofchaos.org) describe the loop:
- Define steady state as a measurable output, such as request success rate or checkout latency, not CPU usage.
- Form a hypothesis: steady state continues in both the control group and the experimental group.
- Introduce a real-world event: an instance dies, a dependency slows down, a zone becomes unreachable.
- Look for a difference between control and experiment. A difference means the hypothesis is wrong and you have something to fix.
Every experiment gets a short written plan before it runs:
name: my-app-redis-latency
hypothesis: with 300ms added latency to Redis, my-app p99 stays under 800ms and error rate under 1%
steady_state: dashboard https://grafana.example.com/d/my-app, panels "p99 latency" and "5xx ratio"
fault: NetworkChaos delay 300ms on my-app pods, traffic to redis only
scope: namespace my-app-staging, one pod first
duration: 5m
abort_if: 5xx ratio above 2% for 1m, or any page fires
rollback: delete the NetworkChaos object
owner: on-call engineer for my-app
Prerequisites
- Observability good enough to see the steady state in near real time. If you cannot see the impact, you cannot stop it.
- Fix known weaknesses first. If you already know a component has no failover, an experiment only confirms it. Experiments are for things you believe work.
- The on-call engineer knows an experiment is running, and the people running it can stop it in seconds.
- Run during working hours with the owning team present.
Limit the blast radius
Start in staging. Move to production only after the experiment passes in staging, and start there with the smallest scope: one pod, one instance, a small share of traffic. Widen the scope one step at a time.
On Kubernetes, Chaos Mesh can be restricted to opted-in namespaces. Install it with controllerManager.enableFilterNamespace=true and annotate only the namespaces where experiments are allowed:
helm repo add chaos-mesh https://charts.chaos-mesh.org
helm upgrade --install chaos-mesh chaos-mesh/chaos-mesh -n chaos-mesh --create-namespace \
--set chaosDaemon.runtime=containerd \
--set chaosDaemon.socketPath=/run/containerd/containerd.sock \
--set controllerManager.enableFilterNamespace=true
kubectl annotate namespace my-app-staging chaos-mesh.org/inject=enabled
Everything else is protected even if someone writes a selector that is too broad.
Kubernetes examples
Kill one random pod of my-app:
apiVersion: chaos-mesh.org/v1alpha1
kind: PodChaos
metadata:
name: my-app-pod-kill
namespace: my-app-staging
spec:
action: pod-kill
mode: one
selector:
namespaces: [my-app-staging]
labelSelectors:
app: my-app
pod-kill uses a grace period of 0 by default, so the pod gets no chance to shut down cleanly. That is closer to a node crash than kubectl delete pod, which sends SIGTERM and waits for the pod's termination grace period. Both are worth testing, but they test different things: crash recovery versus graceful shutdown.
Add latency to traffic from my-app to Redis:
apiVersion: chaos-mesh.org/v1alpha1
kind: NetworkChaos
metadata:
name: my-app-redis-latency
namespace: my-app-staging
spec:
action: delay
mode: one
selector:
namespaces: [my-app-staging]
labelSelectors:
app: my-app
direction: to
target:
mode: all
selector:
namespaces: [my-app-staging]
labelSelectors:
app: redis
delay:
latency: "300ms"
jitter: "50ms"
duration: "5m"
Deleting the object ends the fault: kubectl -n my-app-staging delete networkchaos my-app-redis-latency. Before a pod-kill experiment, check that the deployment has more than one replica, a readiness probe, and a PodDisruptionBudget. If not, the result is already known.
AWS example with automatic stop
AWS Fault Injection Service (FIS) supports stop conditions: CloudWatch alarms that end the experiment automatically if they go into alarm. This template stops one tagged staging instance and starts it again after five minutes:
{
"description": "Stop one my-app instance in staging",
"roleArn": "arn:aws:iam::111122223333:role/fis-my-app",
"targets": {
"myAppInstances": {
"resourceType": "aws:ec2:instance",
"resourceTags": { "app": "my-app", "env": "staging" },
"filters": [{ "path": "State.Name", "values": ["running"] }],
"selectionMode": "COUNT(1)"
}
},
"actions": {
"stopOne": {
"actionId": "aws:ec2:stop-instances",
"parameters": { "startInstancesAfterDuration": "PT5M" },
"targets": { "Instances": "myAppInstances" }
}
},
"stopConditions": [
{
"source": "aws:cloudwatch:alarm",
"value": "arn:aws:cloudwatch:eu-central-1:111122223333:alarm:my-app-5xx-high"
}
]
}
aws fis create-experiment-template --cli-input-json file://stop-one.json
aws fis start-experiment --experiment-template-id <template-id>
Give the FIS role permissions only on resources with the right tags, so a mistake in the target definition cannot reach production.
During and after
During the run, one person watches the steady-state metrics and has the stop command ready. Anyone can call abort, and nobody argues about it until afterwards.
After the run, write down the result next to the plan: hypothesis confirmed or not, what the metrics showed, what surprised you. A failed hypothesis produces tickets, like a postmortem without the outage. A passed one can be scheduled to run regularly, so later changes that break the fallback are caught.
Tools
- Chaos Mesh and LitmusChaos: Kubernetes-native, both CNCF projects.
- AWS Fault Injection Service and Azure Chaos Studio: managed services for their clouds.
- Gremlin: commercial, multi-platform.
- Chaos Monkey: Netflix's instance terminator, built to run with Spinnaker.
Checklist
- Written plan: hypothesis, steady-state metric, scope, duration, abort conditions, owner.
- Staging first, then the smallest possible scope in production.
- Namespace or tag restrictions so experiments cannot reach unintended targets.
- Automatic stop conditions where the tool supports them, a manual stop ready everywhere.
- On-call informed, team present, working hours.
- Results written down and failed hypotheses turned into tickets.
