From Postmortems to Pre-mortems: Preventing Repeat Incidents

From Postmortems to Pre-mortems: Preventing Repeat Incidents

Reading time1 min
#chaos engineering#devops#incident management

From Postmortems to Pre-mortems: Preventing Repeat Incidents

Most teams write postmortems. Fewer teams see their incident count go down because of them. The same kinds of failures come back: an expired certificate, a dependency timeout without a fallback, a config change that skipped review. This post covers three practices that break that loop: making postmortem actions actually happen, running a pre-mortem before risky work, and keeping a failure-mode document per service that you test on purpose.

Why incidents repeat

  • Action items die in a separate backlog. They are written in the postmortem doc, nobody owns them, and they lose every prioritization meeting.
  • The fix targets the trigger, not the condition. The bad config value is reverted, but nothing stops the next bad value from reaching production.
  • Lessons stay local. One service learns that a cache outage takes it down. The three other services using the same cache never hear about it.
  • Nobody looks across postmortems. Each one is read once, then archived.

Make postmortems produce finished work

  • Every action item becomes a ticket in the owning team's normal backlog, with an owner and a due date. Link the tickets from the postmortem.
  • Mark items that prevent recurrence differently from nice-to-haves, and review open ones in a recurring meeting until they are closed.
  • Tag each incident with contributing factors from a short fixed list, for example config-change, capacity, dependency-timeout, expired-credential, deploy. Count the tags every quarter. Recurring tags show where engineering time will pay off.
  • When a postmortem finds a failure mode, ask which other services share it, and file tickets for them too.

Run a pre-mortem before risky work

The pre-mortem is a technique described by psychologist Gary Klein in Harvard Business Review in 2007. Instead of asking "what could go wrong?", which invites polite optimism, you state that the project has already failed and ask why.

Run one before a migration, a major launch, a new region, or any change that is hard to roll back:

  1. Get the people who will build and operate the change in a room for under an hour.
  2. State the scenario: "It is a month after launch and this went badly. What happened?"
  3. Everyone writes reasons on their own for a few minutes, without discussion.
  4. Go around the room and collect one reason per person per round until the lists are empty.
  5. Group duplicates and rank by likelihood and impact.
  6. For the top items, decide on a mitigation, a test, or an explicit acceptance of the risk, each with an owner.

The silent writing step matters. It keeps the most senior person in the room from setting the direction for everyone else.

Keep a failure-mode document per service

Pre-mortems and postmortems both produce the same kind of knowledge: how this service behaves when something it depends on breaks. Store it in one place, next to the code, in a format that can be reviewed in pull requests:

# docs/failure-modes.yaml in the my-app repository
service: my-app
owner: team-payments
failure_modes:
  - id: redis-slow
    dependency: redis (session cache)
    failure: latency above 1s or connection timeouts
    expected: sessions fall back to the database, error rate unchanged
    detection: alert MyAppRedisLatencyHigh
    runbook: https://runbooks.example.com/my-app/redis
    test: toxiproxy latency toxic in staging
    last_tested: 2026-09-14
    result: passed
  - id: payment-provider-down
    dependency: payment provider API
    failure: HTTP 5xx or timeouts on all requests
    expected: checkout shows a retry message, orders stay in pending state
    detection: alert MyAppPaymentErrors
    runbook: https://runbooks.example.com/my-app/payment-provider
    test: not tested yet
    last_tested: null
    result: null

Each entry is a hypothesis: "when X fails, the service does Y, and alert Z shows it". An entry that has never been tested is an assumption, and the document makes that visible.

Test the hypotheses

Start in staging with simple tools. Toxiproxy sits between the service and a dependency and injects latency, timeouts or connection drops on demand:

# proxy local port 26379 to the real Redis
toxiproxy-cli create -l localhost:26379 -u localhost:6379 redis

# point my-app at localhost:26379, then add one second of latency
toxiproxy-cli toxic add -t latency -a latency=1000 redis

# watch error rate and latency on the dashboard, then remove the toxic
toxiproxy-cli toxic remove -n latency_downstream redis

Check the result against the expected field. If the service did what the document says, update last_tested. If it did not, you found a future incident ahead of time: file the fix and keep the entry marked as failed until it passes. Once an experiment is stable in staging, it can move to production with a limited blast radius, using tools like Chaos Mesh or AWS Fault Injection Service.

Close the loop

  • Every postmortem adds or updates failure-mode entries for the services involved.
  • Every pre-mortem adds the risks it found, with the agreed mitigation.
  • Untested entries and failed tests go into the same review as open postmortem actions.

Checklist

  • Postmortem actions are tickets with owners, due dates and a recurring review.
  • Incidents are tagged with contributing factors, and the tags are counted every quarter.
  • A pre-mortem runs before migrations, launches and hard-to-reverse changes.
  • Each service has a failure-mode document in its repository.
  • Each failure mode has an expected behavior, a detection method and a test.
  • Untested and failed entries are tracked like any other reliability work.