Writing Runbooks People Can Actually Follow
A runbook gets read by someone who is tired, under pressure, and often not the person who wrote it. If following it requires knowledge that only lives in a senior engineer's head, it is not a runbook yet. This post covers how to structure one, how to get people to it from the alert, how to turn the common fixes into buttons, and how to keep the whole set current.
One paging alert, one runbook
Every alert that can wake someone up links to a runbook. In Prometheus the convention is a runbook_url annotation, which most alert routing setups can pass through to Slack or the paging tool:
- alert: MyAppHighErrorRate
expr: |
sum(rate(http_requests_total{job="my-app", code=~"5.."}[5m]))
/ sum(rate(http_requests_total{job="my-app"}[5m])) > 0.05
for: 10m
labels:
severity: page
annotations:
summary: "my-app returns more than 5% errors"
runbook_url: "https://runbooks.example.com/my-app/high-error-rate"
A check in CI keeps that rule from eroding. This one uses yq v4 to list alerting rules without a runbook link:
yq '.groups[].rules[] | select(has("alert") and .annotations.runbook_url == null) | .alert' rules/*.yaml
Fail the pipeline if the output is not empty.
Structure
Use the same layout for every runbook, so people know where to look:
# MyAppHighErrorRate
What it means: more than 5% of my-app requests return 5xx for 10 minutes.
Impact: users see errors on checkout.
Dashboard: https://grafana.example.com/d/my-app?var-env=production
Access needed: kubectl access to the prod cluster, Grafana viewer.
Escalation: #team-payments, then the secondary on-call.
## 1. Check for a recent deploy
Look at the latest my-app deploy in CI, or run:
kubectl -n my-app rollout history deployment/my-app
If a deploy finished shortly before the alert, go to step 4.
## 2. Check pod health
kubectl -n my-app get pods -l app=my-app
Expected: all pods Running and READY. If pods are restarting, check logs:
kubectl -n my-app logs deploy/my-app --previous --tail=100
## 3. Check dependencies
Database and payment provider panels on the dashboard above.
If the payment provider is failing, follow the payment-provider runbook.
## 4. Roll back (safe, reversible)
kubectl -n my-app rollout undo deployment/my-app
kubectl -n my-app rollout status deployment/my-app --timeout=5m
## Verify
Error rate on the dashboard is back under 5% for 10 minutes.
## If nothing above helped
Escalate. Do not restart the database.
Rules for the content:
- Exact commands with real names. Namespace, deployment, cluster context. "Restart the app server" is not a step.
- Expected output after each step. The reader needs to know whether the step worked before moving on.
- Explicit decision points. "If X, go to step 4" instead of a paragraph of possibilities.
- Mark destructive steps and put a check before them.
- List the access needed at the top. Finding out at step 3 that you lack permissions wastes the most time.
- Link dashboards with variables already set, so the reader lands on the right environment.
Make the common fixes clickable
If the same command is run during most incidents, turn it into a job with an audit trail instead of a command people paste. A GitHub Actions workflow with workflow_dispatch gives you a button, inputs, a log, and approvals through environments:
name: restart-my-app
on:
workflow_dispatch:
inputs:
environment:
type: choice
options: [staging, production]
required: true
jobs:
restart:
runs-on: ubuntu-latest
environment: ${{ inputs.environment }}
steps:
- name: Restart and wait for rollout
env:
KUBECONFIG_DATA: ${{ secrets.KUBECONFIG }}
run: |
echo "$KUBECONFIG_DATA" > "$RUNNER_TEMP/kubeconfig"
export KUBECONFIG="$RUNNER_TEMP/kubeconfig"
kubectl -n my-app rollout restart deployment/my-app
kubectl -n my-app rollout status deployment/my-app --timeout=5m
Store KUBECONFIG as an environment secret, so staging and production use different credentials. Add required reviewers on the production environment if the action should need a second person. The runbook step then becomes a link to the workflow. Rundeck, AWX or any other job runner works the same way.
Where runbooks live
Keep them in git, next to the alert rules or the service code, and publish them as a static site. Then:
- a change to an alert and its runbook go through the same pull request,
- history shows who changed a step and why,
- a link checker such as lychee can run in CI and catch dead dashboard links.
A wiki works too, as long as it has an owner and review. Runbooks without either drift fastest.
Keep them current
- Owner and last-reviewed date at the top of every runbook.
- Fix it during the incident. The responder notes every wrong or missing step, and correcting the runbook is an action item in the postmortem.
- Test with someone new. During a game day in staging, ask the newest team member to follow the runbook without help. Every question they ask is a missing line.
- Delete runbooks for alerts that no longer exist. A stale runbook is worse than none, because people trust it.
Checklist
- Every paging alert has a
runbook_url, enforced in CI. - Same structure everywhere: meaning, impact, dashboard, access, escalation, steps, verify.
- Exact commands, expected output, explicit decision points.
- Frequent fixes are jobs with a button, a log and approvals.
- Runbooks live in git with owners and review dates.
- Runbook fixes are postmortem action items, and new team members test them.
