Measuring On-Call Load Before It Burns People Out

Measuring On-Call Load Before It Burns People Out

Reading time1 min
#oncall#burnout#devops#metrics#engineering

Measuring On-Call Load Before It Burns People Out

Engineers rarely leave because of one bad night. They leave after months of broken sleep, alerts nobody fixes, and the sense that on-call is a tax only some of the team pays. All of that leaves a trace in data you already have: the paging tool's incident log and your alert rules. This post covers what to measure, how to pull it, and what to do with the result.

Start from the paging log

The paging tool knows what actually reached a human: when, at what urgency, and whether it escalated. Alert rule metrics only tell you what fired. Every paging tool has an API or a CSV export, so pull the last few months of incidents and work from that.

Metrics worth tracking

Per rotation and per week:

  • Pages per shift. Only high-urgency notifications that interrupt someone.
  • Off-hours pages. Pages outside working hours in the responder's local time. These cost the most.
  • Interrupted nights. Count nights with at least one page, not pages. Three pages in one night and one page on three nights feel very different.
  • Actionable ratio. The share of pages where the responder had to do something. This needs a convention, such as a tag or a resolution note, set when the page is resolved.
  • Repeat offenders. The same alert firing again and again. These are the cheapest wins.
  • Time to acknowledge. A rising trend often means people are tired or have started ignoring the pager.
  • Escalations. How often the secondary gets paged because the primary did not answer.
  • Load distribution. Whether a few people carry most of the pages, usually because they are the only ones who know a system.

You need a definition of "too much". Google's SRE book gives two reference points: no more than 25 percent of an SRE's time on on-call, and at most two incidents per 12-hour shift, based on its estimate that an incident with follow-up work takes about six hours. Your limits can differ, but write them down.

Pulling the data

An example with the PagerDuty REST API. It fetches high-urgency incidents for the last 30 days, page by page:

#!/usr/bin/env bash
set -euo pipefail
since=$(date -u -d '30 days ago' +%Y-%m-%dT%H:%M:%SZ)  # macOS: date -u -v-30d ...
offset=0
: > incidents.jsonl
while :; do
  page=$(curl -sf -G https://api.pagerduty.com/incidents \
    -H "Authorization: Token token=${PD_TOKEN}" \
    -H "Accept: application/vnd.pagerduty+json;version=2" \
    --data-urlencode "since=${since}" \
    -d "urgencies[]=high" -d "limit=100" -d "offset=${offset}")
  echo "$page" | jq -c '.incidents[]' >> incidents.jsonl
  [ "$(echo "$page" | jq -r '.more')" = "true" ] || break
  offset=$((offset + 100))
done

Then answer the basic questions with jq:

# pages by hour of day in the team's time zone
TZ=Europe/Berlin jq -r '.created_at | fromdateiso8601 | strflocaltime("%H")' incidents.jsonl \
  | sort | uniq -c

# noisiest alerts
jq -r '.title' incidents.jsonl | sort | uniq -c | sort -rn | head -20

For per-person numbers, join the incident timestamps with the on-call schedule for the same period. The person on call when the page was created is the one who got woken up. Put the weekly results on a dashboard or in the handoff notes so the trend is visible, not just the latest week.

Reducing the load

Measurement is only useful if it changes the alerts.

  • Page only on symptoms that need a human now: user-facing errors, latency, or saturation that will cause an outage soon. Everything else goes to a ticket queue or a chat channel.
  • Use for: so short spikes do not page. If an alert flaps between firing and resolved, keep_firing_for in Prometheus keeps it active for a while after the condition clears.
  • Every paging alert has an owner and a runbook link.
  • Review the noisiest alerts at every on-call handoff: fix the cause, retune, downgrade to a ticket, or delete. The SRE book's guideline is to aim for one alert per incident.

A paging rule written that way:

groups:
  - name: my-app
    rules:
      - alert: MyAppHighErrorRate
        expr: |
          sum(rate(http_requests_total{job="my-app", code=~"5.."}[5m]))
            / sum(rate(http_requests_total{job="my-app"}[5m])) > 0.05
        for: 10m
        labels:
          severity: page
        annotations:
          summary: "my-app returns more than 5% errors for 10 minutes"
          runbook_url: "https://runbooks.example.com/my-app/high-error-rate"

Compare that with an alert on host CPU above some threshold over a single evaluation period. High CPU on its own is not user impact, and one period is enough for any spike to fire it. Alerts like that are a common source of pages that nobody acts on.

The human side

Numbers miss things, so ask directly:

  • A short anonymous survey at the end of each shift: how many times were you woken up, how rested do you feel, what should not have paged you.
  • Time off after a rough night or a rough week, written into the on-call policy instead of left to each manager.
  • A rotation large enough that each person's turn does not come around too often. If the team is too small, fix the rotation before tuning alerts.
  • Compensation for on-call, in money or time, agreed up front.
  • Fix the knowledge gap behind uneven load. If only one person can handle an alert, pair others with them and write the runbook.

Checklist

  • Export the paging log regularly and keep the history.
  • Track pages per shift, off-hours pages, interrupted nights, actionable ratio, repeat alerts, time to acknowledge and escalations.
  • Write down what "too much" means for your team.
  • Review the noisiest alerts at every handoff and act on them.
  • Page on symptoms, with for:, an owner and a runbook link.
  • Run a short anonymous survey per shift and give recovery time after bad nights.