Reducing Alert Noise: Measure, Delete, Group

Reducing Alert Noise: Measure, Delete, Group

Reading time1 min
#devops#observability#alerts#monitoring

Reducing Alert Noise: Measure, Delete, Group

Alert fatigue starts when most notifications need no action. People learn that the channel is noise, mute it, and the one alert that matters gets the same treatment as the hundred that did not. Adding more monitoring makes it worse. The fix is fewer, better alerts and a routing setup that removes duplicates.

Measure before changing anything

You need a list of what actually fires and how often. Good sources:

  • The incident export from your paging tool (PagerDuty, Opsgenie, Grafana OnCall).
  • A webhook receiver in Alertmanager that writes every notification to a log.
  • The ALERTS series that Prometheus keeps for every pending and firing alert.

This query ranks alerts by how long they were firing over the last week:

sort_desc(
  sum by (alertname) (count_over_time(ALERTS{alertstate="firing"}[7d]))
)

Each sample is one rule evaluation where the alert was firing, so the number is proportional to time spent firing. Alerts at the top of this list that nobody acts on are your first candidates.

Then review every alert that notified a human in the last few weeks. For each one, write down:

  1. Did someone have to do something?
  2. Did it have to happen now, or could it wait until working hours?
  3. Was it a duplicate of another alert for the same problem?

An alert that fails the first question should be deleted or turned into a dashboard panel. One that fails the second becomes a ticket, not a page. One that fails the third needs grouping or inhibition.

Page on symptoms, not causes

High CPU, a restarted container, a full disk at 70%, a pending pod: these are causes. Many of them never affect users. Page on what users notice: errors, latency and availability of the main flows. Keep causes on dashboards, where you look at them while debugging a symptom.

groups:
  - name: my-app-symptoms
    rules:
      - alert: MyAppHighErrorRatio
        expr: |
          sum(rate(http_requests_total{job="my-app",code=~"5.."}[5m]))
          /
          sum(rate(http_requests_total{job="my-app"}[5m]))
          > 0.02
        for: 10m
        keep_firing_for: 5m
        labels:
          severity: page
          team: payments
        annotations:
          summary: "my-app 5xx ratio above 2% for 10 minutes"
          runbook_url: "https://runbooks.example.com/my-app/high-error-ratio"
          dashboard: "https://grafana.example.com/d/my-app"

The threshold here is a placeholder. A better version is an SLO burn-rate alert: page when the error budget is burning fast, open a ticket when it is burning slowly. The Google SRE Workbook chapter on alerting on SLOs describes the multi-window setup, and tools like Sloth and Pyrra generate the Prometheus rules for you.

Stop flapping

Two rule settings handle most flapping:

  • for keeps the alert pending until the condition has been true for that long. Short spikes never notify.
  • keep_firing_for (Prometheus 2.42 and later) keeps the alert firing for a while after the condition clears, so a metric hovering around the threshold does not resolve and re-fire every few minutes.

Let Alertmanager group, inhibit and route

route:
  receiver: team-slack
  group_by: [alertname, cluster, namespace]
  group_wait: 30s
  group_interval: 5m
  repeat_interval: 4h
  routes:
    - matchers:
        - severity = "page"
      receiver: oncall-pagerduty
    - matchers:
        - severity = "ticket"
      receiver: ticket-webhook
      repeat_interval: 24h

inhibit_rules:
  - source_matchers:
      - alertname = "ClusterUnreachable"
    target_matchers:
      - severity =~ "page|ticket"
    equal: [cluster]

receivers:
  - name: team-slack
    slack_configs:
      - channel: "#alerts-my-team"
        api_url_file: /etc/alertmanager/secrets/slack-webhook-url
  - name: oncall-pagerduty
    pagerduty_configs:
      - routing_key_file: /etc/alertmanager/secrets/pagerduty-routing-key
  - name: ticket-webhook
    webhook_configs:
      - url: http://alert-to-ticket.monitoring.svc:8080/hook

What each part does:

  • group_by turns fifty alerts from the same incident into one notification. Group by the labels that identify a problem, not by instance or pod.
  • group_wait, group_interval, repeat_interval are shown with their default values. Raise repeat_interval for low-severity routes so open tickets do not get reminders every few hours.
  • Inhibition mutes alerts while a bigger one is active. Here, if a cluster is unreachable, every other alert from that cluster is noise. An alert that matches both sides of the rule cannot inhibit itself.
  • Routing by severity sends pages to the pager and everything else to places that do not wake people up.

Silences are for planned work and known issues. Give them an expiry and a comment, and do not use them as a permanent replacement for fixing or deleting an alert.

One path to the pager

When CloudWatch alarms, Grafana alerts, a SaaS APM and Prometheus all notify on their own, the same failure produces several pages from different tools. Pick one source of truth for each signal and send everything through one router, either Alertmanager or your incident tool with deduplication keys configured. If two tools define the same condition, delete one.

Make every page actionable

Every paging alert should answer three questions in the notification itself: what is broken, how bad it is, and where to start. In practice that means:

  • a summary with the affected service and the measured value,
  • a runbook_url with concrete first steps,
  • a dashboard link scoped to the service,
  • an owner label (team) that matches the route.

If a runbook says "ignore this when X", put X into the alert expression instead.

Keep it clean

Noise comes back as services change. Review the pages from the last on-call shift at every handover. Treat a page that needed no action as a bug, with an owner and a fix: delete the alert, raise the threshold, add for, or move it to a ticket route.

Checklist

  • Export a list of what fired in the last weeks and how often.
  • Delete alerts nobody acts on. Downgrade non-urgent ones to tickets.
  • Page on user-facing symptoms, preferably SLO burn rate.
  • Use for and keep_firing_for against flapping.
  • group_by on incident-level labels, inhibition for cascades, routing by severity.
  • One router to the pager, one tool per signal.
  • Runbook, dashboard and owner on every paging alert.
  • Review pages at every on-call handover.