Reducing Alert Noise: Measure, Delete, Group
Alert fatigue starts when most notifications need no action. People learn that the channel is noise, mute it, and the one alert that matters gets the same treatment as the hundred that did not. Adding more monitoring makes it worse. The fix is fewer, better alerts and a routing setup that removes duplicates.
Measure before changing anything
You need a list of what actually fires and how often. Good sources:
- The incident export from your paging tool (PagerDuty, Opsgenie, Grafana OnCall).
- A webhook receiver in Alertmanager that writes every notification to a log.
- The
ALERTSseries that Prometheus keeps for every pending and firing alert.
This query ranks alerts by how long they were firing over the last week:
sort_desc(
sum by (alertname) (count_over_time(ALERTS{alertstate="firing"}[7d]))
)
Each sample is one rule evaluation where the alert was firing, so the number is proportional to time spent firing. Alerts at the top of this list that nobody acts on are your first candidates.
Then review every alert that notified a human in the last few weeks. For each one, write down:
- Did someone have to do something?
- Did it have to happen now, or could it wait until working hours?
- Was it a duplicate of another alert for the same problem?
An alert that fails the first question should be deleted or turned into a dashboard panel. One that fails the second becomes a ticket, not a page. One that fails the third needs grouping or inhibition.
Page on symptoms, not causes
High CPU, a restarted container, a full disk at 70%, a pending pod: these are causes. Many of them never affect users. Page on what users notice: errors, latency and availability of the main flows. Keep causes on dashboards, where you look at them while debugging a symptom.
groups:
- name: my-app-symptoms
rules:
- alert: MyAppHighErrorRatio
expr: |
sum(rate(http_requests_total{job="my-app",code=~"5.."}[5m]))
/
sum(rate(http_requests_total{job="my-app"}[5m]))
> 0.02
for: 10m
keep_firing_for: 5m
labels:
severity: page
team: payments
annotations:
summary: "my-app 5xx ratio above 2% for 10 minutes"
runbook_url: "https://runbooks.example.com/my-app/high-error-ratio"
dashboard: "https://grafana.example.com/d/my-app"
The threshold here is a placeholder. A better version is an SLO burn-rate alert: page when the error budget is burning fast, open a ticket when it is burning slowly. The Google SRE Workbook chapter on alerting on SLOs describes the multi-window setup, and tools like Sloth and Pyrra generate the Prometheus rules for you.
Stop flapping
Two rule settings handle most flapping:
forkeeps the alert pending until the condition has been true for that long. Short spikes never notify.keep_firing_for(Prometheus 2.42 and later) keeps the alert firing for a while after the condition clears, so a metric hovering around the threshold does not resolve and re-fire every few minutes.
Let Alertmanager group, inhibit and route
route:
receiver: team-slack
group_by: [alertname, cluster, namespace]
group_wait: 30s
group_interval: 5m
repeat_interval: 4h
routes:
- matchers:
- severity = "page"
receiver: oncall-pagerduty
- matchers:
- severity = "ticket"
receiver: ticket-webhook
repeat_interval: 24h
inhibit_rules:
- source_matchers:
- alertname = "ClusterUnreachable"
target_matchers:
- severity =~ "page|ticket"
equal: [cluster]
receivers:
- name: team-slack
slack_configs:
- channel: "#alerts-my-team"
api_url_file: /etc/alertmanager/secrets/slack-webhook-url
- name: oncall-pagerduty
pagerduty_configs:
- routing_key_file: /etc/alertmanager/secrets/pagerduty-routing-key
- name: ticket-webhook
webhook_configs:
- url: http://alert-to-ticket.monitoring.svc:8080/hook
What each part does:
group_byturns fifty alerts from the same incident into one notification. Group by the labels that identify a problem, not byinstanceorpod.group_wait,group_interval,repeat_intervalare shown with their default values. Raiserepeat_intervalfor low-severity routes so open tickets do not get reminders every few hours.- Inhibition mutes alerts while a bigger one is active. Here, if a cluster is unreachable, every other alert from that cluster is noise. An alert that matches both sides of the rule cannot inhibit itself.
- Routing by severity sends pages to the pager and everything else to places that do not wake people up.
Silences are for planned work and known issues. Give them an expiry and a comment, and do not use them as a permanent replacement for fixing or deleting an alert.
One path to the pager
When CloudWatch alarms, Grafana alerts, a SaaS APM and Prometheus all notify on their own, the same failure produces several pages from different tools. Pick one source of truth for each signal and send everything through one router, either Alertmanager or your incident tool with deduplication keys configured. If two tools define the same condition, delete one.
Make every page actionable
Every paging alert should answer three questions in the notification itself: what is broken, how bad it is, and where to start. In practice that means:
- a summary with the affected service and the measured value,
- a
runbook_urlwith concrete first steps, - a dashboard link scoped to the service,
- an owner label (
team) that matches the route.
If a runbook says "ignore this when X", put X into the alert expression instead.
Keep it clean
Noise comes back as services change. Review the pages from the last on-call shift at every handover. Treat a page that needed no action as a bug, with an owner and a fix: delete the alert, raise the threshold, add for, or move it to a ticket route.
Checklist
- Export a list of what fired in the last weeks and how often.
- Delete alerts nobody acts on. Downgrade non-urgent ones to tickets.
- Page on user-facing symptoms, preferably SLO burn rate.
- Use
forandkeep_firing_foragainst flapping. group_byon incident-level labels, inhibition for cascades, routing by severity.- One router to the pager, one tool per signal.
- Runbook, dashboard and owner on every paging alert.
- Review pages at every on-call handover.
