Silences and grouping
On a platform that drains and replaces nodes, one upgrade can fire dozens of alerts at once, none of which need a human. Grouping turns those into a single notification, and a silence mutes the ones you expect while maintenance runs. Get both right, or the on-call rotation learns to ignore the pager.
Group related alerts
Alertmanager groups alerts that share the labels you list in group_by, and sends one notification per group instead of one per alert:
route:
group_by: ["alertname", "cluster", "namespace"]
group_wait: 30s # wait this long to collect the first batch of a new group
group_interval: 5m # wait this long before sending an update for a group
repeat_interval: 4h # re-notify an unresolved group this often
group_wait gives a burst a moment to gather, so a whole pool going NotReady arrives as one notification, not twenty. group_interval paces updates to a group, and repeat_interval re-pages if it is still firing.
Silence maintenance
A node replacement or a cluster upgrade rolls nodes on purpose, and the alerts that follow (a node briefly NotReady, a pod rescheduling) are expected. Silence them for the window rather than paging on your own maintenance. A silence matches alert labels for a set duration. Add one with amtool :
$ amtool silence add \
--alertmanager.url=http://localhost:9093 \
--duration=2h --comment="cluster upgrade, ticket OPS-1234" \
alertname=~"KubeNodeNotReady|KubePodNotReady" cluster="prod-eu"
Scope the silence tightly (the specific alerts and the specific cluster) so you do not blind yourself to a real problem elsewhere while maintenance runs. Set the duration to a little longer than the planned window, and let it expire on its own.
Tip
Before a planned node replacement or cluster upgrade , add the silence, do the work, then confirm it expired. amtool makes this scriptable, so it can be a step in your maintenance runbook.
Inhibition
Inhibition suppresses one alert while another is firing, so one root cause does not page ten times. While a node's KubeNodeNotReady is firing, this rule mutes every warning alert that carries the same node label:
inhibit_rules:
- source_matchers: ["alertname = KubeNodeNotReady"]
target_matchers: ["severity = warning"]
equal: ["node"]
Audit who silenced what
A silence hides real alerts, so keep a record of who added it and why. amtool fills in the author for you, from --author or your username. The reason is yours to write: put the maintenance ticket number in --comment, for example --comment="cluster upgrade, ticket OPS-1234".
Run amtool silence query from time to time. It lists the active silences with the author and comment of each, which is how you find a silence that was forgotten and is still hiding alerts.
SLOs and error budgets
Define service-level objectives on your own metrics and alert on burn rate instead of raw thresholds, so pages mean a real customer problem.
See flows with Hubble
Hubble is on by default and records every connection and drop, so scrape its metrics and use the CLI to watch flows live.