Skip to main content

Silences and grouping

On a platform that drains and replaces nodes, one upgrade can fire dozens of alerts at once, none of which need a human. Grouping turns those into a single notification, and a silence mutes the ones you expect while maintenance runs. Get both right, or the on-call rotation learns to ignore the pager.

Alertmanager groups alerts that share the labels you list in group_by, and sends one notification per group instead of one per alert:

alertmanager-config.yamlyaml
		route:
  group_by: ["alertname", "cluster", "namespace"]
  group_wait: 30s # wait this long to collect the first batch of a new group
  group_interval: 5m # wait this long before sending an update for a group
  repeat_interval: 4h # re-notify an unresolved group this often
	

group_wait gives a burst a moment to gather, so a whole pool going NotReady arrives as one notification, not twenty. group_interval paces updates to a group, and repeat_interval re-pages if it is still firing.

Silence maintenance

A node replacement or a cluster upgrade rolls nodes on purpose, and the alerts that follow (a node briefly NotReady, a pod rescheduling) are expected. Silence them for the window rather than paging on your own maintenance. A silence matches alert labels for a set duration. Add one with :

		$ amtool silence add \
  --alertmanager.url=http://localhost:9093 \
  --duration=2h --comment="cluster upgrade, ticket OPS-1234" \
  alertname=~"KubeNodeNotReady|KubePodNotReady" cluster="prod-eu"
	

Scope the silence tightly (the specific alerts and the specific cluster) so you do not blind yourself to a real problem elsewhere while maintenance runs. Set the duration to a little longer than the planned window, and let it expire on its own.

Tip

Before a planned or , add the silence, do the work, then confirm it expired. amtool makes this scriptable, so it can be a step in your maintenance runbook.

Inhibition

Inhibition suppresses one alert while another is firing, so one root cause does not page ten times. While a node's KubeNodeNotReady is firing, this rule mutes every warning alert that carries the same node label:

alertmanager-config.yamlyaml
		inhibit_rules:
  - source_matchers: ["alertname = KubeNodeNotReady"]
    target_matchers: ["severity = warning"]
    equal: ["node"]
	

Audit who silenced what

A silence hides real alerts, so keep a record of who added it and why. amtool fills in the author for you, from --author or your username. The reason is yours to write: put the maintenance ticket number in --comment, for example --comment="cluster upgrade, ticket OPS-1234".

Run amtool silence query from time to time. It lists the active silences with the author and comment of each, which is how you find a silence that was forgotten and is still hiding alerts.