On a platform that drains and replaces nodes, one upgrade can fire dozens of alerts at once, none of which need a human. Grouping turns those into a single notification, and a silence mutes the ones you expect while maintenance runs. Get both right, or the on-call rotation learns to ignore the pager. ## Group related alerts Alertmanager groups alerts that share the labels you list in `group_by`, and sends one notification per group instead of one per alert: ```yaml title="alertmanager-config.yaml" route: group_by: ["alertname", "cluster", "namespace"] group_wait: 30s # wait this long to collect the first batch of a new group group_interval: 5m # wait this long before sending an update for a group repeat_interval: 4h # re-notify an unresolved group this often ``` `group_wait` gives a burst a moment to gather, so a whole pool going `NotReady` arrives as one notification, not twenty. `group_interval` paces updates to a group, and `repeat_interval` re-pages if it is still firing. ## Silence maintenance A node replacement or a cluster upgrade rolls nodes on purpose, and the alerts that follow (a node briefly `NotReady`, a pod rescheduling) are expected. Silence them for the window rather than paging on your own maintenance. A silence matches alert labels for a set duration. Add one with [`amtool`](/docs/hetzner/apalla/observability/alerting/alert-routing-and-receivers#test-each-route): ```console $ amtool silence add \ --alertmanager.url=http://localhost:9093 \ --duration=2h --comment="cluster upgrade, ticket OPS-1234" \ alertname=~"KubeNodeNotReady|KubePodNotReady" cluster="prod-eu" ``` Scope the silence tightly (the specific alerts and the specific cluster) so you do not blind yourself to a real problem elsewhere while maintenance runs. Set the duration to a little longer than the planned window, and let it expire on its own. > [!TIP] > Before a planned [node replacement](/docs/hetzner/apalla/concepts/operations/self-healing-and-node-replacement) or [cluster upgrade](/docs/hetzner/apalla/clusters/upgrades/upgrade-to-a-new-kubernetes-version), add the silence, do the work, then confirm it expired. `amtool` makes this scriptable, so it can be a step in your maintenance runbook. ## Inhibition Inhibition suppresses one alert while another is firing, so one root cause does not page ten times. While a node's `KubeNodeNotReady` is firing, this rule mutes every `warning` alert that carries the same `node` label: ```yaml title="alertmanager-config.yaml" inhibit_rules: - source_matchers: ["alertname = KubeNodeNotReady"] target_matchers: ["severity = warning"] equal: ["node"] ``` ## Audit who silenced what A silence hides real alerts, so keep a record of who added it and why. `amtool` fills in the author for you, from `--author` or your username. The reason is yours to write: put the maintenance ticket number in `--comment`, for example `--comment="cluster upgrade, ticket OPS-1234"`. Run `amtool silence query` from time to time. It lists the active silences with the author and comment of each, which is how you find a silence that was forgotten and is still hiding alerts.