Alertmanager's routing tree decides which alert goes to whom. You match on the labels an alert carries, `severity`, a team label, a `cluster` label, and send each branch to the right receiver. With the tree set up correctly, a page reaches the on-call for that service, and one client's alerts never land in another client's channel. ## The routing tree Routes are a tree: an alert enters at the root and walks down until it matches a branch. Match on labels, and put the most specific routes first. ```yaml title="alertmanager-config.yaml" route: receiver: default-email group_by: ["alertname", "cluster", "namespace"] routes: - matchers: ["severity = critical"] receiver: pagerduty - matchers: ["team = backend"] receiver: backend-slack - matchers: ["cluster = client-acme"] receiver: acme-slack ``` ## Receivers Each receiver is a destination. Add these receiver configurations to the same `alertmanager-config.yaml` file alongside your routes. Keep every credential in a Secret and reference it, never inline. - **Slack:** an incoming-webhook URL per channel. ```yaml title="alertmanager-config.yaml" receivers: - name: backend-slack slack_configs: - api_url_file: /etc/alertmanager/secrets/slack-backend/url channel: "#backend-alerts" ``` For that file to exist, list the Secret under `secrets` in the `Alertmanager` object from [Set up Alertmanager](/docs/hetzner/apalla/observability/alerting/set-up-alertmanager). The operator mounts each one at `/etc/alertmanager/secrets/`: ```yaml title="alertmanager.yaml" spec: secrets: - slack-backend ``` - **PagerDuty / Opsgenie:** a routing/integration key; use these for the `critical` branch that must page a human. - **Email (SMTP):** a `smtp_*` block for a mailbox; fine for warnings and digests. - **Generic webhook:** a `webhook_configs` URL for anything else, a ticketing system, a custom bot, a chatops handler. ## Per-client routing for agencies For an agency, the `cluster` (or a `client`) label on every alert is what keeps tenants apart. Each cluster remote-writes with its own external label (see [Multi-cluster observability](/docs/hetzner/apalla/observability/multi-cluster/multi-cluster-observability)), so central Alertmanager can route `cluster = client-acme` to the Acme channel and `cluster = client-beta` to Beta's, from one config. A client only ever sees their own alerts, and you see all of them. ## Test each route Routing mistakes are quiet ones. A matcher with a typo never matches anything, and the alerts it was meant to catch fall through to `default-email` without any error. Test the tree before you rely on it. `amtool` is Alertmanager's command-line tool. It ships at `/bin/amtool` in the Alertmanager image, so you can run it with `kubectl exec`. To check a config file before you apply it, get the binary for your own machine from the [Alertmanager release](https://github.com/prometheus/alertmanager/releases) tarball. Start with `check-config`, which catches a typo or a receiver you referenced but never defined: ```console $ amtool check-config alertmanager-config.yaml Checking 'alertmanager-config.yaml' SUCCESS Found: - global config - route - 0 inhibit rules - 4 receivers - 0 templates ``` `config routes test` takes the labels an alert would carry and prints the receiver it reaches: ```console $ amtool config routes test --config.file=alertmanager-config.yaml \ severity=critical cluster=client-acme pagerduty ``` This alert carries both labels, so it matches two branches, and Alertmanager takes the first one it finds. Since `severity = critical` sits above the cluster route, the alert pages `pagerduty` and Acme's channel never sees it. Drop the severity and the same alert goes where you would expect: ```console $ amtool config routes test --config.file=alertmanager-config.yaml \ severity=warning cluster=client-acme acme-slack ``` That may be exactly what you want, or the opposite. Either way you decide it by the order of the branches, so check the order before you rely on it. Running `amtool config routes show` against the same file prints the whole tree the way Alertmanager reads it. To test what is running rather than what you wrote, run the same command inside the pod. The operator renders your config into the container, so that copy is the live one: ```console $ kubectl -n monitoring exec alertmanager-main-0 \ -c alertmanager -- amtool config routes test \ --config.file=/etc/alertmanager/config_out/alertmanager.env.yaml \ severity=critical cluster=client-acme ``` Test the critical path first. If it is misconfigured, no one will get paged, and that is something you want to discover now and not during a real incident.