Alert routing and receivers
Alertmanager's routing tree decides which alert goes to whom. You match on the labels an alert carries, severity, a team label, a cluster label, and send each branch to the right receiver. With the tree set up correctly, a page reaches the on-call for that service, and one client's alerts never land in another client's channel.
The routing tree
Routes are a tree: an alert enters at the root and walks down until it matches a branch. Match on labels, and put the most specific routes first.
route:
receiver: default-email
group_by: ["alertname", "cluster", "namespace"]
routes:
- matchers: ["severity = critical"]
receiver: pagerduty
- matchers: ["team = backend"]
receiver: backend-slack
- matchers: ["cluster = client-acme"]
receiver: acme-slack
Receivers
Each receiver is a destination. Add these receiver configurations to the same alertmanager-config.yaml file alongside your routes. Keep every credential in a Secret and reference it, never inline.
Slack: an incoming-webhook URL per channel.
alertmanager-config.yamlyaml receivers: - name: backend-slack slack_configs: - api_url_file: /etc/alertmanager/secrets/slack-backend/url channel: "#backend-alerts"For that file to exist, list the Secret under
secretsin theAlertmanagerobject from Set up Alertmanager . The operator mounts each one at/etc/alertmanager/secrets/<secret-name>:alertmanager.yamlyaml spec: secrets: - slack-backendPagerDuty / Opsgenie: a routing/integration key; use these for the
criticalbranch that must page a human.Email (SMTP): a
smtp_*block for a mailbox; fine for warnings and digests.Generic webhook: a
webhook_configsURL for anything else, a ticketing system, a custom bot, a chatops handler.
Per-client routing for agencies
For an agency, the cluster (or a client) label on every alert is what keeps tenants apart. Each cluster remote-writes with its own external label (see Multi-cluster observability ), so central Alertmanager can route cluster = client-acme to the Acme channel and cluster = client-beta to Beta's, from one config. A client only ever sees their own alerts, and you see all of them.
Test each route
Routing mistakes are quiet ones. A matcher with a typo never matches anything, and the alerts it was meant to catch fall through to default-email without any error. Test the tree before you rely on it.
amtool is Alertmanager's command-line tool. It ships at /bin/amtool in the Alertmanager image, so you can run it with kubectl exec. To check a config file before you apply it, get the binary for your own machine from the Alertmanager release tarball.
Start with check-config, which catches a typo or a receiver you referenced but never defined:
$ amtool check-config alertmanager-config.yaml
Checking 'alertmanager-config.yaml' SUCCESS
Found:
- global config
- route
- 0 inhibit rules
- 4 receivers
- 0 templates
config routes test takes the labels an alert would carry and prints the receiver it reaches:
$ amtool config routes test --config.file=alertmanager-config.yaml \
severity=critical cluster=client-acme
pagerduty
This alert carries both labels, so it matches two branches, and Alertmanager takes the first one it finds. Since severity = critical sits above the cluster route, the alert pages pagerduty and Acme's channel never sees it. Drop the severity and the same alert goes where you would expect:
$ amtool config routes test --config.file=alertmanager-config.yaml \
severity=warning cluster=client-acme
acme-slack
That may be exactly what you want, or the opposite. Either way you decide it by the order of the branches, so check the order before you rely on it. Running amtool config routes show against the same file prints the whole tree the way Alertmanager reads it.
To test what is running rather than what you wrote, run the same command inside the pod. The operator renders your config into the container, so that copy is the live one:
$ kubectl -n monitoring exec alertmanager-main-0 \
-c alertmanager -- amtool config routes test \
--config.file=/etc/alertmanager/config_out/alertmanager.env.yaml \
severity=critical cluster=client-acme
Test the critical path first. If it is misconfigured, no one will get paged, and that is something you want to discover now and not during a real incident.
Set up Alertmanager
Deploy Alertmanager as its own object, hold its config in a Secret, load rules with PrometheusRule objects, and confirm an alert flows end to end.
Platform alert rules
Some conditions never trigger self-healing on bare-metal pools, so alerting is the only way to act on node integrity and capacity, and this page ships those rules.