Platform alert rules
A health daemon runs on every node and reports problems as node conditions.
On cloud pools a broken node is rebooted and, if that does not help, replaced from a fresh image. On bare metal only a Ready=False node is acted on: the server is rebooted, and if that does not help it is re-provisioned on the same physical machine. Every other problem there, a tampered node, a failed integrity check, a failing disk, is reported but not fixed automatically, so an alert is the only way you act on it.
This page ships those rules, plus the capacity and control-plane rules every cluster should run. You alert on a condition by reading kube_node_status_condition from kube-state-metrics , with node-exporter supplying the capacity metrics. Set both up first.
Node integrity
VerityCorruption, SealedOSTampered, and NodeTampered are the integrity conditions the daemon publishes on every node. Only VerityCorruption is fixed automatically, and only on cloud nodes. The other two are never fixed on any pool. Alert on all three conditions:
apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
name: syself-node-integrity
namespace: monitoring
labels:
prometheus: main
spec:
groups:
- name: syself-node-integrity
rules:
- alert: NodeVerityCorruption
expr: kube_node_status_condition{condition="VerityCorruption",status="true"} == 1
for: 1m
labels: {severity: critical}
annotations:
summary: "sealed OS integrity check failed on {{ $labels.node }}"
description: "A block on the sealed OS failed its integrity check. On bare-metal pools this does not self-heal, so act on it."
- alert: NodeSealedOSTampered
expr: kube_node_status_condition{condition="SealedOSTampered",status="true"} == 1
for: 1m
labels: {severity: critical}
annotations:
summary: "sealed OS tampering on {{ $labels.node }}"
description: "An unauthorized change was detected in the read-only sealed OS image. This never self-heals on any pool, so investigate the node."
- alert: NodeTampered
expr: kube_node_status_condition{condition="NodeTampered",status="true"} == 1
for: 1m
labels: {severity: critical}
annotations:
summary: "node tampering detected on {{ $labels.node }}"
description: "A protected file changed since the boot baseline. This latches; a revert does not clear it. Investigate the node."
Important
NodeTampered stays set once it fires. Reverting the changed file does not clear it, and neither does waiting. Only replacing the node does, so treat the alert as a prompt to investigate.
Ready and the daemon conditions
The daemon sets many more conditions than the three above, DisksFailure and CNIUnhealthy among them. Add a rule for each, in the same shape: kube_node_status_condition{condition="...",status="true"} == 1. Alert on Ready too, so a node that goes NotReady surfaces even when nothing else does. Node health conditions lists them all.
Capacity
Catch a cluster running out of room before a workload does:
- CPU and memory: alert when almost all of a node's schedulable CPU or memory is reserved, or when node-exporter shows sustained high memory with little available.
- Disk: alert on a filesystem past ~85% (a
/varfilling with logs is the classic one), using node-exporter'snode_filesystem_avail_bytes. - PVC full: alert when a volume's used space approaches its capacity, from the kubelet volume metrics.
Control plane
- etcd: alert on rising fsync latency, a growing database size approaching its quota, and frequent leader changes. These are the early signs of a control plane under strain, from the loopback etcd scrape.
- Saturation: alert on controller-manager and scheduler work-queue depth staying high, which means reconciles are backing up.
The collection tier itself
Every rule on this page depends on the collectors still running, so alert on them too. A blind spot here is worse than a broken alert, because a silent collector looks exactly like a healthy cluster: no alerts fire, and the graphs simply stop.
An agent that stops reporting and a remote-write that is failing are the first two to cover. Meta monitoring carries the metrics and the expression for both. Three more are worth adding:
- The write-ahead log keeps growing. Sustained growth means the ingest rate exceeds what the backend will accept, and the buffer is deferring the loss rather than preventing it.
- Active series jump. A steep climb in series is a cardinality spike in progress, usually from one workload. Alerting on the rate of change catches it before the agent is killed, and Troubleshoot the collectors covers finding the source.
- A target is down on the nodes where it should exist. Scope these per node role. A rule that expects etcd on every node fires forever on workers and teaches everyone to ignore it, which is worse than having no rule.
Certificate expiry
Alert before a certificate expires, not after. Two sources: the health daemon's CertRenewalFailing node condition for node-level certificates, and, for your ingress certificates, cert-manager's expiry metric. To catch a node certificate before renewal ever fails, read the autopilot.syself.com/certs annotation on the Node, which lists each certificate with its notAfter date. See Rotate and renew certificates .
Keep all of these in Git as PrometheusRule objects so a rebuilt cluster comes back with the same alerts, and route the critical ones to a pager. See Alert routing and receivers .
Alert routing and receivers
Route alerts by severity, team, and cluster to Slack, PagerDuty, email, or a webhook, so the right person is paged and clients are separated.
SLOs and error budgets
Define service-level objectives on your own metrics and alert on burn rate instead of raw thresholds, so pages mean a real customer problem.