A health daemon runs on every node and reports problems as node conditions. On cloud pools a broken node is rebooted and, if that does not help, replaced from a fresh image. On bare metal only a `Ready=False` node is acted on: the server is rebooted, and if that does not help it is re-provisioned on the same physical machine. Every other problem there, a tampered node, a failed integrity check, a failing disk, is reported but not fixed automatically, so an alert is the only way you act on it. This page ships those rules, plus the capacity and control-plane rules every cluster should run. You alert on a condition by reading `kube_node_status_condition` from [kube-state-metrics](/docs/hetzner/apalla/observability/metrics/kube-state-metrics), with node-exporter supplying the capacity metrics. Set both up first. ## Node integrity `VerityCorruption`, `SealedOSTampered`, and `NodeTampered` are the integrity conditions the daemon publishes on every node. Only `VerityCorruption` is fixed automatically, and only on cloud nodes. The other two are never fixed on any pool. Alert on all three conditions: ```yaml title="node-integrity-rules.yaml" apiVersion: monitoring.coreos.com/v1 kind: PrometheusRule metadata: name: syself-node-integrity namespace: monitoring labels: prometheus: main spec: groups: - name: syself-node-integrity rules: - alert: NodeVerityCorruption expr: kube_node_status_condition{condition="VerityCorruption",status="true"} == 1 for: 1m labels: {severity: critical} annotations: summary: "sealed OS integrity check failed on {{ $labels.node }}" description: "A block on the sealed OS failed its integrity check. On bare-metal pools this does not self-heal, so act on it." - alert: NodeSealedOSTampered expr: kube_node_status_condition{condition="SealedOSTampered",status="true"} == 1 for: 1m labels: {severity: critical} annotations: summary: "sealed OS tampering on {{ $labels.node }}" description: "An unauthorized change was detected in the read-only sealed OS image. This never self-heals on any pool, so investigate the node." - alert: NodeTampered expr: kube_node_status_condition{condition="NodeTampered",status="true"} == 1 for: 1m labels: {severity: critical} annotations: summary: "node tampering detected on {{ $labels.node }}" description: "A protected file changed since the boot baseline. This latches; a revert does not clear it. Investigate the node." ``` > [!IMPORTANT] > `NodeTampered` stays set once it fires. Reverting the changed file does not clear it, and neither does waiting. Only replacing the node does, so treat the alert as a prompt to investigate. ## Ready and the daemon conditions The daemon sets many more conditions than the three above, `DisksFailure` and `CNIUnhealthy` among them. Add a rule for each, in the same shape: `kube_node_status_condition{condition="...",status="true"} == 1`. Alert on `Ready` too, so a node that goes `NotReady` surfaces even when nothing else does. [Node health conditions](/docs/hetzner/apalla/observability/reference/node-health-conditions) lists them all. ## Capacity Catch a cluster running out of room before a workload does: - **CPU and memory:** alert when almost all of a node's schedulable CPU or memory is reserved, or when node-exporter shows sustained high memory with little available. - **Disk:** alert on a filesystem past ~85% (a `/var` filling with logs is the classic one), using node-exporter's `node_filesystem_avail_bytes`. - **PVC full:** alert when a volume's used space approaches its capacity, from the kubelet volume metrics. ## Control plane - **etcd:** alert on rising fsync latency, a growing database size approaching its quota, and frequent leader changes. These are the early signs of a control plane under strain, from the loopback etcd scrape. - **Saturation:** alert on controller-manager and scheduler work-queue depth staying high, which means reconciles are backing up. ## The collection tier itself Every rule on this page depends on the collectors still running, so alert on them too. A blind spot here is worse than a broken alert, because a silent collector looks exactly like a healthy cluster: no alerts fire, and the graphs simply stop. An agent that stops reporting and a remote-write that is failing are the first two to cover. [Meta monitoring](/docs/hetzner/apalla/observability/alerting/set-up-alertmanager#meta-monitoring) carries the metrics and the expression for both. Three more are worth adding: - **The write-ahead log keeps growing.** Sustained growth means the ingest rate exceeds what the backend will accept, and the buffer is deferring the loss rather than preventing it. - **Active series jump.** A steep climb in series is a cardinality spike in progress, usually from one workload. Alerting on the rate of change catches it before the agent is killed, and [Troubleshoot the collectors](/docs/hetzner/apalla/observability/collection/troubleshoot-the-collectors) covers finding the source. - **A target is down on the nodes where it should exist.** Scope these per node role. A rule that expects etcd on every node fires forever on workers and teaches everyone to ignore it, which is worse than having no rule. ## Certificate expiry Alert before a certificate expires, not after. Two sources: the health daemon's `CertRenewalFailing` node condition for node-level certificates, and, for your ingress certificates, cert-manager's expiry metric. To catch a node certificate before renewal ever fails, read the `autopilot.syself.com/certs` annotation on the Node, which lists each certificate with its `notAfter` date. See [Rotate and renew certificates](/docs/hetzner/apalla/network/dns-certs/certificate-rotation-and-renewal). Keep all of these in Git as `PrometheusRule` objects so a rebuilt cluster comes back with the same alerts, and route the `critical` ones to a pager. See [Alert routing and receivers](/docs/hetzner/apalla/observability/alerting/alert-routing-and-receivers).