Every Syself Autopilot node runs a health daemon. It watches the node's services, disks, certificates, clock and network, and it reads the kernel's log. What it finds is written onto the Node object as **NodeConditions**, so you can read the state of a node with `kubectl` and alert on it. Each condition has a name and a status. These conditions report problems, so `True` means the health daemon found one and `False` means it did not. Kubernetes' own `Ready` condition runs the other way round, where `True` is the state you want. Every condition the daemon sets exists on every node from the moment it joins, so a node with nothing wrong lists them all as `False`. ## How to read them on a node ```console $ kubectl get node \ -o jsonpath='{range .status.conditions[*]}{.type}={.status} {.reason}{"\n"}{end}' DisksFailure=False DisksHealthy ReadonlyFilesystem=False FilesystemIsNotReadOnly CertRenewalFailing=False Renewing ClockNotSynchronized=False ClockSynced ... ``` The command prints three fields. The full entry carries these: | Field | Meaning | | -------------------- | --------------------------------------------------------------------- | | `type` | the condition name (for example `DisksFailure`, `ReadonlyFilesystem`) | | `status` | `True` (problem present), `False` (healthy), or `Unknown` | | `reason` | a short machine-readable reason (for example `FilesystemIsReadOnly`) | | `message` | a human-readable detail line | | `lastTransitionTime` | when the status last changed | | `lastHeartbeatTime` | when the health daemon last refreshed the entry, changed or not | To see one condition in detail, select it by name: ```console $ kubectl get node \ -o jsonpath='{range .status.conditions[?(@.type=="ReadonlyFilesystem")]}{@}{end}' | jq . { "lastHeartbeatTime": "2026-09-28T07:50:19Z", "lastTransitionTime": "2026-09-15T11:56:54Z", "message": "no filesystem was remounted read-only", "reason": "FilesystemIsNotReadOnly", "status": "False", "type": "ReadonlyFilesystem" } ``` When the health daemon finds a problem it sets the condition to `True`. What clears it again depends on which kind of check set it. A scheduled check re-runs on its interval and writes whatever it finds, so once the problem is gone the next run sets its condition back to `False`. Conditions from the kernel log are the special case. A kernel message is just a log line: it says a problem happened, and afterwards there is no second line when it gets fixed. Those conditions stay `True` until the health daemon restarts. The health daemon runs on the node, so a node failure takes the daemon with it. The daemon can also fail on its own, and the two cases are handled differently. **The node is gone.** The kubelet stops reporting and Kubernetes moves `Ready` to `Unknown`. On cloud pools Syself Autopilot replaces the machine on its own, with nothing required from you. On bare-metal pools nothing happens automatically, because `Ready=Unknown` there can mean a network partition rather than a dead server, so that case is left to you. **The health daemon is gone but the node is fine.** Nothing detects this for you. The conditions freeze at their last value, so the node keeps showing them as `False` while nothing is checking it. Alert on the health report's `updatedAt` timestamp, which stops updating as soon as the health daemon goes down. ## The health report behind a condition A condition tells you that something is wrong. The health report tells you what. The health daemon writes it to the node's annotations under `autopilot.syself.com/`, split across five annotations: `storage`, `disks`, `certs`, `tamper` and `services`. Each one is refreshed every two minutes. Each is a self-contained JSON document, carrying its own `updatedAt` and `role` fields, so you can read the certificates without pulling the rest. ```json // autopilot.syself.com/storage { "updatedAt": "", "role": "controlplane | worker", "filesystems": [ ... ], "volumeGroups": [ ... ] } // autopilot.syself.com/disks { "updatedAt": "", "role": "...", "disks": [ ... ] } // autopilot.syself.com/certs { "updatedAt": "", "role": "...", "certs": [ ... ] } // autopilot.syself.com/tamper { "updatedAt": "", "role": "...", "tamper": { ... } } // autopilot.syself.com/services { "updatedAt": "", "role": "...", "services": [ ... ] } ``` Read one annotation like this: ```console $ kubectl get node \ -o jsonpath='{.metadata.annotations.autopilot\.syself\.com/disks}' | jq . ``` The other four annotations also work the same way. A key with nothing to report on that node reads `null`. ## How to alert on a condition Conditions serve two purposes. A small set drives remediation while others are informational. The health daemon records them and takes no further action, so alerting on informational ones is yours to configure. [Conditions to alert on](#conditions-to-alert-on) says which is which, per pool type. The path is [`kube-state-metrics`](/docs/hetzner/apalla/observability/metrics/kube-state-metrics). It turns every NodeCondition into a Prometheus metric, `kube_node_status_condition{condition="...",status="true"}`, and you write a Prometheus or Alertmanager rule on it. For example, a rule that fires when `kube_node_status_condition{condition="DisksFailure",status="true"} == 1` pages you when a disk starts failing. A condition you read with `kubectl` only exists for as long as the node does. When a node is replaced its `Machine` and its `Node` are deleted together, so the conditions and the health report go with it. The only history you keep is what your collector scraped while the node was still there. [Detection vs retention](/docs/hetzner/apalla/observability/detection-vs-retention) covers the same concept for logs. How the platform reboots, drains and replaces a node it considers unhealthy is in [Machine health checks and remediation](/docs/hetzner/apalla/servers-and-nodes/maintenance/machine-health-checks-and-remediation). See [set up Prometheus](/docs/hetzner/apalla/observability/metrics/set-up-prometheus) for installing the metrics pipeline, and [Metrics per component](/docs/hetzner/apalla/observability/reference/metrics-per-component) for the daemon's own metrics on port `20257`. ## Conditions to alert on These are the conditions that need you. The daemon sets more than these, but the rest cover the units it repairs itself and Syself's own services on the node, so they are not signals you act on. Alert on all of them. On bare-metal pools the platform acts on `Ready=False` alone, so nothing below is handled for you. On cloud pools four of these also trigger a reboot and then a replacement, marked in the tables, and everything else is still yours to catch. ### Checks that run on a schedule | Condition | What it means | | ------------------------ | ------------------------------------------------------------------------------------------------------------------------------- | | `ServiceNotRecovering` | repeated repair did not fix kubelet, containerd, or the worker proxy. On cloud pools this triggers a reboot, then a replacement | | `KubeletHeartbeatFailed` | kubelet lost its API connection and could not re-establish it | | `DisksFailure` | a physical disk is wearing out or failing | | `DiskWearHigh` | a drive is approaching end of life | | `DiskTemperatureHigh` | a cooling or airflow problem | | `DiskUsageHigh` | logs, images, or data filled a filesystem | | `CertRenewalFailing` | a node certificate expires within 7 days and has not renewed | | `ClockNotSynchronized` | NTP is broken and the clock is drifting | | `CNIUnhealthy` | the cilium-agent crashed or is degraded | | `GpuActivationFailed` | `gpu-activate` did not finish, so containerd cannot use the GPU (GPU nodes only) | ### Conditions from the kernel log The health daemon also watches `/dev/kmsg`, the kernel's message log, and sets a condition when a line matches a problem it knows. | Condition | What it means | | ------------------------ | ------------------------------------------------------------------------------------------------------------------------------------------------------------- | | `ReadonlyFilesystem` | the kernel hit an I/O or filesystem error and remounted read-only to protect data; writes now fail. On cloud pools this triggers a reboot, then a replacement | | `KernelDeadlock` | a critical process is wedged on stuck I/O or a kernel deadlock. On cloud pools this triggers a reboot, then a replacement | | `VerityCorruption` | a block on the sealed read-only OS failed its dm-verity check, from disk corruption or tampering. On cloud pools this triggers a reboot, then a replacement | | `CperHardwareErrorFatal` | a fatal hardware fault in CPU, memory, or PCIe, reported by firmware | | `GpuFallenOffBus` | a GPU dropped off the PCI bus due to power, heat, or a hardware fault (GPU nodes only) | ### Integrity and tamper | Condition | What it means | | ------------------ | ------------------------------------------------------------------------------------------------------------ | | `SealedOSTampered` | a sealed layer was swapped for a different image, or its recorded hash was altered on the writable partition | | `NodeTampered` | a protected file under `/var` was added, removed, or changed since the boot baseline | `NodeTampered` does not clear when you revert the change. Once the daemon has confirmed one, it keeps reporting the node as tampered. `SealedOSTampered` and `VerityCorruption` both cover the sealed OS layers but catch different attacks. `VerityCorruption` fires when the kernel hits a bad block on a read, meaning a sealed layer was changed in place. `SealedOSTampered` fires when a layer's hash no longer matches the recorded one. That catches a layer swapped for a different image, one that is internally consistent under its own hash. See [Node and OS security](/docs/hetzner/apalla/security/node-and-os-security) for what is baselined, and [Respond to a tampered node](/docs/hetzner/apalla/security/respond-to-a-tampered-node) for the runbook. ### Events only (no condition set) Some kernel problems raise a Node event and increment `problem_counter{reason="..."}` without setting a condition: disk and filesystem I/O errors, hung tasks, corrected memory errors, NVIDIA Xid errors, and hardware errors the firmware reports as corrected or recoverable. The event and the counter are the only record. The API server deletes events after three hours. The counter is a metric on the health daemon's endpoint, so whatever your collector scrapes stays in your metrics store for as long as you keep it. Read the event while it is still there. To catch a problem that repeats, alert on the rate of `problem_counter`. `OOMKilling` and `KernelOops` are not emitted. A pod OOM is a workload signal the kubelet already reports. A kernel oops is rare and informational, so both were only noise as Node events. ## Related - [Metrics per component](/docs/hetzner/apalla/observability/reference/metrics-per-component): the health daemon's own metrics on port `20257`, and every other metrics endpoint on a node. - [kube-state-metrics](/docs/hetzner/apalla/observability/metrics/kube-state-metrics): the component that turns these conditions into Prometheus metrics. - [Machine health checks and remediation](/docs/hetzner/apalla/servers-and-nodes/maintenance/machine-health-checks-and-remediation): what the platform does with a node it considers unhealthy, and how to pause it for maintenance. - [Platform labels and annotations](/docs/hetzner/apalla/servers-and-nodes/labels/platform-labels-and-annotations): the node's labels and the `autopilot.syself.com/*` health annotations.