Node health conditions
Every Syself Autopilot node runs a health daemon. It watches the node's services, disks, certificates, clock and network, and it reads the kernel's log. What it finds is written onto the Node object as NodeConditions, so you can read the state of a node with kubectl and alert on it.
Each condition has a name and a status. These conditions report problems, so True means the health daemon found one and False means it did not. Kubernetes' own Ready condition runs the other way round, where True is the state you want. Every condition the daemon sets exists on every node from the moment it joins, so a node with nothing wrong lists them all as False.
How to read them on a node
$ kubectl get node <name> \
-o jsonpath='{range .status.conditions[*]}{.type}={.status} {.reason}{"\n"}{end}'
DisksFailure=False DisksHealthy
ReadonlyFilesystem=False FilesystemIsNotReadOnly
CertRenewalFailing=False Renewing
ClockNotSynchronized=False ClockSynced
...
The command prints three fields. The full entry carries these:
| Field | Meaning |
|---|---|
type | the condition name (for example DisksFailure, ReadonlyFilesystem) |
status | True (problem present), False (healthy), or Unknown |
reason | a short machine-readable reason (for example FilesystemIsReadOnly) |
message | a human-readable detail line |
lastTransitionTime | when the status last changed |
lastHeartbeatTime | when the health daemon last refreshed the entry, changed or not |
To see one condition in detail, select it by name:
$ kubectl get node <name> \
-o jsonpath='{range .status.conditions[?(@.type=="ReadonlyFilesystem")]}{@}{end}' | jq .
{
"lastHeartbeatTime": "2026-09-28T07:50:19Z",
"lastTransitionTime": "2026-09-15T11:56:54Z",
"message": "no filesystem was remounted read-only",
"reason": "FilesystemIsNotReadOnly",
"status": "False",
"type": "ReadonlyFilesystem"
}
When the health daemon finds a problem it sets the condition to True. What clears it again depends on which kind of check set it.
A scheduled check re-runs on its interval and writes whatever it finds, so once the problem is gone the next run sets its condition back to False.
Conditions from the kernel log are the special case. A kernel message is just a log line: it says a problem happened, and afterwards there is no second line when it gets fixed. Those conditions stay True until the health daemon restarts.
The health daemon runs on the node, so a node failure takes the daemon with it. The daemon can also fail on its own, and the two cases are handled differently.
The node is gone. The kubelet stops reporting and Kubernetes moves Ready to Unknown. On cloud pools Syself Autopilot replaces the machine on its own, with nothing required from you. On bare-metal pools nothing happens automatically, because Ready=Unknown there can mean a network partition rather than a dead server, so that case is left to you.
The health daemon is gone but the node is fine. Nothing detects this for you. The conditions freeze at their last value, so the node keeps showing them as False while nothing is checking it. Alert on the health report's updatedAt timestamp, which stops updating as soon as the health daemon goes down.
The health report behind a condition
A condition tells you that something is wrong. The health report tells you what.
The health daemon writes it to the node's annotations under autopilot.syself.com/, split across five annotations: storage, disks, certs, tamper and services. Each one is refreshed every two minutes.
Each is a self-contained JSON document, carrying its own updatedAt and role fields, so you can read the certificates without pulling the rest.
// autopilot.syself.com/storage
{ "updatedAt": "<timestamp>", "role": "controlplane | worker",
"filesystems": [ ... ], "volumeGroups": [ ... ] }
// autopilot.syself.com/disks
{ "updatedAt": "<timestamp>", "role": "...", "disks": [ ... ] }
// autopilot.syself.com/certs
{ "updatedAt": "<timestamp>", "role": "...", "certs": [ ... ] }
// autopilot.syself.com/tamper
{ "updatedAt": "<timestamp>", "role": "...", "tamper": { ... } }
// autopilot.syself.com/services
{ "updatedAt": "<timestamp>", "role": "...", "services": [ ... ] }
Read one annotation like this:
$ kubectl get node <name> \
-o jsonpath='{.metadata.annotations.autopilot\.syself\.com/disks}' | jq .
The other four annotations also work the same way. A key with nothing to report on that node reads null.
How to alert on a condition
Conditions serve two purposes. A small set drives remediation while others are informational. The health daemon records them and takes no further action, so alerting on informational ones is yours to configure. Conditions to alert on says which is which, per pool type.
The path is kube-state-metrics . It turns every NodeCondition into a Prometheus metric, kube_node_status_condition{condition="...",status="true"}, and you write a Prometheus or Alertmanager rule on it. For example, a rule that fires when kube_node_status_condition{condition="DisksFailure",status="true"} == 1 pages you when a disk starts failing.
A condition you read with kubectl only exists for as long as the node does. When a node is replaced its Machine and its Node are deleted together, so the conditions and the health report go with it. The only history you keep is what your collector scraped while the node was still there. Detection vs retention covers the same concept for logs.
How the platform reboots, drains and replaces a node it considers unhealthy is in Machine health checks and remediation .
See set up Prometheus for installing the metrics pipeline, and Metrics per component for the daemon's own metrics on port 20257.
Conditions to alert on
These are the conditions that need you. The daemon sets more than these, but the rest cover the units it repairs itself and Syself's own services on the node, so they are not signals you act on.
Alert on all of them. On bare-metal pools the platform acts on Ready=False alone, so nothing below is handled for you. On cloud pools four of these also trigger a reboot and then a replacement, marked in the tables, and everything else is still yours to catch.
Checks that run on a schedule
| Condition | What it means |
|---|---|
ServiceNotRecovering | repeated repair did not fix kubelet, containerd, or the worker proxy. On cloud pools this triggers a reboot, then a replacement |
KubeletHeartbeatFailed | kubelet lost its API connection and could not re-establish it |
DisksFailure | a physical disk is wearing out or failing |
DiskWearHigh | a drive is approaching end of life |
DiskTemperatureHigh | a cooling or airflow problem |
DiskUsageHigh | logs, images, or data filled a filesystem |
CertRenewalFailing | a node certificate expires within 7 days and has not renewed |
ClockNotSynchronized | NTP is broken and the clock is drifting |
CNIUnhealthy | the cilium-agent crashed or is degraded |
GpuActivationFailed | gpu-activate did not finish, so containerd cannot use the GPU (GPU nodes only) |
Conditions from the kernel log
The health daemon also watches /dev/kmsg, the kernel's message log, and sets a condition when a line matches a problem it knows.
| Condition | What it means |
|---|---|
ReadonlyFilesystem | the kernel hit an I/O or filesystem error and remounted read-only to protect data; writes now fail. On cloud pools this triggers a reboot, then a replacement |
KernelDeadlock | a critical process is wedged on stuck I/O or a kernel deadlock. On cloud pools this triggers a reboot, then a replacement |
VerityCorruption | a block on the sealed read-only OS failed its dm-verity check, from disk corruption or tampering. On cloud pools this triggers a reboot, then a replacement |
CperHardwareErrorFatal | a fatal hardware fault in CPU, memory, or PCIe, reported by firmware |
GpuFallenOffBus | a GPU dropped off the PCI bus due to power, heat, or a hardware fault (GPU nodes only) |
Integrity and tamper
| Condition | What it means |
|---|---|
SealedOSTampered | a sealed layer was swapped for a different image, or its recorded hash was altered on the writable partition |
NodeTampered | a protected file under /var was added, removed, or changed since the boot baseline |
NodeTampered does not clear when you revert the change. Once the daemon has confirmed one, it keeps reporting the node as tampered.
SealedOSTampered and VerityCorruption both cover the sealed OS layers but catch different attacks. VerityCorruption fires when the kernel hits a bad block on a read, meaning a sealed layer was changed in place. SealedOSTampered fires when a layer's hash no longer matches the recorded one. That catches a layer swapped for a different image, one that is internally consistent under its own hash.
See Node and OS security for what is baselined, and Respond to a tampered node for the runbook.
Events only (no condition set)
Some kernel problems raise a Node event and increment problem_counter{reason="..."} without setting a condition: disk and filesystem I/O errors, hung tasks, corrected memory errors, NVIDIA Xid errors, and hardware errors the firmware reports as corrected or recoverable. The event and the counter are the only record.
The API server deletes events after three hours. The counter is a metric on the health daemon's endpoint, so whatever your collector scrapes stays in your metrics store for as long as you keep it. Read the event while it is still there. To catch a problem that repeats, alert on the rate of problem_counter.
OOMKilling and KernelOops are not emitted. A pod OOM is a workload signal the kubelet already reports. A kernel oops is rare and informational, so both were only noise as Node events.
Related
- Metrics per component : the health daemon's own metrics on port
20257, and every other metrics endpoint on a node. - kube-state-metrics : the component that turns these conditions into Prometheus metrics.
- Machine health checks and remediation : what the platform does with a node it considers unhealthy, and how to pause it for maintenance.
- Platform labels and annotations : the node's labels and the
autopilot.syself.com/*health annotations.
Metrics per component
Every Syself Autopilot component's metrics endpoint, its port and scheme, which nodes it exists on, and whether it sits on the pod network or the node's loopback. Health-daemon metrics and how to scrape a host-network node.
Log sources on a sealed node
A sealed Syself node has a read-only root and a writable /var, and this page lists every log stream that lives there and disappears on reprovision.