Skip to main content

Node health conditions

Inspect 1.36

Every Syself Autopilot node runs a health daemon. It watches the node's services, disks, certificates, clock and network, and it reads the kernel's log. What it finds is written onto the Node object as NodeConditions, so you can read the state of a node with kubectl and alert on it.

Each condition has a name and a status. These conditions report problems, so True means the health daemon found one and False means it did not. Kubernetes' own Ready condition runs the other way round, where True is the state you want. Every condition the daemon sets exists on every node from the moment it joins, so a node with nothing wrong lists them all as False.

How to read them on a node

		$ kubectl get node <name> \
  -o jsonpath='{range .status.conditions[*]}{.type}={.status} {.reason}{"\n"}{end}'
DisksFailure=False DisksHealthy
ReadonlyFilesystem=False FilesystemIsNotReadOnly
CertRenewalFailing=False Renewing
ClockNotSynchronized=False ClockSynced
...
	

The command prints three fields. The full entry carries these:

Field Meaning
type the condition name (for example DisksFailure, ReadonlyFilesystem)
status True (problem present), False (healthy), or Unknown
reason a short machine-readable reason (for example FilesystemIsReadOnly)
message a human-readable detail line
lastTransitionTime when the status last changed
lastHeartbeatTime when the health daemon last refreshed the entry, changed or not

To see one condition in detail, select it by name:

		$ kubectl get node <name> \
  -o jsonpath='{range .status.conditions[?(@.type=="ReadonlyFilesystem")]}{@}{end}' | jq .
{
  "lastHeartbeatTime": "2026-09-28T07:50:19Z",
  "lastTransitionTime": "2026-09-15T11:56:54Z",
  "message": "no filesystem was remounted read-only",
  "reason": "FilesystemIsNotReadOnly",
  "status": "False",
  "type": "ReadonlyFilesystem"
}
	

When the health daemon finds a problem it sets the condition to True. What clears it again depends on which kind of check set it.

A scheduled check re-runs on its interval and writes whatever it finds, so once the problem is gone the next run sets its condition back to False.

Conditions from the kernel log are the special case. A kernel message is just a log line: it says a problem happened, and afterwards there is no second line when it gets fixed. Those conditions stay True until the health daemon restarts.

The health daemon runs on the node, so a node failure takes the daemon with it. The daemon can also fail on its own, and the two cases are handled differently.

The node is gone. The kubelet stops reporting and Kubernetes moves Ready to Unknown. On cloud pools Syself Autopilot replaces the machine on its own, with nothing required from you. On bare-metal pools nothing happens automatically, because Ready=Unknown there can mean a network partition rather than a dead server, so that case is left to you.

The health daemon is gone but the node is fine. Nothing detects this for you. The conditions freeze at their last value, so the node keeps showing them as False while nothing is checking it. Alert on the health report's updatedAt timestamp, which stops updating as soon as the health daemon goes down.

The health report behind a condition

A condition tells you that something is wrong. The health report tells you what.

The health daemon writes it to the node's annotations under autopilot.syself.com/, split across five annotations: storage, disks, certs, tamper and services. Each one is refreshed every two minutes.

Each is a self-contained JSON document, carrying its own updatedAt and role fields, so you can read the certificates without pulling the rest.

json
		// autopilot.syself.com/storage
{ "updatedAt": "<timestamp>", "role": "controlplane | worker",
  "filesystems": [ ... ], "volumeGroups": [ ... ] }
 
// autopilot.syself.com/disks
{ "updatedAt": "<timestamp>", "role": "...", "disks": [ ... ] }
 
// autopilot.syself.com/certs
{ "updatedAt": "<timestamp>", "role": "...", "certs": [ ... ] }
 
// autopilot.syself.com/tamper
{ "updatedAt": "<timestamp>", "role": "...", "tamper": { ... } }
 
// autopilot.syself.com/services
{ "updatedAt": "<timestamp>", "role": "...", "services": [ ... ] }
	

Read one annotation like this:

		$ kubectl get node <name> \
  -o jsonpath='{.metadata.annotations.autopilot\.syself\.com/disks}' | jq .
	

The other four annotations also work the same way. A key with nothing to report on that node reads null.

How to alert on a condition

Conditions serve two purposes. A small set drives remediation while others are informational. The health daemon records them and takes no further action, so alerting on informational ones is yours to configure. Conditions to alert on says which is which, per pool type.

The path is . It turns every NodeCondition into a Prometheus metric, kube_node_status_condition{condition="...",status="true"}, and you write a Prometheus or Alertmanager rule on it. For example, a rule that fires when kube_node_status_condition{condition="DisksFailure",status="true"} == 1 pages you when a disk starts failing.

A condition you read with kubectl only exists for as long as the node does. When a node is replaced its Machine and its Node are deleted together, so the conditions and the health report go with it. The only history you keep is what your collector scraped while the node was still there. covers the same concept for logs.

How the platform reboots, drains and replaces a node it considers unhealthy is in .

See for installing the metrics pipeline, and for the daemon's own metrics on port 20257.

Conditions to alert on

These are the conditions that need you. The daemon sets more than these, but the rest cover the units it repairs itself and Syself's own services on the node, so they are not signals you act on.

Alert on all of them. On bare-metal pools the platform acts on Ready=False alone, so nothing below is handled for you. On cloud pools four of these also trigger a reboot and then a replacement, marked in the tables, and everything else is still yours to catch.

Checks that run on a schedule

Condition What it means
ServiceNotRecovering repeated repair did not fix kubelet, containerd, or the worker proxy. On cloud pools this triggers a reboot, then a replacement
KubeletHeartbeatFailed kubelet lost its API connection and could not re-establish it
DisksFailure a physical disk is wearing out or failing
DiskWearHigh a drive is approaching end of life
DiskTemperatureHigh a cooling or airflow problem
DiskUsageHigh logs, images, or data filled a filesystem
CertRenewalFailing a node certificate expires within 7 days and has not renewed
ClockNotSynchronized NTP is broken and the clock is drifting
CNIUnhealthy the cilium-agent crashed or is degraded
GpuActivationFailed gpu-activate did not finish, so containerd cannot use the GPU (GPU nodes only)

Conditions from the kernel log

The health daemon also watches /dev/kmsg, the kernel's message log, and sets a condition when a line matches a problem it knows.

Condition What it means
ReadonlyFilesystem the kernel hit an I/O or filesystem error and remounted read-only to protect data; writes now fail. On cloud pools this triggers a reboot, then a replacement
KernelDeadlock a critical process is wedged on stuck I/O or a kernel deadlock. On cloud pools this triggers a reboot, then a replacement
VerityCorruption a block on the sealed read-only OS failed its dm-verity check, from disk corruption or tampering. On cloud pools this triggers a reboot, then a replacement
CperHardwareErrorFatal a fatal hardware fault in CPU, memory, or PCIe, reported by firmware
GpuFallenOffBus a GPU dropped off the PCI bus due to power, heat, or a hardware fault (GPU nodes only)

Integrity and tamper

Condition What it means
SealedOSTampered a sealed layer was swapped for a different image, or its recorded hash was altered on the writable partition
NodeTampered a protected file under /var was added, removed, or changed since the boot baseline

NodeTampered does not clear when you revert the change. Once the daemon has confirmed one, it keeps reporting the node as tampered.

SealedOSTampered and VerityCorruption both cover the sealed OS layers but catch different attacks. VerityCorruption fires when the kernel hits a bad block on a read, meaning a sealed layer was changed in place. SealedOSTampered fires when a layer's hash no longer matches the recorded one. That catches a layer swapped for a different image, one that is internally consistent under its own hash.

See for what is baselined, and for the runbook.

Events only (no condition set)

Some kernel problems raise a Node event and increment problem_counter{reason="..."} without setting a condition: disk and filesystem I/O errors, hung tasks, corrected memory errors, NVIDIA Xid errors, and hardware errors the firmware reports as corrected or recoverable. The event and the counter are the only record.

The API server deletes events after three hours. The counter is a metric on the health daemon's endpoint, so whatever your collector scrapes stays in your metrics store for as long as you keep it. Read the event while it is still there. To catch a problem that repeats, alert on the rate of problem_counter.

OOMKilling and KernelOops are not emitted. A pod OOM is a workload signal the kubelet already reports. A kernel oops is rare and informational, so both were only noise as Node events.