Skip to main content

Machine health checks and remediation

Inspect 1.36

Syself Autopilot watches every node and replaces the ones that stop being healthy. Most of the time you do nothing.

Every node in a Syself Autopilot cluster runs a health daemon. The daemon continuously checks the node and reports problems as Kubernetes NodeConditions (flags on a Node object that describe its health state), events, and metrics.

What a condition does when it fires depends on the pool. A few conditions trigger automatic machine replacement; the rest are there for visibility and alerting.

What the health check watches

Syself Autopilot watches each node's Ready state, the kubelet's summary of whether the node can still run pods. When a node reports itself broken, or stops reporting at all, for longer than a short grace period, it counts as unhealthy and remediation begins.

On cloud pools, Syself Autopilot also acts on a few of its own health signals that mean the node is unlikely to recover on its own. Bare-metal pools are more conservative: a Ready failure triggers remediation, but the deeper health signals alert a person instead of reprovisioning the server, because a reinstall does not fix a hardware or storage fault, and the same fault can strike several servers at once.

See for what a node reports and how to alert on it.

The daemon and the replacement logic are separate. Keep them distinct.

The daemon reports. It detects a problem and writes it as a NodeCondition, an event, and a metric. That is all it does for most problems. It does not cordon, drain, or replace nodes. It does not set the Ready condition; the kubelet owns that.

Three things act on what it reports.

  • In-place repair restarts a broken unit (kubelet, containerd) on the node.
  • Syself's replacement logic reboots the node and, if that fails, replaces it. It is the only automatic actor that deletes a node.
  • You, by alerting on a condition and deciding what to do. You can also force a replacement at any time: delete the Machine object, and the controllers drain the node and bring up a fresh one in its place.

Which conditions the replacement logic acts on, and how long each has to hold, is in What remediation does.

In-place repair

Before anything heavier runs, a broken critical unit is restarted on the node.

kubelet: the health daemon probes /healthz every 10 seconds. On a failed probe it restarts kubelet.service and waits a 1-minute cooldown before trying again.

containerd: the daemon probes the CRI socket (CRI, the Container Runtime Interface, connects kubelet to the container runtime) every 10 seconds. On failure it restarts containerd.service with a 2-minute cooldown.

systemd also restarts these units on its own (Restart=always).

Every restart increments systemd's NRestarts counter for that unit.

When repair is not working. The service-not-recovering check runs every minute. For the critical units (kubelet, containerd, and on a worker the failover proxy), it sets ServiceNotRecovering = True when a unit is either in the systemd failed state, or flapping (its NRestarts went up by 3 or more since the last check). ServiceNotRecovering is one of the conditions you alert on yourself; see .

The daemon also exports syself_node_unit_restarts{unit} as a Prometheus metric, so you can see repair activity as a time series.

What remediation does

Remediation tries a reboot first. If the node comes back, it stays. If not, the machine is drained (its pods move to other nodes) and then replaced.

flowchart LR
  A["Node unhealthy past its window"]:::platform --> B["Reboot"]:::platform
  B --> C{"Node recovers?"}:::decision
  C -->|Yes| D["Stays in the cluster"]:::platform
  C -->|No| E["Drain, then replace"]:::platform

The last step differs by infrastructure. On cloud, Syself Autopilot reboots the node and, if it does not come back, replaces it with a fresh machine. On bare metal it is more patient, retrying the reboot before falling back to reprovisioning the host, since a dedicated server is worth more effort to recover than a disposable VM.

On bare-metal pools the replacement logic watches only the Ready condition. On cloud pools it also watches four of the daemon's own conditions: ReadonlyFilesystem, KernelDeadlock, ServiceNotRecovering, and VerityCorruption. A reboot clears these conditions and if it does not then a newly provisioned VM does.

Tip

The Ready=False window is long enough that a normal reboot or a short network blip resolves before Syself Autopilot notices. If your planned work (a firmware update, a manual reboot, hardware swap) takes longer than that, pause the Machine first, or Syself Autopilot starts remediating a node you are already fixing. The exact windows are in Cloud pools and Bare-metal pools.

Cloud pools

The control plane and the worker pools use the same windows.

Watched condition Unhealthy when Must hold for Then
Ready Unknown 600s reboot, then replace
Ready False 300s reboot, then replace
ReadonlyFilesystem True 120s reboot, then replace
KernelDeadlock True 120s reboot, then replace
ServiceNotRecovering True 60s reboot, then replace
VerityCorruption True 60s reboot, then replace

Syself reboots the server once, allowing up to 180 seconds for the reboot to complete. If the condition is still true after it has held for the time shown above, Syself replaces the machine with a new VM.

Bare-metal pools

The control plane and the worker pools use the same window.

Watched condition Unhealthy when Must hold for Then
Ready False 300s reboot (up to 2×), then re-provision

On bare metal, Ready=Unknown alerts a human instead of triggering remediation. It is not wired into automatic replacement.

Note

Automatic replacement on bare metal is deliberately narrow right now: only Ready=False triggers it. Re-provisioning the same physical host will not fix a bad disk or failing hardware, so the daemon's own conditions alert a human instead. Widening the set of conditions that trigger automatic action on bare metal is planned for a future Kubernetes minor track.

Syself reboots the physical server up to two times. If the node is still unhealthy after 300 seconds, Syself re-provisions the same physical host (wipe and reinstall). The longer timeout gives hardware more time to stabilise after a reboot before Syself escalates to re-provisioning.

The platform stops before it makes things worse

Remediation is bounded on purpose. Syself Autopilot only replaces nodes while the number of unhealthy machines in a pool stays small, and it never removes a control-plane node in a way that would cost etcd its quorum, the majority of control-plane members the cluster needs to stay writable. A node that is still booting is given time to settle before it counts against that limit at all.

The exact limits, per pool type:

Setting Control plane Cloud worker Bare-metal worker
Remediate only while this many are unhealthy ≤ 1 machine (an absolute count, not a percentage; it caps concurrent remediation. KCP separately refuses any deletion that would cost etcd quorum) [0-2] in the pool ≤ 1
nodeStartupTimeout (grace for a booting node) 1800s 600s 1800s
Drain before delete (nodeDrainTimeout) 180s 180s 180s
Detach volumes before delete (nodeVolumeDetachTimeout) 120s 120s 120s

These values (the startup grace, the drain timeout, the volume-detach timeout, the concurrency gate, and the rollout maxSurge/maxUnavailable) are upstream Cluster API settings. The Cluster Stack ships them as tested defaults on the generated MachineDeployment and KubeadmControlPlane. They are not exposed as topology variables today, and the topology controller reverts a hand edit to a derived object on its next reconcile, so treat them as fixed for now. Making the drain and reschedule behavior tunable per pool is planned for a future release.

Important

If more nodes go unhealthy than the limit allows, automatic replacement stops and waits for you. A fleet-wide problem should not trigger fleet-wide reprovisioning. That is your cue to take over: check the to see what is happening, pause remediation on any node you are working on, and fix the root cause before you let replacement resume.

Before a node is removed its pods are drained, and the drain respects PodDisruptionBudgets, so set one on any workload that must not lose too many replicas at once. See .

Self-healing does not repair hardware

Remediation swaps a node out. It does not fix the physical machine underneath. On bare metal this matters most when a pool's host selector is narrow. If every host the selector matches is already in use, a machine that fails is released and then re-claims the same host, because no other host qualifies. The server is wiped and reinstalled, and a fault in the hardware itself comes straight back.

Re-provisioning clears a software or filesystem problem. It does nothing for a failing disk, bad memory, or a dying NIC. So do not lean on self-healing to route around broken hardware. Detect a hardware fault directly instead:

  • Alert on the hardware-related . On bare metal these page a human rather than trigger a reboot, precisely because a reboot would not help.
  • Or run your own observability against the servers (SMART data, memory, temperature, link state) and alert from there.

When a server is genuinely broken, take it out of the pool: set maintenanceMode: true on its HetznerBareMetalHost (or remove it), and open a Hetzner support case for the hardware. Alerting on the node conditions is the reliable way to catch a hardware fault.

Pause remediation during maintenance

When you plan to disrupt a node yourself (a firmware update, an in-place reboot, hardware work), pause remediation so Syself Autopilot does not mistake your maintenance for a failure. Pausing a Machine stops the management cluster from reconciling it at all, not only from remediating it.

Annotate the Machine

In the management cluster:

		$ kubectl annotate machine <machine-name> -n <namespace> cluster.x-k8s.io/paused=true
	

Verify the annotation is set

Before you start work, check the value:

		$ kubectl get machine <machine-name> -n <namespace> \
  -o jsonpath='{.metadata.annotations.cluster\.x-k8s\.io/paused}'
	

The output should be true.

Do your maintenance

The Machine is not touched while it is paused, so do the firmware update, reboot, or hardware work now.

Remove the pause when you are done

		$ kubectl annotate machine <machine-name> -n <namespace> cluster.x-k8s.io/paused-
	
Warning

Double-check the annotation for typos. If it is not set correctly, the node can be flagged unhealthy and reprovisioned while you work on it.

When remediation should not fire

Pause first for anything that makes a healthy node look unhealthy on purpose: rebooting it, taking it off the network, or running long disruptive maintenance. Leave remediation on for everything else. A node that genuinely fails should be replaced, and that is the whole point of the system.