Syself Autopilot watches every node and replaces the ones that stop being healthy. Most of the time you do nothing. Every node in a Syself Autopilot cluster runs a health daemon. The daemon continuously checks the node and reports problems as Kubernetes NodeConditions (flags on a `Node` object that describe its health state), events, and metrics. What a condition does when it fires depends on the pool. A few conditions trigger automatic machine replacement; the rest are there for visibility and alerting. ## What the health check watches Syself Autopilot watches each node's `Ready` state, the kubelet's summary of whether the node can still run pods. When a node reports itself broken, or stops reporting at all, for longer than a short grace period, it counts as unhealthy and remediation begins. On cloud pools, Syself Autopilot also acts on a few of its own health signals that mean the node is unlikely to recover on its own. Bare-metal pools are more conservative: a `Ready` failure triggers remediation, but the deeper health signals alert a person instead of reprovisioning the server, because a reinstall does not fix a hardware or storage fault, and the same fault can strike several servers at once. See [Node problem detection and conditions](/docs/hetzner/apalla/servers-and-nodes/access/node-problem-detection-and-conditions) for what a node reports and how to alert on it. The daemon and the replacement logic are separate. Keep them distinct. **The daemon reports.** It detects a problem and writes it as a NodeCondition, an event, and a metric. That is all it does for most problems. It does not cordon, drain, or replace nodes. It does not set the `Ready` condition; the kubelet owns that. **Three things act on what it reports.** - **In-place repair** restarts a broken unit (kubelet, containerd) on the node. - **Syself's replacement logic** reboots the node and, if that fails, replaces it. It is the only automatic actor that deletes a node. - **You**, by alerting on a condition and deciding what to do. You can also force a replacement at any time: delete the `Machine` object, and the controllers drain the node and bring up a fresh one in its place. Which conditions the replacement logic acts on, and how long each has to hold, is in [What remediation does](#what-remediation-does). ## In-place repair Before anything heavier runs, a broken critical unit is restarted on the node. **kubelet**: the health daemon probes `/healthz` every 10 seconds. On a failed probe it restarts `kubelet.service` and waits a 1-minute cooldown before trying again. **containerd**: the daemon probes the CRI socket (CRI, the Container Runtime Interface, connects kubelet to the container runtime) every 10 seconds. On failure it restarts `containerd.service` with a 2-minute cooldown. systemd also restarts these units on its own (`Restart=always`). Every restart increments systemd's `NRestarts` counter for that unit. **When repair is not working.** The `service-not-recovering` check runs every minute. For the critical units (kubelet, containerd, and on a worker the failover proxy), it sets `ServiceNotRecovering = True` when a unit is either in the systemd `failed` state, or flapping (its `NRestarts` went up by 3 or more since the last check). `ServiceNotRecovering` is one of the conditions you alert on yourself; see [Node health conditions](/docs/hetzner/apalla/observability/reference/node-health-conditions#conditions-to-alert-on). The daemon also exports `syself_node_unit_restarts{unit}` as a Prometheus metric, so you can see repair activity as a time series. ## What remediation does Remediation tries a reboot first. If the node comes back, it stays. If not, the machine is drained (its pods move to other nodes) and then replaced. ```mermaid flowchart LR A["Node unhealthy past its window"]:::platform --> B["Reboot"]:::platform B --> C{"Node recovers?"}:::decision C -->|Yes| D["Stays in the cluster"]:::platform C -->|No| E["Drain, then replace"]:::platform ``` The last step differs by infrastructure. On cloud, Syself Autopilot reboots the node and, if it does not come back, replaces it with a fresh machine. On bare metal it is more patient, retrying the reboot before falling back to reprovisioning the host, since a dedicated server is worth more effort to recover than a disposable VM. On bare-metal pools the replacement logic watches only the `Ready` condition. On cloud pools it also watches four of the daemon's own conditions: `ReadonlyFilesystem`, `KernelDeadlock`, `ServiceNotRecovering`, and `VerityCorruption`. A reboot clears these conditions and if it does not then a newly provisioned VM does. > [!TIP] > The `Ready=False` window is long enough that a normal reboot or a short network blip resolves before Syself Autopilot notices. If your planned work (a firmware update, a manual reboot, hardware swap) takes longer than that, [pause the Machine first](#pause-remediation-during-maintenance), or Syself Autopilot starts remediating a node you are already fixing. The exact windows are in [Cloud pools](#cloud-pools) and [Bare-metal pools](#bare-metal-pools). ### Cloud pools The control plane and the worker pools use the same windows. | Watched condition | Unhealthy when | Must hold for | Then | | ---------------------- | -------------- | ------------- | -------------------- | | `Ready` | `Unknown` | 600s | reboot, then replace | | `Ready` | `False` | 300s | reboot, then replace | | `ReadonlyFilesystem` | `True` | 120s | reboot, then replace | | `KernelDeadlock` | `True` | 120s | reboot, then replace | | `ServiceNotRecovering` | `True` | 60s | reboot, then replace | | `VerityCorruption` | `True` | 60s | reboot, then replace | Syself reboots the server once, allowing up to 180 seconds for the reboot to complete. If the condition is still true after it has held for the time shown above, Syself replaces the machine with a new VM. ### Bare-metal pools The control plane and the worker pools use the same window. | Watched condition | Unhealthy when | Must hold for | Then | | ----------------- | -------------- | ------------- | ------------------------------------ | | `Ready` | `False` | 300s | reboot (up to 2×), then re-provision | On bare metal, `Ready=Unknown` alerts a human instead of triggering remediation. It is not wired into automatic replacement. > [!NOTE] > Automatic replacement on bare metal is deliberately narrow right now: only `Ready=False` triggers it. Re-provisioning the same physical host will not fix a bad disk or failing hardware, so the daemon's own conditions alert a human instead. Widening the set of conditions that trigger automatic action on bare metal is planned for a future Kubernetes minor track. Syself reboots the physical server up to two times. If the node is still unhealthy after 300 seconds, Syself re-provisions the same physical host (wipe and reinstall). The longer timeout gives hardware more time to stabilise after a reboot before Syself escalates to re-provisioning. ## The platform stops before it makes things worse Remediation is bounded on purpose. Syself Autopilot only replaces nodes while the number of unhealthy machines in a pool stays small, and it never removes a control-plane node in a way that would cost etcd its quorum, the majority of control-plane members the cluster needs to stay writable. A node that is still booting is given time to settle before it counts against that limit at all. The exact limits, per pool type: | Setting | Control plane | Cloud worker | Bare-metal worker | | -------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------- | ----------------- | | Remediate only while this many are unhealthy | `≤ 1` machine (an absolute count, not a percentage; it caps concurrent remediation. KCP separately refuses any deletion that would cost etcd quorum) | `[0-2]` in the pool | `≤ 1` | | `nodeStartupTimeout` (grace for a booting node) | 1800s | 600s | 1800s | | Drain before delete (`nodeDrainTimeout`) | 180s | 180s | 180s | | Detach volumes before delete (`nodeVolumeDetachTimeout`) | 120s | 120s | 120s | These values (the startup grace, the drain timeout, the volume-detach timeout, the concurrency gate, and the rollout `maxSurge`/`maxUnavailable`) are upstream Cluster API settings. The Cluster Stack ships them as tested defaults on the generated `MachineDeployment` and `KubeadmControlPlane`. They are not exposed as topology variables today, and the topology controller reverts a hand edit to a derived object on its next reconcile, so treat them as fixed for now. Making the drain and reschedule behavior tunable per pool is planned for a future release. > [!IMPORTANT] > If more nodes go unhealthy than the limit allows, automatic replacement stops and waits for you. A fleet-wide problem should not trigger fleet-wide reprovisioning. That is your cue to take over: check the [node conditions](/docs/hetzner/apalla/servers-and-nodes/access/node-problem-detection-and-conditions) to see what is happening, [pause remediation](#pause-remediation-during-maintenance) on any node you are working on, and fix the root cause before you let replacement resume. Before a node is removed its pods are drained, and the drain respects PodDisruptionBudgets, so set one on any workload that must not lose too many replicas at once. See [Self-healing and node replacement](/docs/hetzner/apalla/concepts/operations/self-healing-and-node-replacement). ## Self-healing does not repair hardware Remediation swaps a node out. It does not fix the physical machine underneath. On bare metal this matters most when a pool's host selector is narrow. If every host the selector matches is already in use, a machine that fails is released and then re-claims the **same** host, because no other host qualifies. The server is wiped and reinstalled, and a fault in the hardware itself comes straight back. Re-provisioning clears a software or filesystem problem. It does nothing for a failing disk, bad memory, or a dying NIC. So do not lean on self-healing to route around broken hardware. Detect a hardware fault directly instead: - Alert on the hardware-related [node conditions](/docs/hetzner/apalla/servers-and-nodes/access/node-problem-detection-and-conditions). On bare metal these page a human rather than trigger a reboot, precisely because a reboot would not help. - Or run your own observability against the servers (SMART data, memory, temperature, link state) and alert from there. When a server is genuinely broken, take it out of the pool: set `maintenanceMode: true` on its `HetznerBareMetalHost` (or remove it), and open a Hetzner support case for the hardware. Alerting on the node conditions is the reliable way to catch a hardware fault. ## Pause remediation during maintenance When you plan to disrupt a node yourself (a firmware update, an in-place reboot, hardware work), pause remediation so Syself Autopilot does not mistake your maintenance for a failure. Pausing a `Machine` stops the management cluster from reconciling it at all, not only from remediating it. Annotate the Machine In the management cluster: ```console $ kubectl annotate machine -n cluster.x-k8s.io/paused=true ``` Verify the annotation is set Before you start work, check the value: ```console $ kubectl get machine -n \ -o jsonpath='{.metadata.annotations.cluster\.x-k8s\.io/paused}' ``` The output should be `true`. Do your maintenance The Machine is not touched while it is paused, so do the firmware update, reboot, or hardware work now. Remove the pause when you are done ```console $ kubectl annotate machine -n cluster.x-k8s.io/paused- ``` > [!WARNING] > Double-check the annotation for typos. If it is not set correctly, the node can be flagged unhealthy and reprovisioned while you work on it. ## When remediation should not fire Pause first for anything that makes a healthy node look unhealthy on purpose: rebooting it, taking it off the network, or running long disruptive maintenance. Leave remediation on for everything else. A node that genuinely fails should be replaced, and that is the whole point of the system. ## Related - [Node problem detection and conditions](/docs/hetzner/apalla/servers-and-nodes/access/node-problem-detection-and-conditions) - [Node health conditions](/docs/hetzner/apalla/observability/reference/node-health-conditions) - [Self-healing and node replacement](/docs/hetzner/apalla/concepts/operations/self-healing-and-node-replacement) - [Remove a specific node](/docs/hetzner/apalla/servers-and-nodes/maintenance/remove-a-specific-node) - [Verify node integrity](/docs/hetzner/apalla/security/verify-node-integrity) - [Node and OS security](/docs/hetzner/apalla/security/node-and-os-security) - [Platform labels and annotations](/docs/hetzner/apalla/servers-and-nodes/labels/platform-labels-and-annotations) - [Ports and listeners](/docs/hetzner/apalla/security/ports-and-listeners)