Skip to main content
Kubernetes 1.32 is deprecated
Please follow this guide to upgrade.

Where to start troubleshooting

Start with the object and its events

When you are not sure what is wrong, start with the object that is failing and its events. The event on the object tells you what a controller last tried and why it stopped.

Syself Autopilot repairs many faults on its own by replacing the node, so before you dig in, check whether the object is already recovering. See .

The general approach:

  1. Find the object that is unhealthy: the Cluster, a Machine, a Node, a Pod, or a PVC.
  2. Read its status and its events. The events name the reason and the last action.
  3. If that does not clear it, collect what support needs and open a case.

Custom API server domain: check the DNS records

One provisioning trap looks much worse than it is. You set a custom domain for the API server, created a load balancer for it, and the first control plane came up fine. Then nothing else joins. The cluster looks like it is taking forever to provision.

The cause is almost always DNS. The other nodes reach the API server by that custom domain, so they cannot join until the domain resolves to the load balancer. If the DNS records were never created, or they were created but have not propagated yet, the join hangs.

Check, in order:

  1. The DNS records for the custom domain exist and point at the load balancer's IP.
  2. The records have propagated. Resolve the domain from a machine outside your DNS provider and confirm you get the load balancer's address.

Once the domain resolves, the waiting nodes join on their own. No other action is needed.


Read a Machine or Cluster event

Cluster and Machine objects live on the management cluster. Their events show what the controller last tried on your cluster and its nodes.

Note

Run these against the management cluster kubeconfig, where the Cluster and Machine objects live. Node, Pod, and PVC objects live on the workload cluster.

  1. List the objects and their phase:

    				$ kubectl get cluster,machine -n <namespace>
    			
  2. Describe the object and read the Events section at the bottom:

    				$ kubectl describe machine <name> -n <namespace>
    # or
    $ kubectl describe cluster <name> -n <namespace>
    			
  3. Sort every recent event in the namespace by time:

    				$ kubectl get events -n <namespace> --sort-by=.lastTimestamp
    			
  4. Note the exact Reason and Message, and the object they sit on. That string is what support needs first.


Pause automation while you investigate

Sometimes the automation gets in your way. You are looking at a failing Machine, and the controllers keep retrying, or self-healing reboots and replaces the node out from under you before you can read what went wrong. Pause it so it holds still.

Set the cluster.x-k8s.io/paused annotation on the object you want frozen:

				# Freeze one Machine (stops its controllers, including automatic remediation):
$ kubectl annotate machine <name> -n <namespace> cluster.x-k8s.io/paused=true
 
# Or freeze the whole cluster:
$ kubectl annotate cluster <name> -n <namespace> cluster.x-k8s.io/paused=true
			

While an object is paused, its controllers stop reconciling it. Automatic remediation and replacement do not run, so a paused unhealthy node stays put instead of being rebooted or replaced. Nothing is deprovisioned and no workloads move. The object simply stops changing while you look at it.

Resume by removing the cluster.x-k8s.io/paused annotation:

				$ kubectl patch machine <name> -n <namespace> --type=json \
  -p='[{"op":"remove","path":"/metadata/annotations/cluster.x-k8s.io~1paused"}]'
			

Alternatively, run kubectl edit machine <name> -n <namespace> and delete the annotation line manually.

Warning

Pausing turns off self-healing for that object. It does not fix anything by itself, and a node left paused will not recover on its own. Pause to investigate, then resume as soon as you are done.

Pausing is different from maintenanceMode, which deprovisions a bare-metal host and reschedules its workloads. Use pause when you want the automation to stop touching a node without removing it. See for the maintenance-mode path and when each fits.


When to stop and contact support

Stop and open a case when:

  • You have read the events and the object is still unhealthy.
  • The control plane or etcd is at risk (API unreachable, quorum in doubt) and you should not act alone.
  • There is a risk of data loss.
  • The fix needs Syself to act on the node operating system or the bare-metal server itself.

Email support with the cluster name, the object, and the exact Reason and Message from its events. See .