Skip to main content

Plan maintenance and run prechecks

Run these checks before any cluster upgrade, whether you are bumping the Cluster Stack release within a minor version or moving to a new Kubernetes minor version. It is much easier to find a problem before the rollout starts than to diagnose it while the rollout is running.

Prerequisites

  • Access to the management cluster and the workload cluster ( configured for both)

Prechecks

Verify control plane health

A cluster should always have healthy control planes, especially before an upgrade. The rollout does not start while a control plane node is unhealthy. It waits until that node is healthy again and then proceeds.

Check control plane node status:

		$ kubectl get nodes --selector='node-role.kubernetes.io/control-plane' -o wide
	

All control plane nodes must be Ready. If any are NotReady or show a problem condition, resolve that first.

Check for pending certificate signing requests

Pending CSRs (Certificate Signing Requests) can stall node join during a rollout. New nodes request certificates when they boot; if those requests are not approved quickly, the node takes longer to join or may time out.

List pending CSRs on the workload cluster (the cluster the new nodes are joining, not the management cluster):

		$ kubectl get csr
	

Any CSR in Pending state that has been waiting for more than a few minutes needs attention before you start a rollout. Approve legitimate ones or investigate why they are stalled.

Review PodDisruptionBudgets

The upgrade drain (moving all pods off a node before it is replaced) respects PodDisruptionBudgets (PDBs). A PDB limits how many pods of a group can be unavailable at once. If a PDB is configured too strictly, the drain stalls and the rollout waits.

List all PDBs in the workload cluster:

		$ kubectl get pdb -A
	

For each PDB, check:

  • ALLOWED DISRUPTIONS is at least 1. If it is 0, the drain cannot proceed.
  • The PDB actually has matched pods. A PDB with no matching pods causes no problem, but review it anyway.
		$ kubectl describe pdb -A
	

A PDB with maxUnavailable: 0, or minAvailable equal to the number of replicas, blocks the drain. When that happens, the system waits up to 180 seconds for the drain to finish, then gives up and moves on. This timeout is set in the cluster configuration (nodeDrainTimeoutSeconds: 180). After the timeout, the stack removes the node anyway, which can disrupt the pods that were still running on it. Adjust PDB settings to allow at least one pod to be unavailable per group, or increase replicas so the PDB can be satisfied during the drain.

Tip

PDBs are the right tool for protecting availability during upgrades. Set maxUnavailable: 1 (or a percentage that leaves at least one pod running) and run enough replicas that the PDB can be satisfied during a drain. A PDB with maxUnavailable: 0 across all replicas blocks upgrades entirely.

Verify worker node capacity for bare metal pools

Cloud worker pools use maxSurge: 1: the stack creates the new node first, then removes the old one. Total capacity does not drop during the upgrade.

Bare metal worker pools use maxSurge: 0: the stack frees the old physical host first, then reprovisions it. During each step, your pool runs one fewer worker node.

For bare metal pools, check that your workloads can tolerate one fewer node per pool during the upgrade:

		$ kubectl get nodes --selector='node.kubernetes.io/worker=true' -o wide
$ kubectl top nodes
	

If resource usage is high, either scale up the pool before the upgrade (temporarily) or wait for a lower-traffic period.

Confirm all nodes are Ready

Start upgrades only when all nodes are Ready. An existing unhealthy node complicates rollout tracking and may trigger automatic node replacement during the upgrade.

		$ kubectl get nodes -o wide
	

If a node is NotReady, investigate and resolve the problem before starting the upgrade.

The checks below apply only when you move to a new Kubernetes minor version, specifically the move to 1.36. Skip them for a stack version bump within the same minor.

Take a backup before you start

The upgrade to 1.36 replaces the node operating system, and once the roll is under way it cannot be rolled back (see ). Take a backup first so you can rebuild if something goes wrong. Capture two layers:

  • GitOps state. Confirm the cluster manifests and application manifests are committed and the repository is current. Re-syncing from Git is how you recover anything Git owns.
  • Runtime state and volume data. Back up workloads and PersistentVolumeClaims that Git cannot reconstruct (data, hand-applied objects) with Velero.

See .

Check your node ceiling if you use a custom pod CIDR

Kubernetes 1.36 gives each node a /24 pod subnet, about 254 usable addresses, which is what the per-node maxPods: 220 cap is sized against. The most nodes a cluster can hold is set by how many bits sit between your pod CIDR prefix and /24:

text
		node ceiling = 2 ^ (24 − pod-CIDR prefix length)
	

The default pod CIDR is wide enough that this is not a concern (a /11 pod CIDR holds 2 ^ (24 − 11) = 8192 nodes). A narrow custom pod CIDR lowers the ceiling: a /16 holds 2 ^ (24 − 16) = 256 nodes, a /18 holds 64. Check your pod CIDR:

		$ kubectl get cluster <name> -o jsonpath='{.spec.clusterNetwork.pods.cidrBlocks}'
	

If your node count is near the ceiling, plan your growth around it, or talk to us about widening the pod CIDR before you upgrade.