Skip to main content

Metrics per component

Inspect 1.36

Every Syself Autopilot component reports its own metrics on an HTTP endpoint. The address it binds decides how you scrape it and whether it is ever reachable from the internet.

The endpoints fall into two groups. A component on the pod network binds a pod IP and sits behind a ClusterIP Service, so an ordinary scrape reaches it. The rest bind 127.0.0.1 on the node, so nothing outside that machine can reach them. The kubelet is the exception: it binds every interface, and the host firewall is what keeps it private.

127.0.0.1 is the loopback address, and binding to it keeps an endpoint off the network. Inside a pod, 127.0.0.1 means the pod itself, not the node, so an ordinary scrape cannot reach a node's loopback endpoint. A host-network collector runs on the node itself, so it can.

Every endpoint on this page is collected by the , which runs on each node's own network namespace and therefore reaches both groups. Your own workloads are the exception: they go to the instead, for the reasons in .

This page is the inventory. For what the series from each component actually tell you, see and .

Health-daemon metrics

A health daemon runs on every node. It turns node problems into Kubernetes NodeConditions and events, and it exports Prometheus metrics on 127.0.0.1:20257. The endpoint is plain HTTP with no auth, which is why it stays on loopback.

Most of the problems the daemon detects already show up as NodeConditions, and alerting on those is the normal path. For advanced use cases you can scrape the metrics shown below.

Metric Type What it tells you
problem_counter counter How many times a problem fired, with a reason label for the kind. For a NodeCondition it moves at the same moment problem_gauge is set, so alert on the gauge instead. What only this metric shows is the kernel events that set no condition at all, such as IOError or MemoryReadError. on the conditions page covers them.
syself_node_unit_restarts{unit} gauge, per unit systemd's restart count since boot for each watched unit. Every node has kubelet.service, containerd.service and syself-tunnel-agent.service. Workers also have syself-proxy.service. Control planes also have syself-tunnel-server.service and syself-auditsrc.service. Three or more restarts of kubelet.service or containerd.service between two health checks set the ServiceNotRecovering condition and so does the unit entering systemd's failed state. On a worker syself-proxy.service does the same. The condition means systemd's own restarts are not fixing the problem. The other units set NodeServiceDown instead, because their failure does not make the node unhealthy.
problem_gauge gauge Which conditions are currently set on the node. The type label is the condition, the reason label is the cause, and a value of 1 means it is set. Alert on it directly, for example problem_gauge{type="DisksFailure"} == 1.
Tip

Read the condition first with the health report annotations, then correlate with the metric. See for the one-liners that print the autopilot.syself.com/* report.

Reachable on the pod network

These expose metrics through a Service or their pod IP, so any collector on the pod network reaches them. The System Alloy scrapes the ones running on its own node, which spreads the work across agents without scraping any endpoint twice.

Component Service Port What it tells you
kube-apiserver the kubernetes Service 443 Request rates, latencies, and error codes. Scrape /metrics with a bearer token that has get on nonResourceURLs: ["/metrics"]. It also listens on each control plane's loopback, which is where the System Alloy scrapes it.
CCM ccm-metrics (headless) 8233 Hetzner API call counts and errors, rate-limit headroom, load-balancer reconcile results.
hubble-relay hubble-relay-metrics 9966 Relay health, connected peers, gRPC traffic.
CSI controller csi-controller-metrics (headless) 9189 Volume provision, attach, detach, and resize results, and the Hetzner API calls behind them.
CoreDNS kube-dns 9153 DNS query and error rates.
metrics-server metrics-server 443 Usually read through the metrics.k8s.io aggregated API, not scraped directly.

Reachable only over the node's loopback

These bind 127.0.0.1, so only a collector sharing the node's network namespace reaches them. The kubelet is the exception: it binds every interface, so it answers on both the node IP and loopback, but the host firewall admits port 10250 only from other nodes and the metrics-server pod, so an ordinary Prometheus pod is dropped. Its serving certificate names the node IP and hostname rather than 127.0.0.1, so a loopback scrape has to skip certificate verification.

The Nodes column is also the gating rule: a job for a control-plane-only or worker-only component has to be restricted to those nodes, or it fails permanently on the rest.

Component Loopback port Scheme Nodes
etcd 2381 HTTP, no auth control planes
kube-controller-manager 10257 HTTPS, bearer token control planes
kube-scheduler 10259 HTTPS, bearer token control planes
kubelet 10250 HTTPS, bearer token every node (SAN is the node IP, so the scraper uses insecure_skip_verify)
cadvisor 10250 HTTPS, bearer token every node; the kubelet's /metrics/cadvisor path, for per-container CPU, memory, and network
KubeGate 8080 HTTP, no auth control planes
Cilium agent 9962 HTTP every node
Cilium operator 9963 HTTP only nodes running it; others refuse the connection
Hubble 9965 HTTP every node
CSI node plugin 9189 HTTP cloud workers only; no Service, no container port. Host network because it reads the Hetzner metadata service at 169.254.169.254, which a pod network namespace cannot reach; it binds metrics to loopback so they stay off the worker's public IP
syself-agent (health daemon) 20257 HTTP every node
syself-tunnel-agent 8182 HTTP every node
syself-tunnel-server 8181 HTTP control planes
syself-proxy (failover) 9587 HTTP workers
containerd 1338 HTTP every node; serves /v1/metrics
Note

Several loopback endpoints carry no authentication (etcd on 2381, KubeGate on 8080, syself-agent on 20257). The host-network agent that reaches them is a privileged collector: keep it in a namespace only your platform team can write to, and review its scrape config like a firewall rule.

Important

The Cilium operator, hubble-relay, and the CSI controller each declare more than one container port, and Kubernetes pod discovery emits one target per port. A config that rewrites the address to a fixed metrics port without first selecting that port scrapes the endpoint twice and doubles every rate computed from it.

Why every listener is loopback or authenticated

Nodes carry public IPs and there is no private network. Almost every metrics endpoint is unauthenticated (etcd 2381, KubeGate 8080, the health daemon 20257 and Cilium) and serves plain HTTP with no token, so a host-network process that bound 0.0.0.0 would publish that data on a routable internet-facing address. Binding 127.0.0.1 means the host firewall is not the only thing in front of these endpoints: if a firewall rule is ever misconfigured or fails to apply, a loopback listener is still unreachable.

Warning

If a scrape cannot reach a metrics port, do not open that port in a host-firewall policy (a CiliumClusterwideNetworkPolicy). It will not help. These ports bind 127.0.0.1, so nothing is listening on the node's routable address and the connection is refused either way. Scrape from a host-network DaemonSet instead.

What to scrape and how

Scrape the loopback endpoints from a host-network DaemonSet. Set hostNetwork: true and dnsPolicy: ClusterFirstWithHostNet so the collector shares each node's network namespace and dials 127.0.0.1 directly. That path never leaves the node, so it needs no firewall rule.

A scraper pod that is not on the host network cannot reach these ports. Its packet leaves the pod namespace and arrives at the node's host endpoint, where the default-deny host firewall drops anything not explicitly allowed. That is why a ServiceMonitor pointed at 127.0.0.1 on the node cannot work, and why the loopback endpoints have no Service to select.

The pod-network endpoints are the opposite: they are ordinary ClusterIP Services, reachable from any namespace, and a normal Service or ServiceMonitor picks them up. Syself does not ship the Prometheus Operator CRDs, so if you use ServiceMonitors you install the operator and write your own.

Not scrapeable

  • Cilium Envoy's metrics listener is off. These clusters , so Envoy does no L7 work and its metrics on 9964 would be empty. That listener is also the one the upstream chart cannot bind to loopback, so rather than publish an empty endpoint on the node address, those Envoy metrics are disabled. Nothing useful is lost.
  • GPU nodes ship no GPU metrics on their own. To get GPU utilization and memory, install the NVIDIA DCGM exporter yourself and scrape it like any application metric.