Metrics per component
Every Syself Autopilot component reports its own metrics on an HTTP endpoint. The address it binds decides how you scrape it and whether it is ever reachable from the internet.
The endpoints fall into two groups. A component on the pod network binds a pod IP and sits behind a ClusterIP Service, so an ordinary scrape reaches it. The rest bind 127.0.0.1 on the node, so nothing outside that machine can reach them. The kubelet is the exception: it binds every interface, and the host firewall is what keeps it private.
127.0.0.1 is the loopback address, and binding to it keeps an endpoint off the network. Inside a pod, 127.0.0.1 means the pod itself, not the node, so an ordinary scrape cannot reach a node's loopback endpoint. A host-network collector runs on the node itself, so it can.
Every endpoint on this page is collected by the System Alloy , which runs on each node's own network namespace and therefore reaches both groups. Your own workloads are the exception: they go to the Application Alloy instead, for the reasons in Using Alloy for observability .
This page is the inventory. For what the series from each component actually tell you, see Control-plane metrics and Data-plane and cluster component metrics .
Health-daemon metrics
A health daemon runs on every node. It turns node problems into Kubernetes NodeConditions and events, and it exports Prometheus metrics on 127.0.0.1:20257. The endpoint is plain HTTP with no auth, which is why it stays on loopback.
Most of the problems the daemon detects already show up as NodeConditions, and alerting on those is the normal path. For advanced use cases you can scrape the metrics shown below.
| Metric | Type | What it tells you |
|---|---|---|
problem_counter | counter | How many times a problem fired, with a reason label for the kind. For a NodeCondition it moves at the same moment problem_gauge is set, so alert on the gauge instead. What only this metric shows is the kernel events that set no condition at all, such as IOError or MemoryReadError. Events only on the conditions page covers them. |
syself_node_unit_restarts{unit} | gauge, per unit | systemd's restart count since boot for each watched unit. Every node has kubelet.service, containerd.service and syself-tunnel-agent.service. Workers also have syself-proxy.service. Control planes also have syself-tunnel-server.service and syself-auditsrc.service. Three or more restarts of kubelet.service or containerd.service between two health checks set the ServiceNotRecovering condition and so does the unit entering systemd's failed state. On a worker syself-proxy.service does the same. The condition means systemd's own restarts are not fixing the problem. The other units set NodeServiceDown instead, because their failure does not make the node unhealthy. |
problem_gauge | gauge | Which conditions are currently set on the node. The type label is the condition, the reason label is the cause, and a value of 1 means it is set. Alert on it directly, for example problem_gauge{type="DisksFailure"} == 1. |
Tip
Read the condition first with the health report annotations, then correlate with the metric. See Node health conditions for the one-liners that print the autopilot.syself.com/* report.
Reachable on the pod network
These expose metrics through a Service or their pod IP, so any collector on the pod network reaches them. The System Alloy scrapes the ones running on its own node, which spreads the work across agents without scraping any endpoint twice.
| Component | Service | Port | What it tells you |
|---|---|---|---|
| kube-apiserver | the kubernetes Service | 443 | Request rates, latencies, and error codes. Scrape /metrics with a bearer token that has get on nonResourceURLs: ["/metrics"]. It also listens on each control plane's loopback, which is where the System Alloy scrapes it. |
| CCM | ccm-metrics (headless) | 8233 | Hetzner API call counts and errors, rate-limit headroom, load-balancer reconcile results. |
| hubble-relay | hubble-relay-metrics | 9966 | Relay health, connected peers, gRPC traffic. |
| CSI controller | csi-controller-metrics (headless) | 9189 | Volume provision, attach, detach, and resize results, and the Hetzner API calls behind them. |
| CoreDNS | kube-dns | 9153 | DNS query and error rates. |
| metrics-server | metrics-server | 443 | Usually read through the metrics.k8s.io aggregated API, not scraped directly. |
Reachable only over the node's loopback
These bind 127.0.0.1, so only a collector sharing the node's network namespace reaches them. The kubelet is the exception: it binds every interface, so it answers on both the node IP and loopback, but the host firewall admits port 10250 only from other nodes and the metrics-server pod, so an ordinary Prometheus pod is dropped. Its serving certificate names the node IP and hostname rather than 127.0.0.1, so a loopback scrape has to skip certificate verification.
The Nodes column is also the gating rule: a job for a control-plane-only or worker-only component has to be restricted to those nodes, or it fails permanently on the rest.
| Component | Loopback port | Scheme | Nodes |
|---|---|---|---|
| etcd | 2381 | HTTP, no auth | control planes |
| kube-controller-manager | 10257 | HTTPS, bearer token | control planes |
| kube-scheduler | 10259 | HTTPS, bearer token | control planes |
| kubelet | 10250 | HTTPS, bearer token | every node (SAN is the node IP, so the scraper uses insecure_skip_verify) |
| cadvisor | 10250 | HTTPS, bearer token | every node; the kubelet's /metrics/cadvisor path, for per-container CPU, memory, and network |
| KubeGate | 8080 | HTTP, no auth | control planes |
| Cilium agent | 9962 | HTTP | every node |
| Cilium operator | 9963 | HTTP | only nodes running it; others refuse the connection |
| Hubble | 9965 | HTTP | every node |
| CSI node plugin | 9189 | HTTP | cloud workers only; no Service, no container port. Host network because it reads the Hetzner metadata service at 169.254.169.254, which a pod network namespace cannot reach; it binds metrics to loopback so they stay off the worker's public IP |
syself-agent (health daemon) | 20257 | HTTP | every node |
syself-tunnel-agent | 8182 | HTTP | every node |
syself-tunnel-server | 8181 | HTTP | control planes |
syself-proxy (failover) | 9587 | HTTP | workers |
| containerd | 1338 | HTTP | every node; serves /v1/metrics |
Note
Several loopback endpoints carry no authentication (etcd on 2381, KubeGate on 8080, syself-agent on 20257). The host-network agent that reaches them is a privileged collector: keep it in a namespace only your platform team can write to, and review its scrape config like a firewall rule.
Important
The Cilium operator, hubble-relay, and the CSI controller each declare more than one container port, and Kubernetes pod discovery emits one target per port. A config that rewrites the address to a fixed metrics port without first selecting that port scrapes the endpoint twice and doubles every rate computed from it.
Why every listener is loopback or authenticated
Nodes carry public IPs and there is no private network. Almost every metrics endpoint is unauthenticated (etcd 2381, KubeGate 8080, the health daemon 20257 and Cilium) and serves plain HTTP with no token, so a host-network process that bound 0.0.0.0 would publish that data on a routable internet-facing address. Binding 127.0.0.1 means the host firewall is not the only thing in front of these endpoints: if a firewall rule is ever misconfigured or fails to apply, a loopback listener is still unreachable.
Warning
If a scrape cannot reach a metrics port, do not open that port in a host-firewall policy (a CiliumClusterwideNetworkPolicy). It will not help. These ports bind 127.0.0.1, so nothing is listening on the node's routable address and the connection is refused either way. Scrape from a host-network DaemonSet instead.
What to scrape and how
Scrape the loopback endpoints from a host-network DaemonSet. Set hostNetwork: true and dnsPolicy: ClusterFirstWithHostNet so the collector shares each node's network namespace and dials 127.0.0.1 directly. That path never leaves the node, so it needs no firewall rule.
A scraper pod that is not on the host network cannot reach these ports. Its packet leaves the pod namespace and arrives at the node's host endpoint, where the default-deny host firewall drops anything not explicitly allowed. That is why a ServiceMonitor pointed at 127.0.0.1 on the node cannot work, and why the loopback endpoints have no Service to select.
The pod-network endpoints are the opposite: they are ordinary ClusterIP Services, reachable from any namespace, and a normal Service or ServiceMonitor picks them up. Syself does not ship the Prometheus Operator CRDs, so if you use ServiceMonitors you install the operator and write your own.
Not scrapeable
- Cilium Envoy's metrics listener is off. These clusters block L7 network policy , so Envoy does no L7 work and its metrics on
9964would be empty. That listener is also the one the upstream chart cannot bind to loopback, so rather than publish an empty endpoint on the node address, those Envoy metrics are disabled. Nothing useful is lost. - GPU nodes ship no GPU metrics on their own. To get GPU utilization and memory, install the NVIDIA DCGM exporter yourself and scrape it like any application metric.
Related
- Ports and listeners : every port a node opens, not just the metrics ones.
- Machine health checks and remediation : which conditions trigger replacement, and the window each one has to hold for.
- Collect logs with the log collector : the diagnostics bundle to send when a metric points at a problem you cannot resolve.
Data-plane and cluster component metrics
The platform's networking, DNS, storage, and cloud-integration cluster components each report their own metrics, some on loopback and some on the pod network, and this page covers what each one tells you.
Node health conditions
Every NodeCondition the health daemon can set, what triggers each one, how to read them and the health report on a node, and how to alert on them with kube-state-metrics.