Skip to main content

Alert on dropped packets

A rise in dropped packets is a signal worth paging on: a workload is being blocked, or something is probing it. Alert on hubble_drop_total, and keep a queryable log of the drops so you can go from "drops are up" to "which pod, to where, blocked by what." This assumes you scrape the Hubble metrics from .

Alert on the drop rate

hubble_drop_total counts dropped packets, and its reason and protocol labels tell you what kind. Compare the current rate against the cluster's own recent rate, so the constant background of denied internet traffic does not fire the alert:

hubble-drop-rules.yamlyaml
		apiVersion: monitoring.coreos.com/v1
kind: PrometheusRule
metadata:
  name: hubble-drops
  namespace: monitoring
  labels:
    prometheus: main
spec:
  groups:
    - name: hubble-drops
      rules:
        - alert: HubbleDropsRising
          expr: |
            sum(rate(hubble_drop_total[5m])) by (reason, protocol)
              > 2 * sum(rate(hubble_drop_total[1h] offset 1h)) by (reason, protocol)
          for: 10m
          labels: {severity: warning}
          annotations:
            summary: "packet drops doubled ({{ $labels.reason }}, {{ $labels.protocol }})"
            description: "Drops are running at more than twice the rate of the previous hour. Check whether a policy is blocking a legitimate connection, or something is probing the cluster."
	

Policy misconfig or intrusion

The alert only tells you that drops went up and not what caused them. A host firewall blocking an internet scan and a network policy blocking your own workload look identical in the metric. To find the source, open the flow log or run hubble observe:

  • A policy misconfig: the source is one of your own workloads and the connection is legitimate. Fix it in .
  • A probe or attack: the source is outside the cluster, or a workload is reaching an external destination it never used before.

Keep a queryable flow log

Hubble keeps flows in memory, so the history is gone when the agent restarts or the node is replaced. Turn on Hubble flow export to keep a durable record you can query, and ship the file like any other log.

There is no topology variable for this, so configure it with a CiliumNodeConfig object. It leaves the cilium-config ConfigMap that Syself Autopilot manages untouched, so the setting survives an upgrade.

hubble-flow-export.yamlyaml
		apiVersion: cilium.io/v2
kind: CiliumNodeConfig
metadata:
  namespace: kube-system
  name: hubble-flow-export
spec:
  nodeSelector:
    matchLabels: {} # Apply to all nodes
  defaults:
    hubble-export-file-max-size-mb: "50"
    hubble-export-file-max-backups: "5"
    hubble-export-file-path: "/var/run/cilium/hubble/events.log"
    # Keep the volume down: export these verdicts, not every packet. The key is
    # "allowlist", one word; "allow-list" is silently ignored and you get
    # everything. The value is a bare JSON object, never a JSON array.
    hubble-export-allowlist: '{"verdict":["DROPPED","ERROR"]}'
	

Apply the configuration and restart the Cilium agents to load it:

		$ kubectl apply -f hubble-flow-export.yaml
$ kubectl -n kube-system rollout restart daemonset/cilium
	

Then tail the file like any other log source. The already mounts /var/run/cilium/hubble/ and has a source block for events.log ready to enable.

Each line is a JSON flow with source and destination identity, port, and verdict, so you can ask which pod was denied reaching which service, and by which policy. Widen the allow list to include FORWARDED for a full connection log rather than only drops.

Note

Flow export produces high volume. A busy node produces a large stream even filtered to drops, and widening it to FORWARDED multiplies that by every connection in the cluster. Size the retention on this stream separately from your other logs, and treat it as a diagnostic you turn up during an investigation rather than something to keep for a year.