Skip to main content

Log retention and sizing

Inspect 1.36

On-node logs are bounded on purpose: each stream rotates at a fixed cap, deleting old entries as new ones arrive, so a busy node cannot fill its own /var. Those caps tell you how much history the node holds before it rotates away, which is your deadline for shipping logs off-node, and they are the starting point for sizing your own store.

The on-node caps#

Stream Cap Detail
auditd ~400 MB 50 MB per file, 8 files kept
Kubernetes API audit ~1 GB or 90 days 100 MB per file, 10 files, whichever limit comes first
journald 1 GB total SystemMaxUse=1G, 50 MB per journal file
Container logs rotated by the kubelet a size cap per container with a few rotated files kept

These are hot buffers, not archives. On a busy node the API audit log can reach its cap in well under 90 days, so do not treat "90 days" as a retention guarantee; treat the off-node store as the real retention.

Estimate ingest per node#

To size your store, estimate what one node produces per day, then multiply by node count and add headroom for incident-time spikes:

  • Container logs dominate, and vary entirely with your workloads. Measure a representative node rather than guessing.
  • journald is steady and modest once you filter to the units you keep.
  • Audit is bounded by the caps above; the API audit stream is the largest of the three on a busy control plane.

A spike matters more than the average: during an incident, error logging can jump several-fold, so size ingesters and buffers for the peak, not the mean.

Retention tiers in your store#

Keep tiers rather than one flat retention, the same idea as metrics:

  • Hot, queryable logs for days to weeks, for live troubleshooting.
  • Compacted, cheap logs for the long window compliance requires, on object storage. Loki's compactor merges and applies retention per tier.

Cost versus compliance#

The two pulls are opposite: cost wants short retention and aggressive filtering, compliance wants the audit trail kept for a year or more. Resolve it by treating the two kinds of logs differently. Keep operational logs (container, journald) hot for a short window and drop the noisy ones, and keep the three audit streams for the full compliance window in cheap, write-once object storage. That way you pay for long retention only on the small, high-value streams, not on everything. See .