SLOs and error budgets
An alert on "CPU over 80%" pages you for something that may not matter. An alert on "we are burning through our error budget" pages you for something a customer experiences.
This page covers the second kind. You decide what the service should deliver, measure it from your own metrics, and page only when you are missing that target fast enough to matter.
SLI, SLO, and error budget
- An SLI (service-level indicator) is a measurement of the service, from your own metrics: the fraction of requests served under 300 ms, or the fraction that returned success.
- An SLO (service-level objective) is the target for that indicator over a window: 99.9% of requests succeed over 30 days.
- The error budget is what the SLO allows you to miss: at 99.9%, 0.1% of requests can fail before you are out of budget. As long as budget remains, you are within target.
Build SLIs from app metrics
You need the app's own request metrics for this, so instrument it first (Custom application metrics ). Two common SLIs:
- Availability:
sum(rate(http_requests_total{code!~"5.."}[5m])) / sum(rate(http_requests_total[5m])) - Latency: the fraction of requests under your threshold, from a histogram:
histogram_quantileor a bucket ratio.
Precompute these with recording rules so the dashboards and alerts read a cheap pre-aggregated series instead of recomputing a heavy query every evaluation.
Alert on burn rate, over multiple windows
Burn rate is how fast you are spending the budget. A burn rate of 1 uses the whole budget exactly at the end of the SLO window, so on a 30-day SLO it lasts the full 30 days. A burn rate of 14 spends it 14 times faster, so the same budget is gone in about two days.
Alert when a fast burn holds over a short window and a slower burn holds over a longer one, so a brief spike does not page and a sustained problem does.
Writing these rules by hand is error-prone, because one SLO expands into a dozen or more recording and alerting rules. Sloth is an open-source generator for them: you declare the objective in a short SLO spec, and it emits the finished PrometheusRule.
Dashboards and fatigue
Put the budget on a dashboard: how much is left, and the burn rate over time. A team that can see the budget can release faster while there is room and slow down when it runs low. Only page on things tied to an SLO a customer cares about. Everything else is a ticket or a dashboard, not a page.
Platform alert rules
Some conditions never trigger self-healing on bare-metal pools, so alerting is the only way to act on node integrity and capacity, and this page ships those rules.
Silences and grouping
Group related alerts into one notification and silence known noise during maintenance, so on-call is not buried when a node pool rolls.