Skip to main content

SLOs and error budgets

An alert on "CPU over 80%" pages you for something that may not matter. An alert on "we are burning through our error budget" pages you for something a customer experiences.

This page covers the second kind. You decide what the service should deliver, measure it from your own metrics, and page only when you are missing that target fast enough to matter.

SLI, SLO, and error budget

  • An SLI (service-level indicator) is a measurement of the service, from your own metrics: the fraction of requests served under 300 ms, or the fraction that returned success.
  • An SLO (service-level objective) is the target for that indicator over a window: 99.9% of requests succeed over 30 days.
  • The error budget is what the SLO allows you to miss: at 99.9%, 0.1% of requests can fail before you are out of budget. As long as budget remains, you are within target.

Build SLIs from app metrics

You need the app's own request metrics for this, so instrument it first ( ). Two common SLIs:

  • Availability: sum(rate(http_requests_total{code!~"5.."}[5m])) / sum(rate(http_requests_total[5m]))
  • Latency: the fraction of requests under your threshold, from a histogram: histogram_quantile or a bucket ratio.

Precompute these with recording rules so the dashboards and alerts read a cheap pre-aggregated series instead of recomputing a heavy query every evaluation.

Alert on burn rate, over multiple windows

Burn rate is how fast you are spending the budget. A burn rate of 1 uses the whole budget exactly at the end of the SLO window, so on a 30-day SLO it lasts the full 30 days. A burn rate of 14 spends it 14 times faster, so the same budget is gone in about two days.

Alert when a fast burn holds over a short window and a slower burn holds over a longer one, so a brief spike does not page and a sustained problem does.

Writing these rules by hand is error-prone, because one SLO expands into a dozen or more recording and alerting rules. Sloth⁠ is an open-source generator for them: you declare the objective in a short SLO spec, and it emits the finished PrometheusRule.

Dashboards and fatigue

Put the budget on a dashboard: how much is left, and the burn rate over time. A team that can see the budget can release faster while there is room and slow down when it runs low. Only page on things tied to an SLO a customer cares about. Everything else is a ticket or a dashboard, not a page.