Measuring Availability Honestly: SLIs, Error Budgets and SLO Debt

Measuring Availability Honestly: SLIs, Error Budgets and SLO Debt

Reading time1 min
#devops#slo#availability#engineering#cloud

Measuring Availability Honestly: SLIs, Error Budgets and SLO Debt

A dashboard can show 99.9% availability for a month in which users had a bad time. Usually the number is measuring the wrong thing, or the SLO exists on paper and does not affect any decisions. SLO debt is the gap that builds up when the published target and the real reliability drift apart and nothing changes.

What 99.9% allows

For a 30-day window, 99.9% means 43.2 minutes of full downtime. Over a year it is about 8.8 hours. For a request-based SLO it means one failed request in a thousand. That is a tight budget, and it is easy to stay inside it on paper while missing it in reality.

How availability numbers get inflated

  • Measuring the process, not the requests. up == 0 or a ping on /health tells you the process answers. A service that returns 500 for checkout and 200 for /health looks fully available.
  • Averaging across endpoints. High-volume cheap requests such as static assets and health checks dilute failures of the requests that matter.
  • Counting only server-side 5xx. Load balancer errors, client timeouts, and requests that never arrived because of DNS, TLS or a regional outage are invisible to the application.
  • Ignoring latency. A checkout that takes 30 seconds is a failure for the user, even with a 200 status.
  • Exclusions. "Planned maintenance" and "third-party outage" removed from the number, while users still saw the errors.
  • Dependencies. A request that needs several services in series cannot be more available than their product. Three dependencies at 99.9% each give at most about 99.7% if their failures are independent.

Define SLIs from the user's side

An SLI is good events divided by valid events. Pick a few critical user journeys, such as login, search and checkout. Give each an availability SLI and a latency SLI, and measure them as close to the user as you can: the load balancer or ingress, not only the application. Ingress-nginx, for example, exposes request counts by status and a request duration histogram.

Availability as a recording rule:

groups:
  - name: slo-checkout
    rules:
      - record: slo:sli_error:ratio_rate5m
        expr: |
          sum(rate(http_requests_total{job="my-app", route="/api/checkout", code=~"5.."}[5m]))
          /
          sum(rate(http_requests_total{job="my-app", route="/api/checkout"}[5m]))
        labels:
          slo: checkout-availability

Latency as the share of requests faster than a threshold that is also a histogram bucket boundary:

sum(rate(http_request_duration_seconds_bucket{job="my-app", route="/api/checkout", le="0.5"}[5m]))
/
sum(rate(http_request_duration_seconds_count{job="my-app", route="/api/checkout"}[5m]))

Writing these rules for several windows and services by hand gets repetitive. Sloth and Pyrra generate them from a short SLO definition, and keeping that definition in git gives you review and history.

Error budget and burn rate

The error budget is 1 - SLO, so 0.1% for a 99.9% target. Burn rate is the observed error ratio divided by the budget. A burn rate of 1 uses exactly the whole budget over the SLO window.

Remaining budget for the last 30 days:

1 - (
  sum(increase(http_requests_total{job="my-app", route="/api/checkout", code=~"5.."}[30d]))
  /
  sum(increase(http_requests_total{job="my-app", route="/api/checkout"}[30d]))
) / 0.001

A 30-day range over raw counters is expensive, so in practice build it from recording rules.

For alerting, the Google SRE Workbook recommends multi-window, multi-burn-rate alerts. For a 30-day SLO: page on a burn rate of 14.4 over 1 hour (2% of the budget gone in an hour) or 6 over 6 hours (5%), each confirmed by a short window so the alert resolves quickly. Open a ticket on a burn rate of 1 over 3 days.

- alert: CheckoutErrorBudgetBurn
  expr: |
    (
      slo:sli_error:ratio_rate1h{slo="checkout-availability"} > (14.4 * 0.001)
      and
      slo:sli_error:ratio_rate5m{slo="checkout-availability"} > (14.4 * 0.001)
    )
    or
    (
      slo:sli_error:ratio_rate6h{slo="checkout-availability"} > (6 * 0.001)
      and
      slo:sli_error:ratio_rate30m{slo="checkout-availability"} > (6 * 0.001)
    )
  labels:
    severity: page

This assumes recording rules for the 30m, 1h and 6h windows built the same way as the 5m one.

When the budget keeps running out

If the budget is gone month after month and nothing changes, the SLO is decoration. There are two honest ways out.

Repay the debt. Agree on an error budget policy with product owners before you need it. For example: while the budget is exhausted, risky releases pause and the team works on the largest sources of errors until the SLI recovers. Postmortem action items get owners and due dates, and someone tracks them.

Reset the target. If the system cannot deliver 99.9% and users are fine with less, publish a lower target that you actually meet, and plan the work to raise it. An SLO you always miss carries less information than a lower one you hit.

Contractual SLAs should stay looser than the internal SLO, so you have room to react before a contract is breached.

Find what consumes the budget

Break budget consumption down by route, then by cause:

topk(5, sum by (route) (increase(http_requests_total{job="my-app", code=~"5.."}[7d])))

Typical work that buys budget back: canary releases with automatic rollback on burn rate (Argo Rollouts and Flagger can do this), timeouts and circuit breakers around dependencies, removing single points of failure, and keeping capacity headroom for peaks.

Report it

A short monthly report per journey is enough: SLI against target, budget consumed, the incidents that used most of it, and the status of their action items. The value is in the conversation it forces, not in the document.

Checklist

  • SLIs per user journey, measured as close to the user as possible.
  • Count slow requests as bad, not only errors.
  • No silent exclusions. If something is excluded, write down why.
  • Check that the SLO is achievable given the dependencies.
  • Burn-rate alerts instead of raw thresholds.
  • An agreed error budget policy, applied when the budget runs out.
  • Either repay the debt or lower the target. Do not keep publishing a number you miss.