What p99 Latency Hides and How to Measure the Tail Properly

What p99 Latency Hides and How to Measure the Tail Properly

Reading time1 min
#devops#latency#metrics#performance#kubernetes

What p99 Latency Hides and How to Measure the Tail Properly

A common situation: the p99 latency panel is flat and green, and users still say the product is slow. Usually the metric is correct about what it measures. It just measures something narrower than what the user experiences, or it is aggregated in a way that throws the tail away.

This post goes through the usual causes and how to fix each one.

Percentiles cannot be averaged

This query looks reasonable and is wrong:

avg(http_request_duration_seconds{job="my-app", quantile="0.99"})

It takes per-pod p99 values from a summary and averages them. The result is not the p99 of anything. One overloaded pod with a terrible tail gets diluted by the healthy ones. Summaries compute quantiles inside the client, so they cannot be combined across instances correctly. The Go client also computes them over a sliding window (10 minutes by default), which smooths short spikes.

Use histograms and aggregate the buckets before computing the quantile:

histogram_quantile(
  0.99,
  sum by (le) (rate(http_request_duration_seconds_bucket{job="my-app"}[5m]))
)

Bucket boundaries limit what you can see

histogram_quantile does not know the real values. It finds the bucket that contains the quantile and interpolates linearly inside it. With buckets at 0.5 and 1 second, a p99 of "0.8" means "somewhere between 0.5 and 1".

It gets worse at the top. If the quantile falls into the +Inf bucket, histogram_quantile returns the upper bound of the highest finite bucket. With a top bucket of 1 second, requests that take 8 seconds still show up as a p99 of 1 second. The graph looks capped because it is.

Fixes:

  • Put bucket boundaries around your SLO threshold and above your longest timeout.
  • Use native histograms, which have exponential buckets with much finer resolution. They are stable since Prometheus 3.8 and need scrape_native_histograms: true in the scrape config.
  • Stop relying on percentiles for alerting. Count the share of requests slower than the threshold instead. If the threshold is a bucket boundary, this is exact:
1 - (
  sum(rate(http_request_duration_seconds_bucket{job="my-app", le="0.5"}[5m]))
  /
  sum(rate(http_request_duration_seconds_count{job="my-app"}[5m]))
)

The share of requests slower than 500 ms is easier to act on than a p99 value, and it maps directly to a latency SLO.

The server does not see all of the latency

A histogram in the application starts the clock when the handler runs. It misses:

  • time in the load balancer, the ingress controller and the kernel accept queue,
  • TLS handshakes and connection setup,
  • time waiting for a free worker or thread, if the middleware sits after that queue,
  • network time to the client,
  • requests the client gave up on. If the client times out at 2 seconds and the server finishes at 6, the user saw an error and the server recorded a successful slow request, or nothing at all if the work was cancelled.

Measure at more than one point. Ingress or load balancer metrics show queueing in front of the app. Synthetic checks and real user monitoring show what the client sees. When those disagree with the application metrics, the gap is your missing latency.

Retries hide failures and add time

A call that times out and succeeds on retry looks fine in the error rate. The caller waited for the timeout plus the second attempt. If the client retries on its own, the backend sees two requests and has no idea they belong together.

Record latency at the caller, including retries, and count retries as their own metric. A rising retry rate is often the earliest sign that a dependency is getting slow.

Fan-out makes the tail common

If one page load calls N backends and waits for all of them, it is as slow as the slowest call. The chance that at least one of N independent calls is slower than its p99 is 1 - 0.99^N. For 10 calls that is about 10%. For 100 calls it is about 63%. Dean and Barroso described this in "The Tail at Scale" (2013).

The same applies to users. Someone who makes a hundred requests in a session will hit the backend p99 regularly. A rare tail per request is not rare per user.

Windows and resolution smooth the spikes

rate(...[5m]) averages over five minutes, and a dashboard with a wide step adds more averaging. A 20-second stall every few minutes can disappear entirely. When investigating, use shorter ranges (Grafana's $__rate_interval is a sane lower bound) and look at the histogram as a heatmap panel. A heatmap shows a bimodal distribution, such as cache hits and misses, that a single percentile line hides.

Trace the slow requests

Percentiles tell you that something is slow. Traces tell you where. With a low head sampling rate the slow requests are exactly the ones you rarely capture. Tail sampling in the OpenTelemetry Collector decides after the trace is complete, so you can keep every slow or failed trace and a small share of the rest:

processors:
  tail_sampling:
    decision_wait: 10s
    policies:
      - name: slow-requests
        type: latency
        latency:
          threshold_ms: 500
      - name: errors
        type: status_code
        status_code:
          status_codes: [ERROR]
      - name: baseline
        type: probabilistic
        probabilistic:
          sampling_percentage: 5

Add the processor to the traces pipeline. Tail sampling needs all spans of a trace on the same Collector instance, so with more than one replica put a tier with the loadbalancing exporter, routing by trace ID, in front of it.

Exemplars connect the two views: a histogram bucket in Grafana links to an example trace ID that landed in it.

Checklist

  • Never average percentiles. Aggregate histogram buckets, then compute the quantile.
  • Check that the top finite bucket is above your longest timeout.
  • Consider native histograms for better resolution.
  • Alert on the share of requests over a threshold, not on a percentile.
  • Measure at the edge and at the client, not only in the app.
  • Record caller-side latency including retries, and count retries.
  • Use heatmaps and short ranges when investigating.
  • Tail-sample slow and failed traces, and enable exemplars.