High-Cardinality Metrics in Prometheus: Finding and Controlling Series Growth

High-Cardinality Metrics in Prometheus: Finding and Controlling Series Growth

Reading time1 min
#devops#prometheus#monitoring#metrics#kubernetes

High-Cardinality Metrics in Prometheus: Finding and Controlling Series Growth

In Prometheus every unique combination of metric name and label values is a separate time series. Cardinality is the number of those series. It is the main driver of Prometheus memory use and query cost, and it grows by accident: one label with an unbounded set of values can turn a single metric into hundreds of thousands of series.

This post covers where series come from, how to find the ones that matter, and what limits to put in place.

Where series come from

The series count of a metric is roughly the product of the number of distinct values of each label, times the number of targets exposing it. Histograms multiply it again. A classic histogram stores one series per bucket, plus +Inf, _sum and _count. With the Go client's default buckets that is 14 series for every label combination.

The usual sources of trouble:

  • IDs in labels. User ID, session ID, request ID, order ID, trace ID, email. The value set is unbounded and grows with traffic.
  • Raw URL paths. /users/48213 instead of the route template /users/{id}.
  • Free text. Error messages or exception strings as label values.
  • Churn. Labels like pod or container_id change on every deploy. Series that stopped receiving samples still sit in the in-memory head block until it is compacted, so frequent rollouts cost memory even when the number of live series looks stable.
  • Exporters and SDKs that expose everything. Auto-instrumentation and some exporters attach many attributes by default. Each one becomes a label.

Finding the offenders

Start with the total:

prometheus_tsdb_head_series

Then look at which metric names and labels contribute most. The Prometheus UI has this under Status, TSDB Status. The same data is in the API:

curl -s 'http://prometheus.example.com:9090/api/v1/status/tsdb?limit=20' \
  | jq '.data.seriesCountByMetricName, .data.labelValueCountByLabelName'

labelValueCountByLabelName is usually the fastest way to spot an ID that leaked into a label: it shows up with a value count far above everything else.

To see which scrape jobs bring the most series:

sort_desc(sum by (job) (scrape_samples_post_metric_relabeling))

To check how many values one suspect label has on one metric:

count(count by (route) (http_requests_total{job="my-app"}))

A query like topk(20, count by (__name__) ({__name__=~".+"})) also works, but it touches every series in the head block. On a large instance it is slow and memory hungry, so prefer the TSDB status endpoint.

For offline analysis, promtool tsdb analyze /prometheus/data reads a block from disk and reports the highest-cardinality metric names, label names, and the label pairs most involved in churn.

Fix it in the instrumentation

The best fix is not emitting the series in the first place. A metric like this one creates a new series for every package:

var deliveries = promauto.NewCounterVec(prometheus.CounterOpts{
    Name: "package_deliveries_total",
    Help: "Delivered packages.",
}, []string{"package_id", "sender", "destination"})

Keep labels that you would group or filter by in a dashboard, with a value set you can list in advance:

var deliveries = promauto.NewCounterVec(prometheus.CounterOpts{
    Name: "package_deliveries_total",
    Help: "Delivered packages by region and result.",
}, []string{"region", "result"})

deliveries.WithLabelValues(regionOf(pkg.Destination), "delivered").Inc()

The package ID goes into a log line or a span attribute, where per-item detail belongs.

A few more rules that prevent most problems:

  • Use the matched route template from your HTTP framework, never the raw path.
  • Map error types to a short fixed list (timeout, connection_refused, invalid_input), not the message text.
  • To jump from a latency spike to a specific request, attach the trace ID as an exemplar, not a label. Exemplar storage needs --enable-feature=exemplar-storage.
  • Consider native histograms. They store the whole distribution in one series per label combination instead of one series per bucket. They are stable since Prometheus 3.8, and scraping them must be turned on with scrape_native_histograms: true.
  • With OpenTelemetry SDKs, use Views to drop attributes you do not need on a metric before it is exported.

Guardrails in Prometheus

Instrumentation will regress at some point, so put limits on the scrape side too:

scrape_configs:
  - job_name: my-app
    sample_limit: 50000
    label_limit: 30
    label_value_length_limit: 200
    static_configs:
      - targets: ["my-app:8080"]
    metric_relabel_configs:
      - source_labels: [__name__]
        regex: "my_app_debug_.*"
        action: drop
      - regex: "session_id"
        action: labeldrop

Things to know about these settings:

  • If a scrape exceeds sample_limit after relabeling, the whole scrape fails and the target's up becomes 0. Pick a limit with headroom above the current value and alert on prometheus_target_scrapes_exceeded_sample_limit_total so a failed scrape gets noticed.
  • metric_relabel_configs runs after the scrape. The target still generates and sends the data. It protects Prometheus, not the application.
  • labeldrop is only safe when the remaining labels still identify each series. If two series differ only by the dropped label, they collide and samples get rejected as duplicates.

When you really need more series

A single Prometheus scales vertically: memory grows with active series and churn. Measure it on your own instance by comparing process_resident_memory_bytes with prometheus_tsdb_head_series over time. When one instance is not enough, split scrape jobs across several Prometheus servers, or send data to a horizontally scalable backend such as Grafana Mimir, Thanos or VictoriaMetrics.

Recording rules make dashboards faster, but they do not reduce ingestion. The raw series are still stored. If you only ever query the aggregate, drop the raw labels at the source.

If what you need is per-user or per-order analysis, metrics are the wrong tool. Logs, traces or an analytical database handle unbounded dimensions; a time series database does not.

Checklist

  • Watch prometheus_tsdb_head_series and alert on unusual growth.
  • Check the TSDB status page after every new service or exporter.
  • No IDs, raw paths or free text in label values.
  • Route templates, bounded error classes, trace IDs as exemplars.
  • sample_limit and label limits on every scrape job, with an alert when they trigger.
  • Drop unused metrics and labels with metric_relabel_configs.
  • Shard or move to Mimir, Thanos or VictoriaMetrics when one instance runs out of memory, after removing what you do not need.