Estimating the Cost of a Self-Hosted Observability Stack

Estimating the Cost of a Self-Hosted Observability Stack

Reading time1 min
#devops#observability#cloud#kubernetes#terraform

Estimating the Cost of a Self-Hosted Observability Stack

Claims like "a few hundred dollars of infrastructure replaced a five-figure SaaS bill" can be true or wildly wrong. It depends on your telemetry volume and on whether engineering time is counted. You can estimate both sides from your own numbers before you migrate anything.

Understand what the SaaS bill is based on

Observability vendors price on some mix of hosts or containers, GB of logs ingested and indexed, spans ingested and indexed, custom metric series, retention and seats. Open the usage page and find the one or two dimensions that make up most of the bill.

Then check whether you can shrink that dimension without moving:

  • drop debug logs and health check access logs at the agent,
  • sample traces instead of keeping all of them,
  • remove custom metrics and high-cardinality tags nobody queries,
  • shorten retention on data that is only used for debugging.

Reducing volume is often the cheapest option, and you need to do it anyway before self-hosting.

Measure your telemetry volume

Metrics. Active series and ingested samples per second:

prometheus_tsdb_head_series
rate(prometheus_tsdb_head_samples_appended_total[1h])

If you have no Prometheus yet, run a temporary one against your targets for a day, or use the custom metric count from the vendor.

Logs. GB per day. The vendor's usage page is the most reliable number. Log agents also expose bytes sent per output, which tells you which sources produce the most.

Traces. Spans per second and average span size. The OpenTelemetry Collector's internal metrics count accepted spans per receiver.

Take these numbers at a peak period, not an average day.

Turn volume into infrastructure

Metrics storage. The Prometheus documentation gives the formula: disk space equals retention in seconds, times samples per second, times bytes per sample, and Prometheus averages 1 to 2 bytes per sample after compression. For 30 days at the upper bound, in GB:

rate(prometheus_tsdb_head_samples_appended_total[1d]) * 2 * 30 * 86400 / 1e9

Metrics memory. Grows with active series and churn. Do not use a rule of thumb. Run Prometheus with your real targets, divide process_resident_memory_bytes by prometheus_tsdb_head_series, and extrapolate with headroom.

Logs in Loki. Loki stores compressed chunks in object storage and indexes only labels. Storage is daily volume times compression ratio times retention. Measure the ratio on a sample of your own logs, since it depends on the content. Compute depends mostly on queries: searching text over long time ranges means reading many chunks.

Traces in Tempo. Also object storage: spans per second after sampling, times average span size, times retention.

Then price it with your provider's calculator:

  • compute and memory for Prometheus, Loki, Tempo, Grafana and the agents,
  • block storage for Prometheus,
  • object storage per GB-month, plus request charges, which add up when many small objects are written,
  • cross-zone traffic. Replicated ingesters send every write to other zones, and AWS and GCP bill traffic between zones.

The costs that are easy to forget

  • Engineering time. Initial setup, upgrades across several components, capacity planning, and on-call for the monitoring stack itself. Estimate hours per month and multiply by a loaded hourly cost. For a small team this is often larger than the infrastructure.
  • High availability. Two Prometheus replicas, or replication in Loki and Mimir, multiply some of the costs above.
  • Monitoring the monitoring. If Prometheus is down, nothing alerts. kube-prometheus-stack ships an always-firing Watchdog alert; route it to an external dead man's switch service.
  • Config as code. Dashboards and alert rules belong in git, not only in the Grafana database.
  • Security. SSO in front of Grafana, TLS, network policies, tenant separation if several teams share it.
  • Features. Check what you would lose: service maps, profiling, real user monitoring, anomaly detection. Some have open-source counterparts (Pyroscope for profiling, Grafana Faro for frontend monitoring, Tempo's metrics generator for service graphs), each one more component to run.

A minimal self-hosted stack

  • kube-prometheus-stack Helm chart: Prometheus Operator, Prometheus, Alertmanager, Grafana, node-exporter and kube-state-metrics.
  • Loki in single binary mode with object storage, which is enough for small volumes.
  • Tempo with object storage for traces.
  • Grafana Alloy or the OpenTelemetry Collector as the agent. Promtail reached end of life in March 2026, so do not build new setups on it.
  • Mimir or Thanos only when you need months of metrics retention or one view across clusters.

Run Prometheus with persistent storage and explicit retention. A plain Deployment without a volume loses all data on restart.

prometheus:
  prometheusSpec:
    retention: 15d
    retentionSize: 45GB
    storageSpec:
      volumeClaimTemplate:
        spec:
          accessModes: ["ReadWriteOnce"]
          resources:
            requests:
              storage: 50Gi
helm repo add prometheus-community https://prometheus-community.github.io/helm-charts
helm upgrade --install monitoring prometheus-community/kube-prometheus-stack \
  --namespace monitoring --create-namespace -f monitoring-values.yaml

retentionSize should stay below the volume size so Prometheus deletes old blocks before the disk fills up.

Compare like with like

Put both options in one table you build from your own data: the SaaS bill after volume reduction on one side, and infrastructure plus engineering time plus the value of lost features on the other. Redo it when volume grows, since both sides scale differently. A hybrid is common: self-host high-volume metrics and logs, and keep a vendor for the features that are expensive to rebuild.

Checklist

  • Find the billing dimensions that dominate the current bill.
  • Cut volume first: debug logs, unused metrics, unsampled traces.
  • Measure series, samples per second, log GB per day and spans per second at peak.
  • Size metrics storage with the documented 1 to 2 bytes per sample, and measure memory on a real instance.
  • Include object storage requests and cross-zone traffic.
  • Count engineering time, HA and lost features.
  • Use persistent storage, explicit retention and an external dead man's switch.