Switching APM Vendors Without Losing Historical Context
Teams change APM vendors for cost, features or consolidation. The usual worry is losing years of history. In practice most of that history cannot be moved, and most of it does not need to be. What you need is a plan for the few things you actually use, plus an overlap period where both tools run side by side.
What can and cannot be moved
- Raw traces are sampled and kept for a short time, often days or weeks depending on the plan. There is no useful way to import them into another vendor: IDs, formats and sampling decisions do not carry over.
- Metrics are aggregated and kept longer. You can export them through the vendor's query API at hourly or daily resolution.
- Backfilling into the new vendor is often limited. Many ingestion APIs reject or restrict samples with old timestamps, and backfilled data may be billed like new data. Check the target's ingestion limits before planning around it.
- Dashboards, alerts, SLO definitions and deploy markers are configuration, not data. They are often the most valuable part, and they have to be rebuilt in any case.
A realistic goal: export the aggregates you use, keep the old tool readable for its remaining retention, and rebuild dashboards and alerts on the new one.
Decide which history you use
Write down the questions history answers for you. Typical ones:
- What did traffic look like during last year's seasonal peak?
- How have latency and error rate per service trended, for capacity planning?
- What were the SLO results for past periods, for reports or contracts?
- What was the performance baseline before a large release?
Each needs a handful of series per service at hourly or daily resolution: request rate, error rate, a couple of latency percentiles. That is a small export.
Export aggregates to a store you control
For example, with New Relic's NerdGraph API and a user API key:
QUERY="SELECT count(*), percentile(duration, 50, 95, 99) FROM Transaction WHERE appName = 'my-app' SINCE 90 days ago TIMESERIES 1 day"
jq -n --arg q "$QUERY" \
'{query: "{ actor { account(id: 1234567) { nrql(query: \($q | tojson)) { results } } } }"}' \
| curl -s https://api.newrelic.com/graphql \
-H 'Content-Type: application/json' \
-H "API-Key: ${NEW_RELIC_USER_KEY}" \
-d @- > my-app-daily-90d.json
Other vendors have equivalent query APIs. Loop over services and time ranges, respect the API rate limits, and save the raw responses before transforming anything.
Where to keep the result:
- Files in object storage (CSV or Parquet) with bucket versioning. Cheap and readable by anything.
- A database you own, such as ClickHouse or your data warehouse, if people will query it often.
- Prometheus, if you want it next to current metrics in Grafana.
promtool tsdb create-blocks-from openmetrics history.om ./blocksbuilds TSDB blocks from an OpenMetrics file, which you then copy into the data directory. Blocks older than the configured retention are deleted, so check retention first.
For each exported series, record the query, the unit and how the vendor computed it. A p95 from one vendor is not directly comparable with a p95 from another. They may use different sampling, different histogram algorithms, and different start and end points for "duration".
Run both tools in parallel
Send the same telemetry to both vendors for an overlap period long enough to cover your normal business cycle, such as month-end processing or weekly peaks. Most APM vendors accept OTLP, directly or through their agent, so the OpenTelemetry Collector can fan out to both:
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
http:
endpoint: 0.0.0.0:4318
processors:
memory_limiter:
check_interval: 1s
limit_percentage: 80
spike_limit_percentage: 20
batch: {}
exporters:
otlphttp/old:
endpoint: https://otlp.old-vendor.example.com
headers:
api-key: ${env:OLD_VENDOR_API_KEY}
otlphttp/new:
endpoint: https://otlp.new-vendor.example.com
headers:
api-key: ${env:NEW_VENDOR_API_KEY}
service:
pipelines:
traces:
receivers: [otlp]
processors: [memory_limiter, batch]
exporters: [otlphttp/old, otlphttp/new]
metrics:
receivers: [otlp]
processors: [memory_limiter, batch]
exporters: [otlphttp/old, otlphttp/new]
The header names and endpoints differ per vendor. Check each vendor's OTLP documentation.
During the overlap, compare the same numbers in both tools: request rate, error rate and latency per service. They will not match exactly. Find out why before trusting either one. Common reasons are a different definition of an error (whether 4xx responses count), different sampling, latency measured at a different point, and units in milliseconds in one tool and seconds in the other.
Make the next switch cheaper
If the old setup used a proprietary agent, this is the moment to move instrumentation to OpenTelemetry SDKs and auto-instrumentation. The vendor then becomes an exporter in the Collector config. Use the OpenTelemetry semantic conventions for names and attributes, for example the http.server.request.duration histogram in seconds, so dashboards and alerts depend less on one backend.
Rebuild what is used, delete the rest
- Inventory dashboards and alerts in the old tool. If it shows view counts or last-viewed dates, use them.
- Migrate the dashboards people open and the alerts that page. Most of the rest can go.
- Run the new alerts in parallel with the old ones and compare what fires before switching the pager over.
- Point deploy markers from CI at the new tool on day one, so the new history has release context.
Keep read access to the old tool
Check the contract end date and the retention of the old data. If possible, keep a minimal plan with read access until your exports are verified and the overlap period is over. Only then turn the old agents off.
Checklist
- List the questions history must answer, and export only the series they need.
- Save raw API responses, then transform. Record query, unit and method per series.
- Store exports in object storage, a database you own, or Prometheus.
- Fan out with the OpenTelemetry Collector for a full business cycle.
- Explain the differences between the two tools before trusting the new numbers.
- Move instrumentation to OpenTelemetry so the next migration is a config change.
- Rebuild used dashboards and alerts only. Keep old read access until exports are verified.
