Benchmarking ARM and x86 Cloud Instances with Real Workloads

Benchmarking ARM and x86 Cloud Instances with Real Workloads

Reading time1 min
#cloud#performance#architecture#benchmarking#devops

Benchmarking ARM and x86 Cloud Instances with Real Workloads

All major clouds offer ARM instances: AWS Graviton, Azure Cobalt and Ampere-based VMs, Google Cloud Axion and Tau T2A. Their list price per vCPU-hour is usually lower than comparable x86 instances. Whether they are cheaper for your service depends on how many requests one instance handles within your latency target. A generic benchmark will not tell you that. Your own workload will.

What actually differs

  • What a vCPU is. On most x86 instance types a vCPU is one hardware thread, and two of them share a physical core through SMT. Graviton and Ampere Altra processors have no SMT, so each vCPU is a full core. CPU-bound services can behave quite differently at the same vCPU count.
  • Optimized code paths. Libraries for compression, cryptography, image processing, JSON parsing and ML inference often have hand-tuned x86 SIMD code (AVX2, AVX-512). Their ARM paths (NEON, SVE) may be newer, slower or missing.
  • Runtime versions. Recent Go, Java, .NET, Node.js and Python releases have mature arm64 support. Older ones may not. Upgrade before benchmarking, or you measure the old runtime.
  • Native dependencies. Python wheels and Node.js native modules may be missing for arm64 and get compiled from source, or fail.
  • Generations. Memory bandwidth, cache sizes and clock speeds differ between processor generations on both sides. A result for one generation does not carry over to the next.

Build it properly

Build a multi-arch image so both platforms run native code:

docker buildx build --platform linux/amd64,linux/arm64 \
  -t registry.example.com/my-app:1.4.0 --push .
docker buildx imagetools inspect registry.example.com/my-app:1.4.0

The second command lists the platforms in the manifest. Never benchmark an image that runs under QEMU emulation. It works, but it is much slower and says nothing about the hardware. Run uname -m inside the container to confirm it reports aarch64 on ARM nodes.

Design the benchmark

  • Change one thing. Same OS image, kernel, runtime version, application build and configuration. Compare the same instance size class from current generations.
  • Compare two ways. At the same vCPU count, and at roughly the same hourly price. Report both.
  • Use real traffic. A realistic load test with the production request mix, real data sizes and real dependencies. A single endpoint returning a constant says nothing about database drivers, TLS, serialization or garbage collection.
  • Fix the rate, measure latency. Use an open model (k6 arrival-rate executors) and step the request rate up. For each instance type, find the highest rate where p99 latency and error rate stay within your SLO. That is its capacity. Maximum throughput from a closed loop is not.
  • Warm up. Discard the first minutes while JIT compilers and caches settle.
  • Repeat. Run on several instances at different times. Cloud instances vary, and a single run can mislead. Report the spread, not one number.
  • Check the load generator. It must not be the bottleneck, and it must not be a burstable instance that runs out of CPU credits halfway through.

Compare with production traffic

The most reliable benchmark is real traffic. On Kubernetes, add an arm64 node pool and run a second Deployment pinned to it, behind the same Service:

spec:
  template:
    metadata:
      labels:
        app: my-app
        arch: arm64
    spec:
      nodeSelector:
        kubernetes.io/arch: arm64

The Service selects app: my-app, and each Deployment has its own arch label in its selector, so the selectors do not overlap. Deployment selectors are immutable, so if the existing one selects only app: my-app, create a new amd64 Deployment with the extra label and retire the old one. Copy the arch pod label onto the scraped metrics, for example with podTargetLabels: [arch] in a ServiceMonitor, and compare the two groups:

histogram_quantile(0.99,
  sum by (arch, le) (rate(http_request_duration_seconds_bucket{job="my-app"}[10m]))
)
sum by (arch) (rate(process_cpu_seconds_total{job="my-app"}[10m]))
/
sum by (arch) (rate(http_requests_total{job="my-app"}[10m]))

The second query gives CPU seconds per request, which stays comparable even when the two groups receive different shares of traffic. Run it long enough to cover daily peaks.

Turn results into cost per request

For each instance type:

cost per million requests = hourly price / (requests per second at SLO * 3600) * 1,000,000

Use the prices you actually pay after discounts and commitments, not list prices. Check that your commitments (Savings Plans, committed use discounts, reservations) apply to the instance family you are moving to.

Add the one-time and running costs of supporting two architectures: multi-arch builds in CI (QEMU builds are slow, native ARM runners cost money), testing on both, and time spent on architecture-specific bugs. For a small service, these can outweigh the savings.

For AWS Lambda, compare billed duration and price per GB-second for each architecture. Lambda scaling is governed by concurrency limits, not by CPU architecture.

Common mistakes

  • Benchmarking an emulated image or a debug build.
  • Comparing an old x86 generation with a new ARM one, or the other way around.
  • Old runtime versions without arm64 optimizations.
  • Synthetic single-endpoint tests instead of the real request mix.
  • One run, one instance, one number.
  • Ignoring the cost of building and supporting two architectures.

Checklist

  • Multi-arch image, verified native on both platforms.
  • Current runtime and library versions.
  • Same everything except the architecture.
  • Realistic load, open model, stepped rates, capacity measured at the SLO.
  • Several runs on several instances.
  • Production canary on an arm64 node pool, compared by latency and CPU per request.
  • Cost per million requests at real prices, plus the cost of supporting both architectures.