Load Testing with Realistic Traffic in k6

Load Testing with Realistic Traffic in k6

Reading time1 min
#load testing#devops#performance#infrastructure

Load Testing with Realistic Traffic in k6

A test that hammers one endpoint at maximum rate measures "hello world" RPS: how fast the service answers a cheap, cached request over a few warm connections. The number is real, but it says little about how the system behaves when actual users arrive.

Why simple tests overstate capacity

  • Only cheap requests. Production traffic includes logins, writes, searches, uploads and the occasional heavy report. Those set the limit, not the health check.
  • Same IDs every time. The same product stays in every cache. Real traffic has a long tail of keys and a lower hit ratio.
  • Few long-lived connections. Kubernetes Services balance per connection, not per request. A load generator with a handful of keep-alive connections can pin all its traffic to a few pods.
  • Closed model. A fixed number of virtual users that each wait for a response before sending the next request. When the system slows down, the generator slows down with it and reports lower throughput instead of growing queues. Real users keep arriving whether you are slow or not. This is the coordinated omission problem.
  • Small test data. A staging database with a fraction of production rows makes every query look fast.
  • No background work. Cron jobs, queue consumers, cache expiry and deploys all run in production at the same time as user traffic.

Get the traffic mix from production

Start with the endpoint mix. From an access log in the combined format:

awk '{print $6, $7}' access.log \
  | sed -E 's/"//; s/\?.*//; s#/[0-9]+#/{id}#g' \
  | sort | uniq -c | sort -rn | head -20

Or from metrics, at your peak hour:

sort_desc(sum by (method, route) (rate(http_requests_total{job="my-app"}[1h])))

Also collect:

  • the typical sequences of calls (traces or session logs show what a user does after login),
  • how IDs are distributed: which share of requests hit popular items,
  • payload sizes for uploads and writes,
  • the peak rate and how fast traffic ramps up to it.

Model flows, not endpoints

In k6, each user journey becomes a scenario. Arrival-rate executors give you an open model: k6 starts new iterations at the configured rate no matter how slow the responses are.

import http from 'k6/http';
import { check, sleep } from 'k6';
import { SharedArray } from 'k6/data';

const BASE = __ENV.BASE_URL || 'https://staging.example.com';
const products = new SharedArray('products', () => JSON.parse(open('./products.json')));
const accounts = new SharedArray('accounts', () => JSON.parse(open('./accounts.json')));

export const options = {
  scenarios: {
    browse: {
      executor: 'ramping-arrival-rate',
      exec: 'browse',
      startRate: 5,
      timeUnit: '1s',
      preAllocatedVUs: 100,
      maxVUs: 1000,
      stages: [
        { target: 50, duration: '10m' },
        { target: 50, duration: '30m' },
        { target: 150, duration: '2m' },
        { target: 50, duration: '5m' },
      ],
    },
    checkout: {
      executor: 'constant-arrival-rate',
      exec: 'checkout',
      rate: 2,
      timeUnit: '1s',
      duration: '47m',
      preAllocatedVUs: 50,
      maxVUs: 300,
    },
  },
  thresholds: {
    http_req_failed: ['rate<0.01'],
    'http_req_duration{scenario:browse}': ['p(95)<400'],
    'http_req_duration{scenario:checkout}': ['p(95)<1000'],
    dropped_iterations: ['count<1'],
  },
};

const pick = (list) => list[Math.floor(Math.random() * list.length)];
const json = { 'Content-Type': 'application/json' };

export function browse() {
  http.get(`${BASE}/`, { tags: { name: 'home' } });
  sleep(1 + Math.random() * 3);

  const search = http.get(`${BASE}/api/search?q=${encodeURIComponent(pick(products).term)}`, {
    tags: { name: 'search' },
  });
  check(search, { 'search ok': (r) => r.status === 200 });
  sleep(1 + Math.random() * 3);

  http.get(`${BASE}/api/products/${pick(products).id}`, { tags: { name: 'product' } });
}

export function checkout() {
  const login = http.post(`${BASE}/api/login`, JSON.stringify(pick(accounts)), { headers: json });
  if (!check(login, { 'login ok': (r) => r.status === 200 })) return;
  const headers = { ...json, Authorization: `Bearer ${login.json('token')}` };

  http.post(`${BASE}/api/cart/items`, JSON.stringify({ productId: pick(products).id, qty: 1 }), {
    headers,
    tags: { name: 'add-to-cart' },
  });
  sleep(2 + Math.random() * 5);

  const order = http.post(`${BASE}/api/orders`, JSON.stringify({ payment: 'test-card' }), {
    headers,
    tags: { name: 'create-order' },
  });
  check(order, { 'order created': (r) => r.status === 201 });
}

Run it with k6 run -e BASE_URL=https://staging.example.com load.js.

Notes on the script:

  • Rates, stages and thresholds are placeholders. Take the rates from your production peak and the thresholds from your SLOs.
  • sleep adds think time. With arrival-rate executors it makes iterations longer, so k6 needs more VUs. Set maxVUs high enough.
  • dropped_iterations counts iterations k6 could not start because it ran out of VUs. If that threshold fails, the generator did not deliver the planned load and the results are not valid.
  • The name tag groups URLs with IDs into one metric. Without it, every product URL becomes its own time series in the k6 results.

Data and environment

  • Use enough distinct products and accounts that the cache hit ratio is close to production. Pre-create test accounts.
  • Size the test database like production, or at least the tables on the hot path.
  • Stub external payment and email providers, but give the stubs a realistic response time.
  • Match production instance types, replica counts, autoscaling settings and connection pool limits.
  • Run long enough for caches, JIT compilers and autoscalers to settle, and run one test right after a deploy to see cold behavior.

Watch the generator and the system

The load generator can be the bottleneck. If CPU or network on the k6 machine is saturated, you are measuring the generator. For larger tests, use several machines or the k6 Operator on Kubernetes.

Correlate k6 results with server-side metrics: latency and errors per route, CPU throttling, database connections, queue depth. k6 can send its metrics to Prometheus through remote write, so both sides fit on one Grafana dashboard.

The useful result is not "maximum RPS before errors". It is the highest arrival rate at which the thresholds still pass, plus which resource saturated first.

Test types worth running

  • Load: expected peak with the real mix.
  • Stress: past the peak, to see how the system fails and recovers.
  • Spike: a sudden jump, to test autoscaling and queueing.
  • Soak: hours at normal load, to find leaks and pool exhaustion.

Checklist

  • Endpoint mix, flows and ID distribution from production.
  • Open model with arrival-rate executors.
  • Think time, varied test data, authenticated sessions.
  • Thresholds from SLOs, plus a check on dropped_iterations.
  • Production-like data size and infrastructure.
  • Generator monitored, server metrics correlated.
  • Report capacity at the SLO and the first saturated resource.