Finding Latency Added by Configuration Defaults
When a request is slower than profiling the code can explain, the missing time is often in configuration: name resolution, connection handling, CPU limits, timeouts. Each one adds a little per hop, and a request that crosses five services pays it five times.
Start with a latency budget
Pick one important request and write down the end-to-end target. Split it across the hops: edge, gateway, service, database, external APIs. Then measure what each hop actually takes. The difference between the budget and the measurement tells you where to look.
Do the same with timeouts. They should get shorter down the call chain. A gateway that gives up after 3 seconds gains nothing from a backend that waits 5 seconds and then retries.
Break one request down
curl reports where the time goes for a single request:
curl -o /dev/null -s -w \
'dns=%{time_namelookup} connect=%{time_connect} tls=%{time_appconnect} ttfb=%{time_starttransfer} total=%{time_total}\n' \
https://api.example.com/v1/orders/123
The values are cumulative seconds from the start. A large dns points at the resolver. A large gap between connect and tls is handshake cost, which matters if production clients do not reuse connections. The gap between tls and ttfb is the server. Run it from your laptop and from a pod inside the cluster and compare. For calls between services, a distributed trace gives the same breakdown per hop.
DNS search domains
Pods with the default ClusterFirst DNS policy get ndots:5 and several search domains in /etc/resolv.conf. A name with fewer than five dots, such as api.example.com, is first tried with every search domain appended (api.example.com.my-ns.svc.cluster.local and so on), for both A and AAAA records, before the real name is queried. With a warm cache these NXDOMAIN answers are fast. With a cold cache or a busy CoreDNS each one is a round trip. If a UDP packet is lost, glibc waits for the resolver timeout, 5 seconds by default, before trying again.
Lower ndots for pods that mostly call external names:
spec:
dnsConfig:
options:
- name: ndots
value: "2"
Check that the short service names your app uses still resolve after the change. NodeLocal DNSCache also helps, since it answers from a cache on the node. To measure, watch coredns_dns_request_duration_seconds and the share of NXDOMAIN in coredns_dns_responses_total.
Connection reuse and pool sizes
A new connection costs one round trip for TCP and one more for a TLS 1.3 handshake (two for TLS 1.2) before the request is even sent. Clients that do not reuse connections pay this on every call. Common causes:
- An HTTP client created per request, so the pool is thrown away each time.
- A pool that is too small. Go's
http.Transportkeeps only 2 idle connections per host by default. Under higher concurrency the extra connections are closed after use and reopened on the next burst. - In Go, response bodies that are not fully read and closed, which prevents reuse.
- A client idle timeout longer than the server's or load balancer's. The client picks a connection the server already closed and has to retry.
t := http.DefaultTransport.(*http.Transport).Clone()
t.MaxIdleConnsPerHost = 64
t.IdleConnTimeout = 50 * time.Second // below the upstream's idle timeout
client := &http.Client{Transport: t, Timeout: 2 * time.Second}
Create the client once and share it. A quick check on the client side: thousands of sockets in TIME_WAIT towards the same backend (ss -tan state time-wait | wc -l) mean connections are not being reused.
Nagle and delayed ACK
If a program writes a request in two small writes on a socket with Nagle's algorithm enabled, the second write waits for the ACK of the first. The receiver may delay that ACK, and on Linux the minimum delayed ACK timeout is 40 ms. The symptom is latency that jumps in steps of roughly 40 ms. Most HTTP clients and runtimes set TCP_NODELAY (Go does by default), so look at custom protocols, older libraries and hand-written socket code.
CPU limits and throttling
A container with a CPU limit gets a quota per CFS period, which is 100 ms by default. A multi-threaded process can use the whole quota early in the period and is then paused until the next one starts. Requests in flight wait. Average CPU usage looks low because the throttling happens in short bursts.
sum by (namespace, pod) (rate(container_cpu_cfs_throttled_periods_total{container!=""}[5m]))
/
sum by (namespace, pod) (rate(container_cpu_cfs_periods_total{container!=""}[5m]))
If a latency-sensitive service is throttled in a large share of periods, raise or remove its CPU limit and keep the request. Also make sure the runtime's thread count matches the limit. Go 1.25 and later set GOMAXPROCS from the cgroup CPU limit. Older Go versions need go.uber.org/automaxprocs, and the JVM has -XX:ActiveProcessorCount.
Timeouts and retries that multiply
A long timeout does not slow down successful requests, but it decides how long callers wait when a dependency is slow, and how long threads and connections stay busy. Retries at several layers multiply: three layers that each try three times can send up to 27 requests to the bottom service for one user action, and the user waits for all the timeouts in between.
Retry at one layer, usually the one closest to the user. Use a retry budget or cap, add jitter, and propagate deadlines so downstream services stop working on requests the caller already abandoned. gRPC propagates deadlines for you if you pass the context along.
Smaller items worth checking
- Synchronous logging to a slow destination, or debug level left on in production.
- Sidecar proxies: compare mesh metrics with application metrics to see what each hop adds.
- Per-request work that belongs at startup: reading config, building clients, loading certificates.
- Cross-zone or cross-region calls caused by routing or replica placement.
Checklist
- Write a latency budget per hop and timeouts that shrink down the chain.
- Break requests down with
curl -wand traces. - Tune
ndotsor add NodeLocal DNSCache for pods that call external names. - One shared HTTP client, pool sized for concurrency, idle timeout below the server's.
- Look for 40 ms steps that point to Nagle and delayed ACK.
- Track CFS throttling and set
GOMAXPROCSor the JVM equivalent to the CPU limit. - Retry at one layer, with a budget and deadline propagation.
