Debugging Latency over Hybrid Cloud VPN Connections

Debugging Latency over Hybrid Cloud VPN Connections

Reading time1 min
#devops#hybrid-cloud#vpn#latency#debugging#case-studies

Debugging Latency over Hybrid Cloud VPN Connections

When a service talks across a site-to-site VPN between an office or data center and a cloud region, "it's slow" can mean many things: a long path, a hairpin through a distant hub, fragmented packets, a saturated tunnel, or an application that makes fifty round trips per request. This post goes through them in the order that is cheapest to check.

Know the floor

Light in fiber covers roughly 200 km per millisecond, so every 100 km of fiber adds about 1 ms of round-trip time. Real cables don't follow straight lines, so the actual floor is higher. A transatlantic round trip will never be a few milliseconds, no matter what you configure.

Measure the RTT between the two sites over the internet, outside the tunnel, and compare it with the RTT through the tunnel. If the tunnel is much slower than the internet path between the same places, traffic is taking a detour.

Step 1: Split the request into phases

Find out whether time goes into the network or into the application:

curl -o /dev/null -s -w \
  'dns      %{time_namelookup}\nconnect  %{time_connect}\ntls      %{time_appconnect}\nttfb     %{time_starttransfer}\ntotal    %{time_total}\n' \
  https://api.internal.example.com/health
  • High dns: a resolver far away, or a forwarding chain through another region.
  • High connect: TCP handshake, roughly one RTT. This is the network path.
  • tls minus connect: TLS handshake, more round trips.
  • High ttfb with low connect: the server is slow, or it calls something else across the VPN.

The last case is common. An application server in the cloud that queries an on-premises database ten times per request pays ten RTTs. No network tuning fixes that. Move the service next to its data, batch the queries, or cache.

Step 2: Check the path

ip route get 10.20.0.10           # which interface and gateway the host uses
mtr -rwzc 100 10.20.0.10          # per-hop latency and loss, 100 probes
mtr -rwzc 100 -T -P 443 10.20.0.10  # same over TCP, if ICMP is filtered

Look for a hairpin: traffic from a European office to a European cloud region that goes through a hub in another continent because all tunnels terminate there. Then check the cloud side: VPC route tables or Transit Gateway route tables on AWS, gcloud compute routes list and Cloud Router learned routes on GCP, effective routes on Azure.

Also check the return path. With several tunnels, the request can go through one and the reply through another (asymmetric routing). Stateful firewalls drop such replies, and the result looks like random timeouts. With BGP, influence path selection on both sides (AS path prepending, MED, local preference) so traffic prefers the nearest tunnel in both directions.

Step 3: MTU and MSS

IPsec adds headers, so the usable packet size inside the tunnel is smaller than 1500 bytes. If path MTU discovery is broken (ICMP blocked somewhere), large packets are dropped silently. Typical symptoms: ping works, small API calls work, large responses or TLS handshakes with big certificate chains hang or are very slow.

Find the largest packet that passes without fragmentation:

# Linux: 1372 bytes of payload + 28 bytes of headers = 1400-byte packet
ping -M do -s 1372 -c 3 10.20.0.10

# macOS
ping -D -s 1372 -c 3 10.20.0.10

Lower the size until it works, then clamp TCP MSS on the on-premises VPN device or router so TCP never sends packets that are too large:

iptables -t mangle -A FORWARD -p tcp --tcp-flags SYN,RST SYN -j TCPMSS --clamp-mss-to-pmtu

Each cloud provider documents the MTU and MSS values its VPN supports. Use those as the upper limit.

Step 4: Tunnel capacity and loss

Each VPN tunnel has a documented bandwidth limit, and a single TCP flow over a long RTT is further limited by window size. Under load, a full tunnel shows as queueing delay and then loss, and TCP throughput drops sharply with even small loss on long paths.

  • Watch tunnel metrics. On AWS, TunnelState, TunnelDataIn and TunnelDataOut in the AWS/VPN namespace. On GCP, the Cloud VPN tunnel metrics in Cloud Monitoring.
  • Test throughput directly with iperf3 -s on one side and iperf3 -c 10.20.0.10 -t 30 -P 4 on the other.
  • To spread traffic across several tunnels, you need ECMP. On AWS that means terminating VPNs on a Transit Gateway with BGP. A virtual private gateway doesn't do ECMP across VPN connections.

Step 5: DNS

Check which resolver each client uses and where the answer points. On-premises resolvers forwarding to a cloud resolver endpoint in a distant region add latency to every lookup. Split-horizon setups sometimes return the address of a service in another region. Use resolver endpoints in each region you connect to, and test from each site with dig.

Step 6: Line up the time zones

When sites are spread across time zones, each has its own peak hours. A tunnel that is fine at 10:00 in one office can be saturated when two offices overlap. Keep every log, dashboard and alert in UTC. Plot tunnel throughput against latency over a whole week, and mark the business hours of each site. Latency that tracks one site's working day points at capacity, not configuration.

Fixes that tend to work

  • Terminate each site's VPN in the region closest to it instead of one central hub.
  • Use BGP rather than static routes, so failover and path selection are automatic.
  • On AWS, accelerated Site-to-Site VPN sends traffic over the AWS global network from the nearest edge location. It requires a Transit Gateway attachment.
  • For steady high-volume traffic, move to Direct Connect, Cloud Interconnect or ExpressRoute.
  • Reduce round trips in the application. This is often the biggest win.
resource "aws_customer_gateway" "office_eu" {
  bgp_asn    = 65010
  ip_address = "203.0.113.10"
  type       = "ipsec.1"
  tags       = { Name = "office-eu" }
}

resource "aws_vpn_connection" "office_eu" {
  transit_gateway_id  = aws_ec2_transit_gateway.main.id
  customer_gateway_id = aws_customer_gateway.office_eu.id
  type                = "ipsec.1"
  enable_acceleration = true
  tags                = { Name = "office-eu" }
}

static_routes_only defaults to false, so this connection uses BGP.

Checklist

  • Compare tunnel RTT with internet RTT between the same sites.
  • Break requests into DNS, connect, TLS and server time with curl -w.
  • Trace both directions, look for hairpins and asymmetric routes.
  • Test path MTU with DF pings and clamp MSS.
  • Check tunnel throughput, loss and ECMP.
  • Use regional resolvers and regional VPN termination.
  • Keep all telemetry in UTC and compare latency with each site's working hours.