IPv6 Rollout Problems and How Dual-Stack Helps

IPv6 Rollout Problems and How Dual-Stack Helps

Reading time1 min
#ipv6#devops#dual-stack#networking#cloud

IPv6 Rollout Problems and How Dual-Stack Helps

Dual-stack means every layer speaks both IPv4 and IPv6: hosts, load balancers, firewalls, DNS and the application. It is the safest way to add IPv6, because IPv4 keeps working the whole time. It is also where most rollout problems come from, because now there are two paths, and the IPv6 one is the one nobody has tested.

How clients choose

When a hostname has both A and AAAA records and the client has IPv6 connectivity, the client usually prefers IPv6 (RFC 6724 default address selection). Modern browsers and many libraries implement Happy Eyeballs (RFC 8305): they start the IPv6 connection, and if it does not succeed within a short delay (the RFC recommends 250 ms), they race an IPv4 connection and use whichever wins.

So a broken IPv6 path does not break everyone. Clients with Happy Eyeballs see a small delay. Clients without it (older runtimes, some HTTP libraries, scripts, IoT devices) wait for the full connect timeout and may fail. That is why IPv6 problems often show up as "slow for some users" rather than "down".

Common failure points

AAAA published before the IPv6 path works. The record goes live, but the load balancer listener, firewall rule or backend is not ready. Publish AAAA last, after the path is tested end to end.

Firewall rules for one family only. Two opposite failures here:

  • IPv6 is blocked because nobody added rules for it. Cloud security groups treat 0.0.0.0/0 and ::/0 as separate rules, and iptables does not cover IPv6 (ip6tables does).
  • IPv6 is wide open because filtering was only ever configured for IPv4. A host with a global IPv6 address and no IPv6 rules exposes every listening port.

With nftables, an inet family table applies the same rules to both protocols, which avoids drift between two rule sets.

ICMPv6 filtered. IPv6 routers do not fragment packets. Path MTU discovery depends on ICMPv6 "Packet Too Big" messages reaching the sender, and neighbour discovery runs over ICMPv6 too. If a firewall drops ICMPv6 as "ping", the classic symptom is a TCP handshake that succeeds and a large response that hangs. RFC 4890 lists which ICMPv6 types to allow. The IPv6 minimum link MTU is 1280 bytes, so tunnels and VPNs on the path matter.

Services listening on IPv4 only. nginx listen 80; binds IPv4 only, you need listen [::]:80; as well. Applications bound to 0.0.0.0 are not reachable over IPv6. On Linux, a socket bound to :: also accepts IPv4 by default (net.ipv6.bindv6only = 0), which is usually what you want.

Code and data that assume IPv4. Database columns sized for 255.255.255.255, regexes in log parsers, IP allowlists, geo-IP lookups, URLs built as host:port without brackets ([2001:db8::10]:8080). Rate limiting per address does not work for IPv6, because one client usually controls a whole /64. Rate-limit per /64 (or larger) instead.

Address assignment. SLAAC is the one mechanism every client supports; stateful DHCPv6 is not universal. If you run an IPv4-only network, hosts still have IPv6 enabled and will accept router advertisements, so enable RA Guard on access switches either way. ISC DHCP, used in many older DHCPv6 examples, reached end of life in 2022; Kea is its replacement.

Applications that cannot do IPv6

Do not patch around them with wrappers that rewrite addresses. Terminate IPv6 in front of them: a dual-stack load balancer or reverse proxy accepts both families and connects to the backend over IPv4. The application never sees IPv6, apart from the client address, which the proxy passes in X-Forwarded-For or the PROXY protocol. Make sure the application parses those IPv6 addresses correctly.

The reverse case, IPv6-only clients that need IPv4-only destinations, is handled by NAT64 with DNS64, using the well-known prefix 64:ff9b::/96 (RFC 6052).

Rollout order

  1. Inventory. Load balancers, CDNs, firewalls, WAF rules, logging pipeline, rate limiters, allowlists, every component that stores or compares IP addresses.
  2. Enable IPv6 on the network and load balancer, without AAAA. Nothing changes for users yet.
  3. Test the IPv6 path directly, by address or a test hostname that has only an AAAA record:
dig +short AAAA test-v6.example.com
curl -6 -sS -o /dev/null -w '%{http_code} connect=%{time_connect}s total=%{time_total}s\n' https://test-v6.example.com/
curl -4 -sS -o /dev/null -w '%{http_code} connect=%{time_connect}s total=%{time_total}s\n' https://www.example.com/
tracepath -6 test-v6.example.com

Request something large, not only a health check, to catch MTU problems. On the servers, ss -ltn shows whether services listen on [::] or only 0.0.0.0.

  1. Fix logging and metrics first. Split request and error rates by address family, so you can compare them after the switch.
  2. Lower the TTL on the records you will change, and wait for the old TTL to pass.
  3. Publish AAAA on a less critical hostname first, then the main ones.
  4. Watch the IPv4 and IPv6 error rates side by side. If IPv6 is worse, the rollback is removing the AAAA record. With a low TTL that takes effect quickly; resolvers that cached the old answer keep it until it expires.

Kubernetes

Dual-stack networking is GA in Kubernetes since 1.23. The cluster, CNI plugin and nodes must be configured for it. A Service then asks for both families:

apiVersion: v1
kind: Service
metadata:
  name: my-app
spec:
  ipFamilyPolicy: PreferDualStack
  ipFamilies: [IPv4, IPv6]
  selector:
    app: my-app
  ports:
    - port: 80
      targetPort: 8080

Check that NetworkPolicies and any cloud firewall rules cover the IPv6 pod and service ranges, not only the IPv4 ones.

Checklist

  • AAAA records published last, after end-to-end tests.
  • Firewall rules for both families, ideally from one rule set.
  • ICMPv6 allowed per RFC 4890, including Packet Too Big.
  • Services listening on [::] as well as IPv4.
  • IP-handling code, storage, allowlists and rate limits reviewed for IPv6.
  • RA Guard on access networks, even IPv4-only ones.
  • Metrics split by address family, low TTL during the switch, AAAA removal as the rollback.