Kubernetes Network Policies: Common Mistakes and a Safe Rollout
NetworkPolicy is the standard way to restrict pod traffic in Kubernetes. It is also easy to get wrong in ways that do not produce an error: the policy is accepted, and some traffic just stops. Kubernetes has no built-in logging for dropped connections, so the first sign is often a timeout in some other service.
This post covers how policies are evaluated, the mistakes that cause most outages, and a rollout order that avoids them.
How evaluation works
- Pods are open by default. A pod with no policy selecting it accepts and sends anything.
- Selection makes a pod isolated. As soon as any policy selects a pod for ingress, only traffic allowed by some policy can reach it. The same goes for egress. Ingress and egress are isolated separately, based on
policyTypes. - Policies are additive. There is no deny rule and no order. The allowed traffic is the union of all policies that select the pod.
- Both ends count. For pod-to-pod traffic, the source must be allowed to send (egress) and the destination must be allowed to receive (ingress).
- The CNI enforces it. If your network plugin does not implement NetworkPolicy, policies are accepted and do nothing. Check this before anything else.
Some exceptions from the Kubernetes documentation that matter in practice:
- Traffic between a pod and the node it runs on is always allowed. Kubelet liveness and readiness probes keep working under a default-deny policy.
- Policies cover TCP, UDP and SCTP. What happens to ICMP is undefined and depends on the plugin, so
pingis not a valid connectivity test. - Whether a policy change affects connections that are already open is implementation-defined. A test right after applying a policy can pass on an old connection and fail later.
Mistakes that cut traffic
Egress default deny without DNS. The most common one. Pods can no longer resolve names, and every outbound call fails with a lookup error. Always pair egress isolation with a DNS rule:
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: allow-dns
namespace: my-app
spec:
podSelector: {}
policyTypes: ["Egress"]
egress:
- to:
- namespaceSelector:
matchLabels:
kubernetes.io/metadata.name: kube-system
podSelector:
matchLabels:
k8s-app: kube-dns
ports:
- protocol: UDP
port: 53
- protocol: TCP
port: 53
Check the labels of your DNS pods first; k8s-app: kube-dns is common but not universal. If your nodes run NodeLocal DNSCache, pods send DNS to a node-local address instead; check how your CNI treats that traffic.
AND vs OR in from and to. These two look almost the same:
# one element: pods labeled app=frontend IN namespaces labeled team=web
ingress:
- from:
- namespaceSelector:
matchLabels:
team: web
podSelector:
matchLabels:
app: frontend
---
# two elements: ANY pod in team=web namespaces, OR app=frontend pods in this namespace
ingress:
- from:
- namespaceSelector:
matchLabels:
team: web
- podSelector:
matchLabels:
app: frontend
The second one is much wider than most people intend. kubectl describe networkpolicy prints how Kubernetes interpreted the rule; read it after every change.
podSelector alone means "this namespace only". A rule with only podSelector never matches pods in other namespaces. Ingress controllers, Prometheus and service meshes usually run in their own namespaces and need explicit rules. Use the kubernetes.io/metadata.name label, which the control plane sets on every namespace, to select a namespace by name.
Forgetting non-application traffic. Before isolating a namespace, list who talks to it besides other services: ingress controller, metrics scraping, log shippers that pull, admission webhooks (called by the API server), backup tools, and egress to databases or APIs outside the cluster.
ipBlock for things behind NAT or load balancers. Source and destination IPs may be rewritten before or after policy processing, depending on the plugin and cloud. An ipBlock for a load balancer or for traffic from outside the cluster may not match what you expect. Test it on your setup.
A rollout order that works
- Map the flows first. If your CNI has flow visibility (Cilium Hubble, Calico flow logs, cloud VPC flow logs), record real traffic for a while and build the allow list from it, not from memory.
- Write the allow rules before the deny. Apply allow policies for every known flow. Since nothing is isolated yet, they change nothing.
- Add default deny for one namespace at a time:
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: default-deny
namespace: my-app
spec:
podSelector: {}
policyTypes: ["Ingress", "Egress"]
- Test from inside the cluster with a throwaway pod using the same labels as the real client:
kubectl run nettest -n my-app --rm -it --restart=Never \
--labels app=frontend --image=busybox:1.36 -- \
sh -c 'nslookup backend.my-app.svc.cluster.local && wget -qO- -T 3 http://backend:8080/healthz'
- Watch for drops. If your CNI reports policy denials, alert on them for the namespace you just changed. Otherwise, watch error rates and timeouts of the services in it.
- Keep policies with the application. Ship each service's NetworkPolicy in its Helm chart or kustomization, so a new dependency and its allow rule are reviewed in the same pull request.
What NetworkPolicy cannot do
It has no explicit deny rules, no logging, no cluster-wide default policy, no selection of Services by name and nothing TLS-related. CNI-specific policies (CiliumNetworkPolicy, Calico GlobalNetworkPolicy) and the cluster-scoped APIs from SIG Network's network-policy-api project cover some of these. Use them knowing that they tie you to a plugin or to an API that is still evolving.
Checklist
- Confirm the CNI enforces NetworkPolicy.
- Every egress-isolated pod has a DNS rule.
- Namespace and pod selectors in the same element when you mean AND.
- Ingress controller, monitoring and webhooks explicitly allowed.
- Allow rules first, default deny last, one namespace at a time.
- Test with labeled throwaway pods, not
ping. - Policies live next to the application they protect.
