Container Escape Paths in Kubernetes and How to Close Them
A container is a normal Linux process with namespaces, cgroups, dropped capabilities, a seccomp filter and usually an AppArmor or SELinux profile around it. It shares the kernel with every other container on the node. "Escape" means one of those layers was weakened by configuration or broken by a bug, and the process ends up with access to the node.
Most exposure in real clusters comes from configuration, not from new exploits. That is good news, because configuration can be checked and enforced.
The main paths
Privileged containers. privileged: true turns off most of the isolation: all capabilities, host devices, no seccomp. Treat a privileged pod as root on the node.
Host mounts. A hostPath volume of /, /etc, /var/lib/kubelet or the container runtime socket (/run/containerd/containerd.sock, /var/run/docker.sock) gives the container control over the node. Access to the runtime socket is equivalent to root, because it can start new privileged containers.
Host namespaces. hostPID, hostIPC and hostNetwork remove the separation from node processes and node networking. With hostNetwork the pod also reaches everything the node can reach, including services bound to localhost and the cloud metadata endpoint.
Extra capabilities. SYS_ADMIN, SYS_PTRACE, SYS_MODULE and NET_ADMIN each widen what the process can do against the kernel. Most applications need none of them.
Root plus privilege escalation. Without user namespaces, UID 0 in the container is UID 0 for the kernel. If allowPrivilegeEscalation is not false, setuid binaries inside the image can raise privileges. Running as non-root with no_new_privs set removes a whole class of problems.
Kernel and runtime bugs. All containers share one kernel, so a kernel privilege escalation is also a container escape. Container runtimes have had escape bugs too. Public examples are runc CVE-2019-5736, CVE-2024-21626 (fixed in runc 1.1.12) and the set CVE-2025-31133, CVE-2025-52565 and CVE-2025-52881 (fixed in runc 1.2.8, 1.3.3 and 1.4.0-rc.3). On the kernel side, Dirty Pipe (CVE-2022-0847) is the well-known one. Configuration does not fix these. Patching does.
Credentials, not escape. A stolen service account token with broad RBAC, or node IAM credentials from the metadata endpoint, give an attacker the same result without touching the kernel.
One note on sidecars: containers in the same pod share a network namespace, so a sidecar can talk to the main container on localhost. NetworkPolicy does not filter traffic inside a pod. The pod is the trust boundary, so vet sidecar images the same way as the application image.
Enforce a baseline with Pod Security Admission
PodSecurityPolicy was removed in Kubernetes 1.25. The built-in replacement is Pod Security Admission, configured with namespace labels.
apiVersion: v1
kind: Namespace
metadata:
name: my-app
labels:
pod-security.kubernetes.io/enforce: restricted
pod-security.kubernetes.io/enforce-version: latest
pod-security.kubernetes.io/warn: restricted
The baseline level blocks privileged pods, hostPath, host namespaces and dangerous capabilities. The restricted level also requires non-root, allowPrivilegeEscalation: false, dropping all capabilities and a seccomp profile. Before you enforce on an existing namespace, see what would break:
kubectl label --dry-run=server --overwrite ns my-app \
pod-security.kubernetes.io/enforce=restricted
The API server returns a warning for every running pod that would be rejected.
A pod spec that passes restricted:
apiVersion: v1
kind: Pod
metadata:
name: my-app
namespace: my-app
spec:
automountServiceAccountToken: false
securityContext:
runAsNonRoot: true
runAsUser: 10001
seccompProfile:
type: RuntimeDefault
containers:
- name: app
image: registry.example.com/my-app:1.4.2
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop: ["ALL"]
Set the seccomp profile explicitly. Kubernetes runs containers Unconfined unless the pod asks for a profile or the kubelet has seccompDefault enabled.
Find what is already running
kubectl get pods -A -o json | jq -r '
.items[]
| select(
.spec.hostPID or .spec.hostNetwork or .spec.hostIPC
or any(.spec.volumes[]?; .hostPath != null)
or any((.spec.containers + (.spec.initContainers // []))[];
.securityContext.privileged == true)
)
| "\(.metadata.namespace)/\(.metadata.name)"'
Some hits are legitimate: CNI plugins, log shippers, node exporters. Keep them in dedicated namespaces with privileged level, and keep application namespaces on restricted.
More layers
User namespaces. With hostUsers: false in the pod spec, root inside the container maps to an unprivileged UID on the host. The feature is stable since Kubernetes 1.36 and needs Linux 6.3 or newer, containerd 2.0+ or CRI-O 1.25+, and runc 1.2+ or crun 1.9+. It turns many escape bugs into "unprivileged user on the node" instead of root.
Admission policies for the rest. Pod Security Admission does not check where images come from or whether they are signed. Use ValidatingAdmissionPolicy (GA since 1.30), Kyverno or Gatekeeper for registry allowlists, signature checks and label rules.
Metadata endpoint. On AWS, require IMDSv2 and set the hop limit to 1 on nodes, so pods without hostNetwork cannot get node credentials. Give workloads their own identity with IRSA or EKS Pod Identity. Other clouds have equivalent workload identity features.
Sandboxed runtimes. For untrusted or multi-tenant code, run pods under gVisor or Kata Containers through a RuntimeClass. They put another kernel boundary between the workload and the host.
Runtime detection. Falco or Tetragon can alert on a shell started in a container, writes to system paths, access to the runtime socket and unexpected outbound connections. Detection does not prevent anything, but it shortens the time an attacker has.
Patching. Check kernel and runtime versions across nodes:
kubectl get nodes -o wide
The output includes kernel and container runtime versions. Compare them with the advisories for your distribution and node image, and roll nodes regularly instead of patching them in place.
Checklist
restrictedPod Security level on application namespaces,privilegedonly where a component really needs it.- No
hostPath, host namespaces or runtime socket mounts in application pods. - Non-root, no privilege escalation, all capabilities dropped,
RuntimeDefaultseccomp. automountServiceAccountToken: falseunless the pod calls the API.hostUsers: falsewhere the cluster supports it.- IMDSv2 with hop limit 1, workload identity instead of node credentials.
- Sidecars reviewed like application images.
- Nodes, kernel and runc kept current, runtime alerts routed to someone who reads them.
