Container Escape Paths in Kubernetes and How to Close Them

Container Escape Paths in Kubernetes and How to Close Them

Reading time1 min
#devops#containers#security#kubernetes

Container Escape Paths in Kubernetes and How to Close Them

A container is a normal Linux process with namespaces, cgroups, dropped capabilities, a seccomp filter and usually an AppArmor or SELinux profile around it. It shares the kernel with every other container on the node. "Escape" means one of those layers was weakened by configuration or broken by a bug, and the process ends up with access to the node.

Most exposure in real clusters comes from configuration, not from new exploits. That is good news, because configuration can be checked and enforced.

The main paths

Privileged containers. privileged: true turns off most of the isolation: all capabilities, host devices, no seccomp. Treat a privileged pod as root on the node.

Host mounts. A hostPath volume of /, /etc, /var/lib/kubelet or the container runtime socket (/run/containerd/containerd.sock, /var/run/docker.sock) gives the container control over the node. Access to the runtime socket is equivalent to root, because it can start new privileged containers.

Host namespaces. hostPID, hostIPC and hostNetwork remove the separation from node processes and node networking. With hostNetwork the pod also reaches everything the node can reach, including services bound to localhost and the cloud metadata endpoint.

Extra capabilities. SYS_ADMIN, SYS_PTRACE, SYS_MODULE and NET_ADMIN each widen what the process can do against the kernel. Most applications need none of them.

Root plus privilege escalation. Without user namespaces, UID 0 in the container is UID 0 for the kernel. If allowPrivilegeEscalation is not false, setuid binaries inside the image can raise privileges. Running as non-root with no_new_privs set removes a whole class of problems.

Kernel and runtime bugs. All containers share one kernel, so a kernel privilege escalation is also a container escape. Container runtimes have had escape bugs too. Public examples are runc CVE-2019-5736, CVE-2024-21626 (fixed in runc 1.1.12) and the set CVE-2025-31133, CVE-2025-52565 and CVE-2025-52881 (fixed in runc 1.2.8, 1.3.3 and 1.4.0-rc.3). On the kernel side, Dirty Pipe (CVE-2022-0847) is the well-known one. Configuration does not fix these. Patching does.

Credentials, not escape. A stolen service account token with broad RBAC, or node IAM credentials from the metadata endpoint, give an attacker the same result without touching the kernel.

One note on sidecars: containers in the same pod share a network namespace, so a sidecar can talk to the main container on localhost. NetworkPolicy does not filter traffic inside a pod. The pod is the trust boundary, so vet sidecar images the same way as the application image.

Enforce a baseline with Pod Security Admission

PodSecurityPolicy was removed in Kubernetes 1.25. The built-in replacement is Pod Security Admission, configured with namespace labels.

apiVersion: v1
kind: Namespace
metadata:
  name: my-app
  labels:
    pod-security.kubernetes.io/enforce: restricted
    pod-security.kubernetes.io/enforce-version: latest
    pod-security.kubernetes.io/warn: restricted

The baseline level blocks privileged pods, hostPath, host namespaces and dangerous capabilities. The restricted level also requires non-root, allowPrivilegeEscalation: false, dropping all capabilities and a seccomp profile. Before you enforce on an existing namespace, see what would break:

kubectl label --dry-run=server --overwrite ns my-app \
  pod-security.kubernetes.io/enforce=restricted

The API server returns a warning for every running pod that would be rejected.

A pod spec that passes restricted:

apiVersion: v1
kind: Pod
metadata:
  name: my-app
  namespace: my-app
spec:
  automountServiceAccountToken: false
  securityContext:
    runAsNonRoot: true
    runAsUser: 10001
    seccompProfile:
      type: RuntimeDefault
  containers:
    - name: app
      image: registry.example.com/my-app:1.4.2
      securityContext:
        allowPrivilegeEscalation: false
        readOnlyRootFilesystem: true
        capabilities:
          drop: ["ALL"]

Set the seccomp profile explicitly. Kubernetes runs containers Unconfined unless the pod asks for a profile or the kubelet has seccompDefault enabled.

Find what is already running

kubectl get pods -A -o json | jq -r '
  .items[]
  | select(
      .spec.hostPID or .spec.hostNetwork or .spec.hostIPC
      or any(.spec.volumes[]?; .hostPath != null)
      or any((.spec.containers + (.spec.initContainers // []))[];
             .securityContext.privileged == true)
    )
  | "\(.metadata.namespace)/\(.metadata.name)"'

Some hits are legitimate: CNI plugins, log shippers, node exporters. Keep them in dedicated namespaces with privileged level, and keep application namespaces on restricted.

More layers

User namespaces. With hostUsers: false in the pod spec, root inside the container maps to an unprivileged UID on the host. The feature is stable since Kubernetes 1.36 and needs Linux 6.3 or newer, containerd 2.0+ or CRI-O 1.25+, and runc 1.2+ or crun 1.9+. It turns many escape bugs into "unprivileged user on the node" instead of root.

Admission policies for the rest. Pod Security Admission does not check where images come from or whether they are signed. Use ValidatingAdmissionPolicy (GA since 1.30), Kyverno or Gatekeeper for registry allowlists, signature checks and label rules.

Metadata endpoint. On AWS, require IMDSv2 and set the hop limit to 1 on nodes, so pods without hostNetwork cannot get node credentials. Give workloads their own identity with IRSA or EKS Pod Identity. Other clouds have equivalent workload identity features.

Sandboxed runtimes. For untrusted or multi-tenant code, run pods under gVisor or Kata Containers through a RuntimeClass. They put another kernel boundary between the workload and the host.

Runtime detection. Falco or Tetragon can alert on a shell started in a container, writes to system paths, access to the runtime socket and unexpected outbound connections. Detection does not prevent anything, but it shortens the time an attacker has.

Patching. Check kernel and runtime versions across nodes:

kubectl get nodes -o wide

The output includes kernel and container runtime versions. Compare them with the advisories for your distribution and node image, and roll nodes regularly instead of patching them in place.

Checklist

  • restricted Pod Security level on application namespaces, privileged only where a component really needs it.
  • No hostPath, host namespaces or runtime socket mounts in application pods.
  • Non-root, no privilege escalation, all capabilities dropped, RuntimeDefault seccomp.
  • automountServiceAccountToken: false unless the pod calls the API.
  • hostUsers: false where the cluster supports it.
  • IMDSv2 with hop limit 1, workload identity instead of node credentials.
  • Sidecars reviewed like application images.
  • Nodes, kernel and runc kept current, runtime alerts routed to someone who reads them.