Kubernetes Admission Controllers: Enforcing Policy at the API Server
Every create, update and delete request to the Kubernetes API goes through admission after authentication and authorization. Admission is where you can change a request or reject it. For compliance this is the right place: a rule enforced here applies to kubectl, CI, GitOps tools and controllers alike, and nobody can skip it by deploying a different way.
How admission works
The order is fixed:
- Authentication and authorization.
- Mutating admission: built-in mutating plugins, MutatingAdmissionPolicy and mutating webhooks. They can change the object.
- Schema validation of the result.
- Validating admission: built-in validating plugins, ValidatingAdmissionPolicy and validating webhooks. They can only allow or deny.
- The object is stored in etcd.
There are three ways to add rules:
- Built-in plugins compiled into the API server, such as
NamespaceLifecycle,LimitRanger,ResourceQuotaandPodSecurity. A default set is enabled. On managed clusters you generally cannot change the list. - CEL admission policies. ValidatingAdmissionPolicy (GA since Kubernetes 1.30) and MutatingAdmissionPolicy (GA since 1.36) run CEL expressions inside the API server. No extra service to run.
- Webhooks. The API server calls an HTTPS service you run. Most flexible, and the most operational risk. OPA Gatekeeper and Kyverno work this way.
Start with what is built in
Pod Security Admission replaced PodSecurityPolicy, which was removed in Kubernetes 1.25. You enable it with namespace labels:
kubectl label namespace my-app \
pod-security.kubernetes.io/enforce=baseline \
pod-security.kubernetes.io/warn=restricted \
pod-security.kubernetes.io/audit=restricted
enforce rejects pods, warn returns a warning to the client and audit writes an annotation to the audit log. Running warn and audit one level stricter than enforce shows what would break before you tighten it.
ResourceQuota and LimitRange cover resource limits per namespace without any policy engine.
CEL policies for custom rules
For rules that Pod Security Admission does not cover, a ValidatingAdmissionPolicy is usually enough. This one requires all images to come from an internal registry:
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicy
metadata:
name: images-from-internal-registry
spec:
failurePolicy: Fail
matchConstraints:
resourceRules:
- apiGroups: [""]
apiVersions: ["v1"]
operations: ["CREATE", "UPDATE"]
resources: ["pods"]
validations:
- expression: >-
object.spec.containers.all(c, c.image.startsWith('registry.example.com/')) &&
(!has(object.spec.initContainers) ||
object.spec.initContainers.all(c, c.image.startsWith('registry.example.com/')))
message: "Images must come from registry.example.com."
---
apiVersion: admissionregistration.k8s.io/v1
kind: ValidatingAdmissionPolicyBinding
metadata:
name: images-from-internal-registry
spec:
policyName: images-from-internal-registry
validationActions: [Warn, Audit]
matchResources:
namespaceSelector:
matchExpressions:
- key: kubernetes.io/metadata.name
operator: NotIn
values: ["kube-system"]
The binding starts with Warn and Audit. Once the warnings stop, switch to Deny. Deny and Warn cannot be combined in one binding.
Ephemeral containers are added through the pods/ephemeralcontainers subresource, so cover that subresource too if debug containers matter for your policy.
Validate pods, not only Deployments
A rule that only matches deployments misses StatefulSets, DaemonSets, Jobs, CronJobs and bare pods. A rule that matches pods catches everything, but there is a catch: when a Deployment's pod is rejected, kubectl apply succeeds and the error only shows up in the ReplicaSet's events. Engineers see a Deployment that never becomes ready.
Common solutions: match both pods and the workload kinds, or use the engine's support for this. Kyverno generates workload rules from pod rules automatically, and Gatekeeper has an expansion feature that validates the pod template inside a workload.
Running webhooks safely
A webhook is in the request path of everything it matches. Treat it like critical infrastructure.
- failurePolicy. The default is
Fail: if the webhook is unreachable, matching requests are rejected. That is the safe choice for security rules and a source of outages if the webhook is fragile.Ignorekeeps the cluster working but lets requests through unchecked. - Exclude the webhook's own namespace and
kube-systemwith anamespaceSelector. Otherwise a webhook that is down can block the pods that would bring it back. - Timeouts.
timeoutSecondsdefaults to 10 and must be between 1 and 30. Keep the webhook fast. Do not run slow work like a full image vulnerability scan inside it. - Scope narrowly. Match only the resources and operations you need.
matchConditions(CEL) filters requests before the call is made. - High availability. At least two replicas, a PodDisruptionBudget and spread across nodes.
- Side effects. Declare
sideEffects: Nonewhen the webhook has none, so server-side dry run works.
Image scanning and signatures
Scanning images at admission time is slow and depends on an external service. A more reliable split is to scan and sign images in CI, then verify at admission that the image has a valid signature or attestation. Kyverno's verifyImages rules and the Sigstore policy-controller do that verification. Pinning images by digest prevents a tag from being moved to a different image after it was checked.
Testing policies
kubectl apply --dry-run=server -f manifest.yamlruns admission without storing anything, so it shows denials and warnings.- Keep a set of known-good and known-bad manifests in the policy repository and run them in CI against a kind cluster.
- Watch the API server metrics for webhook latency and rejections (
apiserver_admission_webhook_admission_duration_seconds,apiserver_admission_webhook_rejection_count).
Checklist
- Pod Security Admission labels on every namespace,
warnandauditone level stricter thanenforce. - Quotas and limit ranges per namespace.
- Custom rules as ValidatingAdmissionPolicy where CEL is enough; webhooks only when needed.
- Rules match pods and workload kinds, including init containers.
- New rules start in warn or audit mode.
- Webhooks: narrow scope, own namespace excluded, timeouts set, highly available.
- Images scanned and signed in CI, signatures verified at admission.
