Beyond YAML Linting: Static Analysis for Kubernetes and Terraform
A YAML linter checks that a file is well-formed YAML. A manifest can pass that and still have a misspelled field the API server drops, a Service that selects no pods, or no resource requests at all. A Terraform change can be perfectly formatted and still destroy a database.
It helps to think in layers: syntax, schema, best-practice rules, your own policies, and finally analysis of what will actually change. Each layer catches a different class of mistakes.
Layer 1: YAML itself
yamllint does more than indentation. Two default rules catch real bugs:
key-duplicates. Many parsers silently keep only the last of two identical keys, so a secondenv:block replaces the first.truthy. In YAML 1.1, unquotedyes,no,onandoffare booleans. A country codeNOor a valueoffturns intofalse.
# .yamllint
extends: default
rules:
line-length: disable
truthy:
check-keys: false # allows "on:" in GitHub Actions workflows
Layer 2: schema validation
Check every object against the Kubernetes OpenAPI schema. kubeval is no longer maintained; kubeconform is its replacement. Validate the rendered output, not the templates:
helm template my-app ./chart -f values-production.yaml \
| kubeconform -strict -summary \
-schema-location default \
-schema-location 'https://raw.githubusercontent.com/datreeio/CRDs-catalog/main/{{.Group}}/{{.ResourceKind}}_{{.ResourceAPIVersion}}.json'
-strict rejects fields that are not in the schema. That is what catches resource: instead of resources: or readinessprobe instead of readinessProbe, mistakes that otherwise deploy fine and quietly do nothing. Set -kubernetes-version to your cluster's version so manifests using removed APIs fail here instead of at deploy time.
If CI can reach a cluster, kubectl apply --dry-run=server -f manifests/ goes one step further. It runs the real API server validation and admission webhooks, and catches changes to immutable fields.
Layer 3: best-practice rules
kube-linter, kube-score and Polaris check for patterns that cause outages and security problems. kube-linter's default checks include:
dangling-service: a Service whose selector matches no pods in the same set of manifests.mismatching-selector: a Deployment selector that does not match its pod template labels.unset-cpu-requirementsandunset-memory-requirements.latest-tag,privileged-container,run-as-non-root.- PodDisruptionBudget checks that flag budgets which block node drains.
Probe checks are not on by default. Enable them in the config:
# .kube-linter.yaml
checks:
include:
- no-readiness-probe
- no-liveness-probe
kube-linter lint ./chart --config .kube-linter.yaml
kube-linter renders Helm charts by itself when you point it at a chart directory.
Layer 4: your own policies
Generic tools do not know your rules: production needs at least two replicas, images come only from your registry, every workload has a team label. Write those as policy. Conftest runs Rego policies against any structured file. Current versions use Rego v1 syntax (if, contains) by default:
package main
deny contains msg if {
input.kind == "Deployment"
input.spec.replicas < 2
msg := sprintf("%s: production Deployments need at least 2 replicas", [input.metadata.name])
}
deny contains msg if {
input.kind in {"Deployment", "StatefulSet", "DaemonSet"}
some c in input.spec.template.spec.containers
not startswith(c.image, "registry.example.com/")
msg := sprintf("%s: image %s is not from registry.example.com", [input.metadata.name, c.image])
}
helm template my-app ./chart -f values-production.yaml | conftest test --policy policy/k8s -
Kyverno users can run the same policies they enforce in the cluster with the kyverno CLI in CI.
Layer 5: analyze the Terraform plan, not just the code
For Terraform, terraform validate and tflint catch syntax and provider-specific mistakes, such as invalid instance types. Trivy (trivy config, which absorbed tfsec) and Checkov find security misconfigurations like public buckets or unencrypted volumes.
None of them can tell you that this particular change destroys a resource. Only the plan knows that. Export it as JSON and run policy against it:
terraform plan -out=tfplan
terraform show -json tfplan > plan.json
conftest test plan.json --policy policy/terraform
package main
protected := {"aws_db_instance", "aws_rds_cluster", "aws_s3_bucket", "aws_dynamodb_table"}
deny contains msg if {
some rc in input.resource_changes
rc.type in protected
"delete" in rc.change.actions
msg := sprintf("%s would be destroyed (actions: %v)", [rc.address, rc.change.actions])
}
A replacement shows up as ["delete", "create"] or ["create", "delete"] in actions, so this also catches changes that force recreation. For cost changes, Infracost reads the same plan and reports the difference in the pull request.
Where to run what
- Pre-commit hooks: yamllint, formatters,
terraform validate. Fast feedback, easy to skip. - CI, as required checks: everything above, on rendered output and on the plan.
- In the cluster: admission policies (Kyverno, Gatekeeper or the built-in ValidatingAdmissionPolicy) as the last layer, for changes that never went through CI.
Start with schema validation and plan checks. They are cheap and catch the most damaging mistakes. Add custom policies as you find patterns worth enforcing.
Checklist
- yamllint with
key-duplicatesandtruthy. - kubeconform
-stricton rendered manifests, with CRD schemas and your cluster version. - kube-linter or kube-score in CI, with probe checks enabled.
- Team rules written as Rego or Kyverno policies.
- Terraform plan exported to JSON and checked for deletes of stateful resources.
- Admission policies in the cluster as the final guard.
