Rolling Out Policy as Code Without Fighting Your Engineers

Rolling Out Policy as Code Without Fighting Your Engineers

Reading time1 min
#devops#policy-as-code#opa#engineering

Rolling Out Policy as Code Without Fighting Your Engineers

Policy as code means writing rules like "every instance has an owner tag" or "no public S3 buckets" as code that a machine checks, usually with Open Policy Agent (OPA) and its language Rego, or with Kyverno or CEL in Kubernetes. The idea is sound. The rollout is where it goes wrong: rules appear at deploy time, error messages are cryptic, nobody knows who owns a rule, and there is no way to ask for an exception. Engineers then route around the tool, and the security team adds more rules.

The fix is mostly process. This post covers what causes the friction and how to roll out policies so they stick.

What engineers actually object to

  • Late feedback. The first time a rule fires is in the deploy pipeline or at the admission webhook, after review and merge.
  • Unclear messages. "Request denied by policy" without saying which field, why, or how to fix it.
  • No owner. Nobody can explain a rule or change it, so it never gets fixed when it is wrong.
  • No exceptions process, or one that takes weeks.
  • Rules nobody tested. A policy bug blocks every deploy in the company.
  • A new language. Rego is unfamiliar, and reviewing policy changes feels like a different job.

None of this is about compliance itself. It is about treating policy as something done to engineers instead of code they can read, run and change.

Run the same policies early

The same rule should run in three places: on the developer's machine, in CI on the pull request, and at the cluster or cloud API as a backstop. Conftest runs Rego against structured files such as Kubernetes manifests, Terraform plans and Dockerfiles, so you can use one policy repository for all three.

Example: require owner and cost-center tags on new EC2 instances, checked against the Terraform plan.

# policy/tags.rego
package main

required_tags := {"owner", "cost-center"}

deny contains msg if {
  rc := input.resource_changes[_]
  rc.type == "aws_instance"
  "create" in rc.change.actions
  tags := object.get(rc.change.after, "tags_all", {})
  missing := required_tags - {k | tags[k]}
  count(missing) > 0
  msg := sprintf("%s is missing required tags %v. Add them to the resource or to the provider default_tags.", [rc.address, missing])
}
terraform plan -out=tfplan
terraform show -json tfplan > tfplan.json
conftest test tfplan.json

The check uses tags_all, which includes provider default_tags, so it does not complain about tags set centrally. The message says what is missing and where to add it.

Write tests for every rule

Policies are code. They get unit tests and run in CI like any other code.

# policy/tags_test.rego
package main

test_missing_owner_is_denied if {
  count(deny) == 1 with input as {"resource_changes": [{
    "address": "aws_instance.web",
    "type": "aws_instance",
    "change": {"actions": ["create"], "after": {"tags_all": {"cost-center": "42"}}}
  }]}
}

test_tagged_instance_passes if {
  count(deny) == 0 with input as {"resource_changes": [{
    "address": "aws_instance.web",
    "type": "aws_instance",
    "change": {"actions": ["create"], "after": {"tags_all": {"owner": "payments", "cost-center": "42"}}}
  }]}
}
conftest verify --policy ./policy

A test per rule, plus a few real manifests from each team as fixtures, catches most policy bugs before they block anyone.

Warn first, then deny

Every new rule starts as a warning. Conftest has warn rules that print but do not fail the build. Gatekeeper has enforcementAction: warn and dryrun, Kyverno has Audit mode, and ValidatingAdmissionPolicy bindings have Warn and Audit.

While a rule is in warn mode, count how often it fires and for which teams. Fix the false positives, help teams clean up real violations, and announce a date for switching to deny. When that date comes, the switch is a non-event.

Exceptions are part of the design

There will always be a legitimate case that breaks a rule. If there is no official way to handle it, people disable the check. Make exceptions:

  • Requested through a pull request to the policy repository, with a reason.
  • Scoped narrowly: one resource or one namespace, not a whole team.
  • Time-limited, with an expiry date that the policy itself checks.
  • Visible: a list anyone can read.

A simple way is a data file the policies read with --data, so adding an exception is a reviewed one-line change instead of a Rego edit.

Ownership and review

  • Policies live in a repository with a CODEOWNERS file. Platform or security owns the framework; each rule has a named owner.
  • Each rule has a short description, the reason behind it (the control or risk it addresses) and a link to fix instructions.
  • Engineers can propose changes to rules. If a rule is wrong, fixing it should be as easy as fixing a bug in their own service.

Pick the right tool for the layer

  • Kubernetes admission: start with Pod Security Admission and ValidatingAdmissionPolicy (CEL). Use Gatekeeper or Kyverno when you need audit of existing resources, a policy library or more complex logic. Kyverno policies are YAML, which some teams find easier to review than Rego.
  • Terraform and other config files: Conftest in CI, plus cloud-native guardrails (AWS SCPs, Azure Policy, GCP Organization Policies) for the rules that must hold even if someone bypasses the pipeline.

Using one language everywhere is nice, but not at the cost of fighting a tool that does not fit the layer.

Checklist

  • Same policies run locally, in CI and at the API.
  • Every rule has an owner, a reason, tests and an error message that says how to fix it.
  • New rules start in warn or audit mode with a published date for enforcement.
  • Exceptions are reviewed, narrow, time-limited and visible.
  • Policy changes go through pull requests that engineers can open too.
  • Violation counts are tracked per rule, so noisy or useless rules get fixed or removed.