Self-Hosted CI Runners: Faster Builds Without Exposing Secrets

Self-Hosted CI Runners: Faster Builds Without Exposing Secrets

Reading time1 min
#devops#ci#self-hosted#kubernetes#secrets

Self-Hosted CI Runners: Faster Builds Without Exposing Secrets

A hosted runner is a fresh VM for every job, thrown away afterwards. A self-hosted runner is whatever you make it. You get control over hardware, network and caches, and you also take over isolation, patching and credential handling.

This post covers where the speed actually comes from, what can go wrong with secrets, and a setup that keeps both under control.

Find out where the time goes first

Look at step timings for your slowest pipelines before buying hardware:

  • Queue time means not enough runners. Autoscaling fixes it.
  • Dependency and image downloads mean poor caching or a slow network path to registries.
  • CPU-bound compile and test steps benefit from bigger machines or specific architectures.
  • Waiting on external services is not fixed by a faster runner at all.

Where self-hosting helps

  • Network proximity. Runners in the same region or VPC as your registry, artifact storage and internal services.
  • Machine shape. More cores, more memory, ARM, GPUs.
  • Shared caches. Keep caches close to the runners, but outside the runner itself, so runners can stay disposable:
docker buildx build \
  --cache-from type=registry,ref=registry.example.com/my-app:buildcache \
  --cache-to type=registry,ref=registry.example.com/my-app:buildcache,mode=max \
  -t "registry.example.com/my-app:${GIT_SHA}" --push .

The same idea applies to package managers: a pull-through proxy for npm, Maven or PyPI in the runner network, and an object-storage backend for the CI cache (GitLab Runner supports S3 and GCS for [runners.cache]).

The security model

Any job that runs on a runner can read what that runner can read: files left by earlier jobs, credentials on the machine or in its environment, the cloud metadata endpoint and every internal service reachable from its network. GitHub's own documentation says self-hosted runners should almost never be used for public repositories, because anyone can open a pull request and run code on them. A persistent runner can be compromised once and stay compromised for every job after that.

Isolation rules

1. One job per runner. Use ephemeral runners that are destroyed after a single job. On GitHub, Actions Runner Controller (ARC) runner scale sets create a fresh pod per job, and VM-based runners can be registered with --ephemeral or as just-in-time runners. On GitLab, the Kubernetes executor runs every job in a new pod, and the autoscaling executors can do the same with VMs.

2. Separate pools by trust. Builds of unreviewed branches run on a pool with no deploy credentials. Deploy jobs run on a different pool. On GitHub, use runner groups limited to specific repositories and workflows. On GitLab, mark deploy runners as protected so they only take jobs from protected branches and tags. GitLab also keeps caches for protected and unprotected branches apart by default. Leave it that way.

3. No privileged containers and no Docker socket. Mounting /var/run/docker.sock gives every job root on the host. Build images with rootless BuildKit or Buildah instead. In ARC, containerMode.type: dind runs a privileged Docker sidecar. kubernetes mode runs job containers as separate pods instead.

4. Short-lived credentials. Replace static cloud keys with OIDC. The job gets a signed token, the cloud trusts it only for a specific repository and environment, and the credentials expire after the job.

5. Block the metadata endpoint. Job pods should not be able to reach 169.254.169.254 and borrow the node's cloud role. On AWS, also require IMDSv2 with a hop limit of 1 so containers outside the host network cannot reach it.

6. Do not rely on log masking. Masking matches exact strings. A secret printed in base64, split across lines or sent over the network is not masked.

Example: ARC on Kubernetes

Runner scale set values:

githubConfigUrl: https://github.com/example-org
githubConfigSecret: arc-github-app
runnerGroup: deploy
minRunners: 1
maxRunners: 20
containerMode:
  type: kubernetes
  kubernetesModeWorkVolumeClaim:
    accessModes: ["ReadWriteOnce"]
    storageClassName: standard
    resources:
      requests:
        storage: 10Gi

Network policy for the runner namespace:

apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: block-metadata
  namespace: arc-runners
spec:
  podSelector: {}
  policyTypes: ["Egress"]
  egress:
    - to:
        - ipBlock:
            cidr: 0.0.0.0/0
            except:
              - 169.254.169.254/32

This only works if your CNI enforces NetworkPolicy.

A deploy workflow that uses OIDC instead of stored keys:

on:
  push:
    branches: [main]

permissions:
  contents: read
  id-token: write

jobs:
  deploy:
    runs-on: arc-runner-set
    environment: production
    steps:
      - uses: actions/checkout@v7
      - uses: aws-actions/configure-aws-credentials@v6
        with:
          role-to-assume: arn:aws:iam::123456789012:role/my-app-deploy
          aws-region: eu-central-1
      - run: ./scripts/deploy.sh

runs-on is the Helm release name of the runner scale set. In the IAM role trust policy, restrict token.actions.githubusercontent.com:sub to repo:example-org/my-app:environment:production. Only jobs in the production environment can then assume the role, and the environment's branch rules decide which branches get there. GitLab offers the same through id_tokens in .gitlab-ci.yml.

Keep it maintained

  • Update runner images and the runner agent regularly. Old agents stop working when the CI service drops support.
  • Pin third-party actions and images by digest or full commit SHA.
  • Send runner logs and cloud audit logs somewhere jobs cannot delete them.

Checklist

  • Measure where pipeline time goes before scaling hardware.
  • Runners are ephemeral, one job each.
  • Untrusted builds and deploy jobs run on separate runner pools.
  • No Docker socket, no privileged containers.
  • Cloud access through OIDC with a narrow sub condition.
  • Metadata endpoint blocked from job pods.
  • Caches live outside the runner and are split by trust level.
  • No self-hosted runners for public repositories.