Platform Engineering Without a Platform Team

Platform Engineering Without a Platform Team

Reading time1 min
#platform engineering#startup#devops#kubernetes#terraform

Platform Engineering Without a Platform Team

Platform engineering means making the common path self-service and consistent: create a service, build it, deploy it, see its logs and metrics, get a database. Large companies staff a team to build that path. A team of five engineers cannot, but it still needs the result. Without it, every service gets deployed a little differently and only the person who set it up knows how it works.

The small-team version is not a developer portal. It is fewer choices, shared defaults stored as code, and a clear owner for the boring work.

Pick one of everything

Every option you support is something to patch, upgrade, document and debug. Choose one of each and write the choice down:

  • one CI system,
  • one way to deploy (for example, a Helm chart applied from CI),
  • one infrastructure-as-code tool and one state backend,
  • one logging, metrics and alerting stack,
  • one or two languages and their base images.

Exceptions are fine when there is a real reason, but they need an owner who accepts the extra work.

Buy before you build

Managed services are the platform team you do not have to hire. A managed database with automated backups and point-in-time recovery, hosted CI, a managed container platform and a hosted observability stack cost money. Compare that with the engineering hours it takes to run, patch and back up the same things yourself, and count those hours honestly.

Also question whether you need Kubernetes at all. For a handful of stateless services, a simpler managed runtime (ECS, Cloud Run, App Service or a PaaS) removes cluster upgrades, add-on management and node patching from the list. Kubernetes makes sense when you need what it offers, not as a default.

Self-host only what is core to the product or what has no reasonable managed option.

Put the golden path in code

A wiki page explaining how to deploy goes out of date. Templates and shared code do not, because services use them directly.

A service template repository with a Dockerfile, the deploy config, the CI workflow, health and readiness endpoints, structured logging and a metrics endpoint already wired up. A new service starts as a copy of it.

Reusable CI workflows. With GitHub Actions, put the build and deploy logic in one repository and call it from every service:

# my-org/platform/.github/workflows/build-and-deploy.yml
name: build-and-deploy
on:
  workflow_call:
    inputs:
      app:
        type: string
        required: true
      environment:
        type: string
        default: staging

jobs:
  build:
    runs-on: ubuntu-latest
    permissions:
      contents: read
      packages: write
    steps:
      - uses: actions/checkout@v5
      - uses: docker/login-action@v3
        with:
          registry: ghcr.io
          username: ${{ github.actor }}
          password: ${{ secrets.GITHUB_TOKEN }}
      - uses: docker/build-push-action@v6
        with:
          push: true
          tags: ghcr.io/my-org/${{ inputs.app }}:${{ github.sha }}

  deploy:
    needs: build
    runs-on: ubuntu-latest
    environment: ${{ inputs.environment }}
    steps:
      - uses: actions/checkout@v5
      # cluster credentials (OIDC to your cloud provider) go here
      - run: |
          helm upgrade --install ${{ inputs.app }} oci://ghcr.io/my-org/charts/service \
            --version 1.4.0 --namespace ${{ inputs.app }} --create-namespace \
            -f deploy/values-${{ inputs.environment }}.yaml \
            --set image.tag=${{ github.sha }} --wait

Each service's own workflow is then a few lines:

# my-app/.github/workflows/deploy.yml
name: deploy
on:
  push:
    branches: [main]
jobs:
  deploy:
    uses: my-org/platform/.github/workflows/build-and-deploy.yml@v1
    permissions:
      contents: read
      packages: write
    with:
      app: my-app
    secrets: inherit

The called workflow can't get more token permissions than the caller grants, which is why the caller sets them. If the platform repository is private, allow access to its workflows from other repositories in the organization in the repository's Actions settings. Tag releases of the platform repository, so a change to the shared workflow does not reach every service at once.

One shared Helm chart for all services, with per-service values files. Upgrading probes, security context or labels then happens in one place.

Versioned Terraform modules for the things every service needs, such as a database, a bucket or a queue:

module "db" {
  source = "git::https://github.com/my-org/terraform-modules.git//postgres?ref=v2.3.0"

  name        = "my-app"
  environment = "production"
}

The module sets backups, encryption, deletion protection and monitoring once. Services pin a version and upgrade on purpose.

Own the boring work explicitly

Without a platform team, platform work happens when someone gets annoyed enough. Make it a role instead:

  • A rotating platform duty. One engineer per sprint or per month handles upgrades, dependency bumps, certificate and credential rotation, and cost review. Rotating spreads the knowledge.
  • A platform backlog that lives next to product work and gets real capacity in planning.
  • Automated dependency updates with Renovate or Dependabot, for base images, actions, Helm charts and Terraform providers.
  • Upgrades on the calendar. Managed Kubernetes versions, database major versions and runtime versions reach end of support on published dates. Look them up and schedule the work before the deadline.

Measure the path

Use the four DORA metrics as a check on whether the platform helps: deployment frequency, lead time for changes, change failure rate and time to restore service. Your CI/CD system and incident log already have the data. If deploys get slower or riskier as services are added, the golden path needs work.

Also time how long it takes to go from an empty repository to a service running in staging using only the template and the docs. Ask the next new hire to do it and note every point where they had to ask someone.

When to form a platform team

The rotating duty stops working when it constantly takes most of someone's time, or when product teams regularly wait on platform work. At that point a dedicated team pays for itself, and the templates, modules and shared workflows are already its starting point.

Checklist

  • One default for CI, deploy, IaC, observability and languages, written down.
  • Managed services first, self-hosting only with a reason.
  • Service template, reusable CI workflows, a shared Helm chart and versioned Terraform modules.
  • A rotating platform duty and a platform backlog with real capacity.
  • Automated dependency updates and scheduled version upgrades.
  • DORA metrics and time to first deploy as the measure of the platform.