FinOps for DevOps: Turning Cost Reports into Engineering Tickets

FinOps for DevOps: Turning Cost Reports into Engineering Tickets

Reading time1 min
#finops#devops#engineering#cloud#cost-management

FinOps for DevOps: Turning Cost Reports into Engineering Tickets

A cost report on its own changes nothing. It shows spend by service or account, arrives once a month, and has no owner and no proposed fix. Engineers can't act on "EC2 went up". They can act on "the batch workers in the reporting namespace request four times the CPU they use, here is the graph, lowering requests should remove two nodes".

The work is to turn the first kind of statement into the second kind, routinely.

Make costs attributable

Every cost line needs an owner. In order of reliability:

  1. Accounts, projects, subscriptions. One per team and environment where possible. Nothing can be untagged at this level.
  2. Tags and labels. team, environment, service. Apply them through Terraform defaults (default_tags in the AWS provider, default_labels in the Google provider) so nobody has to remember. On AWS, activate them as cost allocation tags in the Billing console, or they won't appear in Cost Explorer.
  3. Kubernetes namespaces. Cloud billing stops at the node. To split a cluster by team you need OpenCost, Kubecost or your provider's cost allocation for GKE or EKS, which split node cost across pods based on their requests and usage.

Decide upfront how to split shared costs (clusters, networking, support plans). Proportional to direct spend is simple and good enough. Write the rule down so nobody argues about it every month.

Get raw data, not only dashboards

Dashboards are fine for looking. For finding work you want data you can query.

  • AWS: Data Exports (CUR 2.0 or FOCUS format) to S3, queried with Athena. For quick checks, the Cost Explorer API.
  • GCP: Cloud Billing export to BigQuery.
  • Azure: Cost Management exports to a storage account.

A quick look at last month by team and service:

aws ce get-cost-and-usage \
  --time-period Start=2025-09-01,End=2025-10-01 \
  --granularity MONTHLY \
  --metrics UnblendedCost \
  --group-by Type=TAG,Key=team Type=DIMENSION,Key=SERVICE \
  --output json

Costs with an empty team value are your first ticket.

Look at changes, not totals

The total bill is a poor signal. Large stable lines are usually known. What finds work:

  • Week over week, by team and service. Sort by absolute change.
  • Anomaly detection. AWS Cost Anomaly Detection learns a baseline and flags unusual spend per service, account or cost category.
  • Budgets with forecast alerts. Alert when the forecast crosses the budget, not after the month is over.
resource "aws_budgets_budget" "payments" {
  name         = "team-payments-monthly"
  budget_type  = "COST"
  limit_amount = "5000"
  limit_unit   = "USD"
  time_unit    = "MONTHLY"

  cost_filter {
    name   = "TagKeyValue"
    values = ["user:team$payments"]
  }

  notification {
    comparison_operator        = "GREATER_THAN"
    threshold                  = 90
    threshold_type             = "PERCENTAGE"
    notification_type          = "FORECASTED"
    subscriber_email_addresses = ["payments-team@example.com"]
  }
}

Send the alert to the owning team's channel, not to a central FinOps inbox.

Write tickets engineers can act on

A useful cost ticket has the same shape as a good bug report.

Title: Reduce CPU requests for report-worker (reporting namespace)

Owner: team-reporting
Current cost: from the cost export, monthly, with the query used
Evidence: requests vs actual usage over 30 days (graph link)
Proposed change: lower CPU requests from 2 to 500m, keep memory as is
Expected effect: fewer nodes in the shared pool; estimate and how it was calculated
Risk: CPU throttling during month-end runs
Validation: p95 job duration and throttling metrics for two weeks
Rollback: revert the Helm values change

Two fields matter most. "How the estimate was calculated" keeps the numbers honest. "Validation" makes sure the saving is checked after the change instead of assumed.

Prioritize like any other backlog

Score each ticket by expected monthly saving, effort and risk. In practice the backlog splits into three groups:

  • Cleanup. Orphaned volumes, idle environments, missing log retention. Low risk, do them in bulk.
  • Rightsizing. Instance sizes, pod requests, database tiers. Medium risk, needs a validation period.
  • Architecture. Cross-zone traffic, NAT gateway usage, storage classes, caching. Higher effort, often the biggest savings, plan them as normal projects.

After a change ships, compare the same cost line before and after over the same number of days. Record the actual result in the ticket. Over time this tells you which kinds of estimates you can trust.

Make it routine

  • A short weekly review per team: top changes in their spend, open anomalies, status of cost tickets.
  • Cost in code review. Infracost can comment the estimated monthly cost change of a Terraform pull request.
  • Defaults that prevent waste. Log retention in modules, lifecycle rules on buckets, ResourceQuota and LimitRange per namespace. Note that once a namespace has a quota on requests.cpu, pods without CPU requests are rejected unless a LimitRange sets defaults.
  • Unit costs. Cost per request, per customer or per job, tracked next to the absolute numbers. Growth that raises the total but lowers the unit cost is fine. A rising unit cost is a ticket.

Checklist

  • Every cost line maps to a team through accounts, tags or namespace allocation.
  • Billing data is exported somewhere you can query.
  • Budgets alert on forecast, anomaly detection is on, alerts go to owners.
  • Cost tickets include evidence, an estimate with its method, risk and validation.
  • Savings are measured after the change, not assumed.
  • Cost review is a regular part of team rituals, not a quarterly surprise.