Cloud Cost Optimization Without Hurting Reliability
Most cloud waste is boring. Resources nobody owns, instances sized for a load test two years ago, volumes left behind by deleted machines, data transfer nobody modeled. None of it needs clever tricks to fix, but the order of work matters. If you buy a three-year commitment before you clean up, you lock the waste in.
This post is the order that works: ownership first, then deletion, then rightsizing, then architecture, then discounts.
1. Make every resource belong to someone
You can't cut costs nobody is responsible for. Tag or label everything with at least a team and an environment, and make Terraform do it so nobody has to remember.
variable "team" {
type = string
validation {
condition = can(regex("^[a-z0-9-]+$", var.team))
error_message = "team must be a lowercase identifier, like payments."
}
}
provider "aws" {
region = "eu-central-1"
default_tags {
tags = {
team = var.team
environment = var.environment
managed-by = "terraform"
}
}
}
A few things that are easy to miss:
- On AWS, tags only show up in Cost Explorer after you activate them as cost allocation tags in the Billing console.
- On GCP, use labels (lowercase keys and values). They appear in the billing export to BigQuery.
- Tags in Terraform don't cover resources created in the console. To require tags at creation time, use an SCP that denies create actions when
aws:RequestTag/teamis missing, or Azure Policy with a "require tag" rule. - Separate accounts, projects or subscriptions per team and environment are the strongest boundary. Tags get forgotten, account boundaries don't.
2. Delete what nothing uses
Start with resources that are clearly orphaned. They carry no reliability risk once you confirm nothing depends on them.
# EBS volumes not attached to any instance
aws ec2 describe-volumes \
--filters Name=status,Values=available \
--query 'Volumes[].[VolumeId,Size,VolumeType,CreateTime]' \
--output table
# Elastic IPs not associated with anything
aws ec2 describe-addresses \
--query 'Addresses[?AssociationId==null].[PublicIp,AllocationId]' \
--output table
Other usual suspects:
- Old EBS snapshots and AMIs nobody restores from.
- Load balancers with no healthy targets or no traffic.
- Dev and preview environments that run all weekend.
- CloudWatch log groups with the default "Never expire" retention.
- Public IPv4 addresses. AWS charges for every public IPv4 address, attached or idle.
AWS Trusted Advisor and Compute Optimizer, and the GCP Recommender (idle VMs, idle disks, unused IPs), will produce most of this list for you. If you are not sure about a volume, snapshot it, delete it, and delete the snapshot later.
3. Rightsize based on real usage
Look at utilization over a full business cycle, including month-end jobs and batch runs, not a quiet Tuesday. Compute Optimizer and the GCP machine type recommendations do this for VMs. Downsize one step at a time and watch latency and error rates for a few days before the next step.
On Kubernetes you pay for nodes, and nodes are sized by pod requests, not by actual usage. Compare the two per namespace:
sum by (namespace) (kube_pod_container_resource_requests{resource="cpu"})
-
sum by (namespace) (rate(container_cpu_usage_seconds_total{container!=""}[5m]))
A large positive number means CPU that is reserved and paid for but not used. The Vertical Pod Autoscaler with updateMode: "Off" gives per-container request recommendations without changing anything. OpenCost or Kubecost will break cluster cost down by namespace and workload.
4. Fix storage and data transfer
These lines grow quietly because no single engineer sees them.
- gp2 to gp3. gp3 is priced lower per GB and lets you set IOPS and throughput separately from size. The change is an online volume modification, no downtime.
- Object storage lifecycle. Move old objects to cheaper storage classes and expire what you don't need. S3 Intelligent-Tiering helps when access patterns are unknown, but it has a per-object monitoring fee, so check object counts first.
- Cross-zone traffic. AWS and GCP both charge for traffic between availability zones in the same region. Chatty services spread across zones, or a replica in another zone serving all reads, add up.
- NAT gateways. You pay per GB processed. Traffic from private subnets to S3 or DynamoDB should go through VPC gateway endpoints, which have no extra charge.
- Egress to the internet. Put a CDN in front of static content and check what is actually being downloaded.
5. Buy commitments last
Once usage is clean and stable, commit to the baseline.
- AWS Savings Plans. Compute Savings Plans apply across instance families, regions, Fargate and Lambda. EC2 Instance Savings Plans and Reserved Instances give a bigger discount for less flexibility.
- GCP committed use discounts. Resource-based (vCPU and memory) or spend-based for specific services. Some machine families also get sustained use discounts automatically.
- Azure reservations and savings plans work on the same idea.
Commit to the floor of your usage, not the average. Unused commitment is pure waste.
Spot instances (AWS) and Spot VMs (GCP) are a different tool: much cheaper capacity that can be taken back. AWS gives a two-minute interruption notice, GCP gives 30 seconds. Use them for stateless workers, CI runners and batch jobs that can retry, and handle the termination signal.
Keeping reliability intact
- Change one thing at a time and keep a rollback path.
- Rightsize to peak plus headroom, not to the average.
- Keep SLO dashboards open for a week after each change.
- Don't remove redundancy (multi-AZ, replicas) to save money without an explicit decision from whoever owns the SLO.
Checklist
- Every resource has an owner tag or lives in an owned account.
- Orphaned volumes, IPs, snapshots and environments are deleted on a schedule.
- Log retention is set everywhere.
- Instance sizes and pod requests match measured usage plus headroom.
- Storage classes, cross-zone traffic and NAT traffic are reviewed.
- Commitments cover only the stable baseline.
