Cloud SQL vs Self-Managed PostgreSQL on Kubernetes: Cost and Latency

Cloud SQL vs Self-Managed PostgreSQL on Kubernetes: Cost and Latency

Reading time1 min
#cloud#sql#devops#kubernetes#latency

Cloud SQL vs Self-Managed PostgreSQL on Kubernetes: Cost and Latency

The choice looks like "convenience versus control", but the real comparison is narrower: what you pay per month, what latency your application sees, and how much operational work your team takes on. All three can be measured. This post covers what to measure and the usual mistakes on both sides.

What "self-managed on Kubernetes" should mean

A Deployment with replicas: 3 and the postgres image is not a database cluster. It is three unrelated databases, and without a PersistentVolumeClaim the data disappears when the pod is rescheduled. A real setup needs an operator that handles replication, failover, backups and volumes. CloudNativePG, Zalando's postgres-operator and Crunchy PGO are the common choices.

apiVersion: postgresql.cnpg.io/v1
kind: Cluster
metadata:
  name: my-app-db
spec:
  instances: 3
  storage:
    size: 200Gi
    storageClass: premium-rwo
  resources:
    requests:
      cpu: "2"
      memory: 8Gi
    limits:
      memory: 8Gi
  affinity:
    enablePodAntiAffinity: true
    topologyKey: topology.kubernetes.io/zone
    podAntiAffinityType: required
  postgresql:
    parameters:
      shared_buffers: 2GB

This gives a primary and two streaming replicas in different zones, with automated failover. You still have to configure backups to object storage, monitoring, and a tested restore procedure.

Cost: what each option bills for

Cloud SQL pricing dimensions:

  • Edition: Enterprise or Enterprise Plus. Enterprise Plus costs more and adds features like a data cache and shorter maintenance downtime.
  • vCPU and memory per hour for the instance tier.
  • Storage per GB per month. Storage can grow automatically but can never be shrunk.
  • High availability adds a standby in another zone, billed at HA rates.
  • Read replicas are full instances with their own compute and storage.
  • Backups beyond the included amount, and network egress.
  • Committed use discounts are available for Cloud SQL.

Self-managed on GKE pays for:

  • Nodes sized for the database plus headroom, usually a dedicated node pool so other workloads don't compete for memory and I/O.
  • A persistent disk per instance. Three instances means three full copies of the data.
  • Traffic between zones for replication, which GCP bills.
  • Backup storage in Cloud Storage.
  • Engineering time: upgrades, operator updates, backup testing, failover drills, on-call. This is usually the largest cost and the one most often left out of the comparison.

Price both options with the Google Cloud pricing calculator using the same vCPU, memory, storage and HA layout. Then add an honest estimate of engineer hours per month for the self-managed one.

Latency: where it actually comes from

Neither option is inherently faster. Differences come from configuration.

  • Network path. An application in one zone talking to a primary in another zone pays a cross-zone round trip on every query. Check where the primary is and where the pods run.
  • Connection setup. New connections cost TLS and authentication round trips. Applications that open a connection per request feel this on both options. Use a pool in the application or PgBouncer.
  • Proxy hops. The Cloud SQL Auth Proxy adds a local hop. It is usually small compared to the query, but measure it, and connect over private IP.
  • Disk. On Cloud SQL and on GCE persistent disks, IOPS and throughput depend on disk type and size. A small disk can be the bottleneck even when CPU is idle.
  • Memory. If the working set doesn't fit in shared_buffers and the OS page cache, reads go to disk. Check the buffer cache hit ratio before buying more CPU.
  • Neighbors. Pods on shared nodes compete for CPU, memory bandwidth and disk. A dedicated node pool removes that variable.

Measure it yourself

Run the same benchmark against both, from a pod in the same cluster and zone as your application, with a dataset of realistic size relative to memory.

# initialize a test dataset (scale factor sets the size)
pgbench -i -s 100 -h 10.10.0.5 -U postgres bench

# 5-minute mixed read/write run, progress every 10 s, per-statement latency
pgbench -c 16 -j 4 -T 300 -P 10 -r -h 10.10.0.5 -U postgres bench

# read-only run
pgbench -S -c 16 -j 4 -T 300 -P 10 -h 10.10.0.5 -U postgres bench

Look at p95 and p99, not only the average. Better still, replay your own slowest queries with EXPLAIN (ANALYZE, BUFFERS) on both. Also test failover: trigger it on each setup and measure how long the application sees errors.

Running Cloud SQL well

  • Private IP, same region as the application, and the application in the primary's zone where possible.
  • A connection pool, since max_connections is tied to instance memory.
  • Query Insights enabled to find slow queries before scaling up.
  • A maintenance window outside peak hours.
  • Plan tier changes. Changing the machine type restarts the instance:
gcloud sql instances patch my-app-db --tier=db-custom-4-16384

Running PostgreSQL on Kubernetes well

  • An operator, not hand-written StatefulSets.
  • Dedicated nodes, memory requests equal to limits, anti-affinity across zones.
  • Continuous WAL archiving and base backups to object storage, with restores tested on a schedule.
  • Monitoring of replication lag, disk usage, connection count and checkpoints.
  • A plan for major version upgrades before you need one.

How to decide

Cloud SQL usually wins when the team is small, the database is important but not unusual, and engineer time is more expensive than the managed service premium. Self-managed makes sense when you need extensions or settings Cloud SQL doesn't allow, already run stateful workloads on Kubernetes with the skills and on-call to match, or have many small databases where per-instance managed pricing adds up.

Checklist

  • Compare the same layout (vCPU, memory, storage, HA) in the pricing calculator.
  • Add engineer hours to the self-managed side.
  • Benchmark both from the application's zone with pgbench and your real queries.
  • Check zone placement, connection pooling, disk size and cache hit ratio before blaming either option.
  • Test failover and restore on whichever option you choose.