Running AWS, GCP and Azure Together: Keeping Multi-Cloud Manageable

Running AWS, GCP and Azure Together: Keeping Multi-Cloud Manageable

Reading time1 min
#multi-cloud#devops#aws#gcp#azure#kubernetes#terraform

Running AWS, GCP and Azure Together: Keeping Multi-Cloud Manageable

Most companies that run on more than one cloud didn't plan it. An acquisition brought an Azure tenant, a data team picked BigQuery, a customer contract required a specific provider. The goal then is not to make the clouds interchangeable. It is to keep each one well run without tripling the work.

Reasons that hold up, and reasons that don't

Reasons that usually hold up:

  • A managed service that only one provider offers and that the business depends on.
  • Customer, regulatory or data residency requirements.
  • An acquisition where migrating would cost more than operating both.

Reasons that usually don't:

  • "Avoiding lock-in" in general. Writing everything to the lowest common denominator costs more than a migration you might never do.
  • Active-active across providers for availability. Replicating data between clouds with acceptable consistency is hard, and most outages are caused by your own changes, not by a provider going down.

What multiplies

Every additional provider adds its own:

  • Identity model (IAM roles and policies, GCP IAM bindings, Entra ID and Azure RBAC).
  • Networking model and limits.
  • Managed services with similar names and different behavior.
  • Billing format and discount programs.
  • Console, CLI, quotas, support process, and things the on-call engineer needs to know at 3 am.

The practices below are about keeping that multiplication under control.

Keep a request path inside one cloud

Pick a primary cloud per workload. Avoid designs where a single user request crosses providers several times. Each crossing adds internet or interconnect latency and egress charges, and egress from every major cloud is billed per GB. Put services that talk to each other a lot in the same cloud, and integrate across clouds asynchronously (queues, batch exports) where you can.

One IaC tool, separate state per provider

Terraform or OpenTofu work well across all three. Don't put everything into one root module. Split state by provider, environment and component, so a GCP change can't block an AWS apply, and so the blast radius stays small.

infra/
  aws/prod/network/
  aws/prod/eks/
  gcp/prod/network/
  gcp/prod/bigquery/
  azure/prod/identity/
  modules/

Use one naming and tagging scheme everywhere. GCP labels only allow lowercase letters, digits, underscores and dashes, so design the scheme around the strictest provider:

locals {
  common_tags = {
    team        = "platform"
    environment = "prod"
    managed-by  = "terraform"
  }
}

provider "aws" {
  region = "eu-central-1"
  default_tags {
    tags = local.common_tags
  }
}

provider "google" {
  project        = "my-gcp-project"
  region         = "europe-west3"
  default_labels = local.common_tags
}

For Azure, pass the same map to the tags argument of each resource or through your modules.

No long-lived keys between clouds

  • CI to cloud. Use OIDC federation from your CI system. GitHub Actions and GitLab CI can authenticate to AWS (IAM OIDC provider), GCP (Workload Identity Federation) and Azure (federated credentials) without stored secrets.
  • Kubernetes to cloud. EKS Pod Identity or IRSA on EKS, Workload Identity Federation on GKE, Microsoft Entra Workload ID on AKS.
  • Cloud to cloud. GCP Workload Identity Federation can trust AWS roles directly. Prefer federation over service account keys copied into another provider's secret store.

Plan the network once

  • Allocate non-overlapping CIDR ranges for every VPC and VNet across all providers and on-premises before you build anything. Renumbering later is painful.
  • Connect with site-to-site IPsec VPN using BGP (AWS Site-to-Site VPN, GCP HA VPN, Azure VPN Gateway) for moderate traffic. Use dedicated links such as GCP Cross-Cloud Interconnect when traffic is large and steady.
  • Run DNS so that each cloud can resolve the others' private zones through forwarding rules, and document which resolver owns which zone.

Kubernetes helps, but doesn't erase differences

Running Kubernetes on all three gives you one deployment model. Load balancer annotations, storage classes, ingress controllers, node images and identity integration still differ. Keep a shared base for manifests and a thin per-cloud overlay with Kustomize or Helm values. Don't pretend the overlays aren't there.

One place for telemetry and cost

  • Observability. Run OpenTelemetry collectors in each environment and send metrics, logs and traces to a single backend. On-call should not need three consoles to follow one incident.
  • Cost. All three providers can export billing data in the FOCUS format (FinOps Open Cost and Usage Specification). Load the exports into one warehouse and report by your shared tags.

Ownership and runbooks

Each cloud needs a named owning team, an account or project structure, a break-glass procedure and runbooks for the common failures. Keep provider-specific knowledge written down, because the person who knows Azure networking will be on holiday when it breaks.

Checklist

  • Write down why each provider is used. If there is no reason, plan to consolidate.
  • Keep request paths inside one cloud, integrate across clouds asynchronously.
  • Separate state per provider and environment, one tagging scheme that fits GCP label rules.
  • Federated identity everywhere, no long-lived keys.
  • Non-overlapping address plan, BGP-based connectivity, documented DNS forwarding.
  • One observability backend and one cost dataset.
  • A named owner and runbooks per provider.