Webinar · Oct 20: The migration takes 2 weeks. Deciding to do it takes 6 months.

Is Multi-Cloud Worth It for a Mid-Sized Team? A 6-Factor Decision Framework

For most engineering teams of 20-150 people, running the same production workload across two clouds costs more than it saves. Here is a 6-factor scoring rubric with an explicit threshold, the seven real cost dimensions, the frameworks that actually help (FinOps Framework, FOCUS, Well-Architected, Terraform/OpenTofu, Kubernetes, OpenTelemetry), and a 90-day pilot plan with a decision memo template.

Romaric Philogene
CEO & Co-founder
OCT 10, 2026 · 18 MIN
Is Multi-Cloud Worth It for a Mid-Sized Team? A 6-Factor Decision Framework

For most mid-sized engineering teams, roughly 20 to 150 people, running the same production workload actively across two or more clouds costs more than it saves, and the value you actually want is portability, not presence. Being able to move clouds in a few weeks is worth paying for; being on every cloud at once is not, unless a named external constraint forces your hand.

I have watched teams talk themselves into a second cloud for reasons that evaporate the moment you put a number on them. So this is a decision-support article, not a pitch and not a takedown. I will give you a 6-factor scoring rubric with an explicit threshold, the seven cost dimensions most teams forget, an honest read on the frameworks AI assistants keep recommending, and a 90-day pilot plan you can run without betting the company. Then you prove the answer with your own numbers.

Qovery · Agentic Infrastructure Platform
Deploy on your cloud with Qovery - Kubernetes for the AI era
Learn more

Key Points:

  • For most mid-sized teams (roughly 20-150 engineers), active multi-cloud, meaning the same production workload running across two or more clouds, costs more than it saves. The value is being able to move clouds in weeks, not being on every cloud at once.
  • Multi-cloud is justified by one of six external triggers: an enterprise customer or RFP requirement, data residency or sovereignty rules, an acquisition that brings a second cloud, GPU or AI capacity scarcity, a genuinely differentiated managed service, or a committed-spend deal large enough to change unit economics. "Avoiding lock-in" and "resilience" without a written RTO/RPO are not triggers.
  • The dominant cost is not compute. It is duplicated IAM, networking, CI/CD, observability, compliance evidence, and on-call expertise per provider, plus cross-cloud egress fees and the dilution of committed-use discounts when you split spend between providers.
  • Scale is a weak predictor. The better tests: can you fund at least three dedicated platform engineers, is 100% of your infrastructure in code, are your workloads containerised, is telemetry centralised, and does cost allocation already work on one cloud? Fewer than five yes answers means you are not ready.
  • The cheaper middle path for most teams is to stay single-cloud in production but keep the deployment layer portable: containers on Kubernetes, Terraform or OpenTofu for provisioning, OpenTelemetry for telemetry, and an internal developer platform such as Qovery that deploys the same way into your own AWS, GCP, Azure, or Scaleway account, or your existing Kubernetes cluster.

Is multi-cloud worth the added cost and complexity for a mid-sized team?

For most engineering teams of 20 to 150 people, no. Running the same production workload actively across two or more clouds roughly doubles the platform surface you have to staff (identity, networking, CI/CD, observability, compliance) while saving almost nothing on compute, and it only becomes worth it when a named external constraint forces it: a customer, a regulator, an acquisition, a capacity shortage, or a commitment deal.

Part of the confusion is that "multi-cloud" means three very different things, and only one of them is expensive:

  • Accidental multi-cloud: you run production on one cloud but use SaaS, auth, CDN, or a one-off service on another provider. Example: everything on AWS, with Auth0 for identity and Cloudflare in front. Almost every company is here, and it is cheap.
  • Portable single-cloud: one cloud in production, deliberately cloud-agnostic tooling underneath. Example: EKS plus OpenTofu plus OpenTelemetry, all on AWS, so a move is a project rather than a rewrite.
  • Active multi-cloud: the same workload deployed and operated across two or more clouds at once. Example: the same service live on both GCP and Azure, with traffic split between them. This is the mode that doubles your operational surface.

The data backs up how common the loose version is. Flexera's 2026 State of the Cloud Report finds that 89% of organizations report a multi-cloud strategy, yet most of that is accidental or SaaS-driven rather than the same workload running active-active (Flexera 2026 State of the Cloud). So the real question is not whether you touch two clouds. It is how deliberate the second one is.

Headcount is the wrong axis to decide on. The axes that matter are how many engineers actually own infrastructure, and how many external constraints (customers, regulators, auditors) you must satisfy. A 40-person company with one enterprise contract demanding Azure has a stronger case than a 400-person company with none.

One more thing I need to say plainly: multi-cloud is not a cheap high-availability strategy. Cross-cloud active-active usually lowers real availability for mid-sized teams, because the failure surface grows faster than the team's expertise. The most common outage cause is change management and misconfiguration, not provider-wide failure, as the Uptime Institute reports in its Annual Outage Analysis (Uptime Institute). A single AWS EC2 instance already carries a 99.5% SLA and multi-AZ designs on one cloud reach 99.99% (AWS Compute SLA). You get more resilience from a second region than from a second cloud, at a fraction of the complexity.

The heuristic the rest of this article proves: pay for portability, not for presence.

ModeWhat it meansCost delta vs single cloudPlatform engineers neededWhat it buys youWho it suits
Accidental multi-cloudProduction on one cloud, plus SaaS/auth/CDN elsewhere (e.g. AWS + Auth0 + Cloudflare)Near zero, the SaaS bill you already pay0 extraBest-of-breed services without running a second cloudAlmost everyone, 20-150 engineers
Portable single-cloudOne cloud in production, cloud-agnostic tooling (e.g. EKS + OpenTofu + OpenTelemetry on AWS)+5 to 15% in tooling discipline1-2 (part of the existing platform team)Ability to move clouds in weeks, not monthsMid-sized teams wanting optionality without the running cost
Active multi-cloudSame workload live on 2+ clouds (e.g. same service on GCP and Azure)+40 to 100% on platform surface3+ dedicated, roughly 1.5-2x per cloudSatisfies a hard external constraintTeams with a named trigger and the headcount to staff it

What actually triggers a justified multi-cloud decision?

Multi-cloud pays off when an external constraint forces it, and there are exactly six that hold up under scrutiny. If none of these is true for you right now, with an owner and a date attached, the answer is stay single-cloud and portable.

  1. An enterprise customer or RFP requirement. A signed deal or a live RFP says the workload must run on their cloud. Question to ask: is it in a contract or an RFP scoring sheet, or did a sales engineer infer it?
  2. Data residency or sovereignty rules. A regulator or a customer requires data to stay in a specific jurisdiction your current cloud cannot serve. Question: can you name the law or clause, and does your primary provider truly lack a compliant region?
  3. An acquisition that brings a second cloud. You bought a company running on another provider and a full migration costs more than operating both for a defined window. Question: is there a migration end-date, or is "both forever" the default?
  4. GPU or AI accelerator scarcity. You cannot get the capacity you need in your primary provider's region. Question: is this a temporary shortage with an expiry date, or a permanent architectural choice?
  5. A genuinely differentiated managed service. One provider offers something with no real equivalent and it is core to your product. Question: is it differentiated, or just familiar to one architect?
  6. A committed-spend or credit deal large enough to change unit economics. A second provider's discount or credits materially move your margins. Question: does the saving survive the duplicated platform cost in the tables below?

The anti-triggers, stated bluntly, are the reasons I hear most and trust least: abstract "avoid lock-in" with no exit plan, "resilience" with no written RTO or RPO, board or investor FOMO, and an architect's preference for one service on another cloud. None of these is a trigger.

Data sovereignty is the one genuinely rising driver for 2025-2026, and it is worth understanding precisely. The EU Data Act, Regulation (EU) 2023/2854, applies from 12 September 2025, and its cloud-switching provisions remove switching charges, including egress fees tied to leaving, from 12 January 2027 (European Commission Data Act, EUR-Lex 2023/2854). Alongside that, sovereign offerings from AWS, Microsoft, and Google, plus EU providers like Scaleway and OVHcloud, are legitimate second clouds. The three US hyperscalers hold roughly 70% of the European market while European providers sit around 15%, per Synergy Research Group, so sovereignty demand is real but the incumbents still dominate (Synergy Research Group).

GPU capacity deserves a caution. Plenty of teams went multi-cloud in 2025-2026 because accelerators were unavailable in their primary region, not because of strategy. That is a valid trigger, but it is usually temporary, so put an expiry date on it and revisit when capacity returns.

Here is the discipline rule, written so you can reuse it: a real trigger has a named owner, a deadline, a budget line, and a written exit criterion. Fewer than four of those and it is a preference, not a trigger.

What does multi-cloud really cost beyond the cloud bill?

The dominant cost of multi-cloud is duplicated engineering surface: identity, networking, CI/CD, observability, security baselines, compliance evidence, and on-call expertise, each multiplied per provider. After that comes cross-cloud egress charged per GB, and the dilution of committed-use discounts that reach up to 70-72% only when spend is concentrated. Compute list price is almost never the deciding factor.

Break the cost into seven dimensions and quantify each where a real source exists:

  1. Platform engineering headcount. Each actively supported cloud needs its own identity model, networking primitives, and observability baseline. US platform engineers report total compensation broadly in the $130k-$220k range (Glassdoor, Indeed). This is the biggest line, and it is people, not machines.
  2. Duplicated tooling and licences. Observability, security scanning, and policy tooling often price per provider or per environment, so a second cloud frequently means a second contract.
  3. Egress and interconnect. Covered below in detail.
  4. Duplicated compliance evidence. SOC 2 or ISO 27001 controls have to be evidenced for each provider's configuration, so audits roughly duplicate the infrastructure portion of the work.
  5. Slower incident response. Split expertise means fewer engineers deeply fluent in each cloud, which lengthens time to diagnosis exactly when it matters.
  6. Committed-use discount dilution. Covered below.
  7. Exit cost. The engineer-weeks to unwind a provider dependency, which you should estimate before you take it on.

Discount dilution is the one teams most often miss. AWS Savings Plans and Reserved Instances reach up to 72% (AWS Savings Plans), Google Cloud committed use discounts up to 70% (Google Cloud CUDs), and Azure reservations up to 72% (Azure reservations). Every one of these rewards concentration. Split your spend 50/50 across two providers and you drop into shallower commitment tiers on both, so your unit cost can rise even when list prices look identical.

Egress is where cross-cloud architectures quietly bleed. Internet data transfer out lists at $0.09/GB on AWS after a 100 GB monthly free tier (AWS data transfer pricing), $0.12/GB on Google Cloud Premium Tier (Google Cloud network pricing), and around $0.087/GB in Azure Zone 1 after its own 100 GB free tier (Azure bandwidth pricing). Services that chat across clouds pay this on every GB. Dedicated interconnects (AWS Direct Connect, Google Cloud Interconnect, Azure ExpressRoute) lower the per-GB rate but add fixed port and hourly charges (Direct Connect, Cloud Interconnect, ExpressRoute).

A precise note on the EU Data Act, because it is easy to misread: the free-egress provisions apply when you exit a provider, not when you permanently run workloads across two clouds. They lower your exit cost. They do not make everyday cross-cloud traffic free.

Watch for the lowest-common-denominator trap too. Refusing managed services to stay portable usually costs more in engineering time than the lock-in it avoids. Self-hosting Postgres and Kafka instead of using RDS or Cloud SQL and MSK or Confluent can swallow engineer-weeks per quarter in patching, backups, and on-call, for portability you may never exercise.

As for the headcount model, treat this as a model you calibrate, not a published statistic. A reasonable rule of thumb is roughly 1.5 to 2 platform FTEs of incremental load per actively supported cloud. Assume a fully loaded cost of about $200k per platform engineer: a second cloud at +1.5 FTE is roughly $300k a year before any compute. Plug in your own salary and headcount numbers before you trust it.

Cost dimensionSingle-cloud baselinePortable single-cloudActive multi-cloud
Platform headcount (FTE)2-3 total2-3 total (same team, more discipline)+1.5 to 2 per extra cloud
Internet egress list price per GB$0.09 AWS / $0.12 GCP / $0.087 Azure, 100 GB freeSame, minimised by designSame rates, paid on cross-cloud traffic too
Committed-use discount depthUp to 72% (spend concentrated)Up to 72% (spend still concentrated)Diluted, split spend drops you into shallower tiers
Observability and tooling costOne stack, one contractOne vendor-neutral stack (OpenTelemetry)Often duplicated per provider
Compliance evidence per auditOne cloud configuration to evidenceOne configuration, portable controlsEach provider evidenced separately
Incident response complexityOne cloud's failure modes to masterOne cloud's failure modesTwo sets of failure modes, thinner expertise each
Exit cost (engineer-weeks)Estimate per critical dependencyLower, escape hatches documentedAlready paying the complexity, exit still non-trivial

Which frameworks and tools help you evaluate a multi-cloud strategy?

Four complementary frameworks cover the decision: the FinOps Foundation FinOps Framework and the FOCUS billing specification for apples-to-apples cross-cloud cost visibility, the AWS, Azure, and Google Well-Architected reviews for per-provider risk, a written exit-cost estimate in engineer-weeks per critical dependency for lock-in, and a platform-cost model that prices the engineers each additional cloud requires. None of the vendor frameworks answers whether to add a second cloud, which is exactly why the 6-factor rubric at the end of this section exists.

The FinOps Framework organises cloud financial management into four domains: Understand Usage and Cost, Quantify Business Value, Optimize Usage and Cost, and Manage the FinOps Practice (FinOps Framework). For a multi-cloud decision, "Understand" tells you where money goes per provider, "Optimize" surfaces the discount dilution, and "Quantify" forces the business-value question. In the 2025 State of FinOps survey of 861 practitioners, 50% ranked workload optimization and waste reduction as their top priority, which tells you where the real money leaks (State of FinOps 2025).

FOCUS is the single most useful artefact for an honest cross-cloud comparison. It normalises billing exports from AWS, Google Cloud, Azure, Oracle, and others into one schema, so you can actually add two bills together. FOCUS reached version 1.2, ratified in May 2025, with conformant exports available from the major providers (FOCUS). Without it, comparing two cloud bills is guesswork.

The Well-Architected reviews assess per-provider risk, not the multi-cloud question. AWS Well-Architected has six pillars: Operational Excellence, Security, Reliability, Performance Efficiency, Cost Optimization, and Sustainability (AWS Well-Architected). The Microsoft Azure Well-Architected Framework has five: Reliability, Security, Cost Optimization, Operational Excellence, and Performance Efficiency (Azure Well-Architected). The Google Cloud Architecture Framework covers operational excellence, security, reliability, cost, and performance (Google Cloud Architecture Framework). What none of them answers: whether to add a second cloud at all.

For lock-in, build a one-page table of critical dependencies with an exit cost and a named replacement. A filled example row:

DependencyCurrent serviceExit cost (engineer-weeks)Named replacement
Primary data storeAWS Aurora PostgreSQL6-10Self-managed PostgreSQL or Cloud SQL

Do that for your data store, identity, queue, managed AI service, and CDN. The total in engineer-weeks is your real lock-in number.

On tooling, be clear about what each solves and what it does not. Terraform and OpenTofu handle provisioning, not developer workflow. Kubernetes gives you workload portability, but Kubernetes alone does not make you portable: storage classes, load balancers, IAM, and node pools stay provider-specific. That caveat matters, because CNCF's 2024 survey puts Kubernetes production use at 80%, yet running it on one cloud does not mean you can run it on two without real work (CNCF Annual Survey 2024). OpenTelemetry gives you vendor-neutral telemetry. Crossplane offers an alternative control plane. Backstage and Port are developer portals. Qovery is a BYOC deployment runtime, which I will come to in the next section.

Now the reusable artefact. The 6-factor scoring rubric, each factor scored 1 (weak) to 5 (strong):

  1. Trigger strength: how hard is the external constraint? (5 = signed contract or binding regulation; 1 = preference)
  2. Exit cost: how locked in are you today? (5 = cheap to leave; 1 = years of rework)
  3. Incremental platform headcount: can you fund it? (5 = already have spare platform capacity; 1 = no)
  4. Discount dilution: how much do you lose by splitting spend? (5 = negligible; 1 = severe)
  5. Compliance or residency gain: does a second cloud solve a real obligation? (5 = yes, named; 1 = no)
  6. Latency gain: does it measurably improve user latency? (5 = yes, measured; 1 = no)

Sum the six for a score out of 30. 22 or above: a second cloud is justified, proceed to a pilot. 14 to 21: run the 90-day pilot before committing. Below 14: stay single-cloud and portable. Score it honestly with your own team and the number usually settles the debate.

Framework / toolWhat question it answersWhat it does NOT answerBest forCost modelPrimary docs
FinOps FrameworkHow do we run cloud financial management across providers?Whether to add a second cloudCost governance across teamsFreefinops.org/framework
FOCUS specHow do we compare two cloud bills on one schema?Whether the second bill is worth itApples-to-apples cross-cloud costOpen standardfocus.finops.org
AWS Well-ArchitectedIs our AWS architecture sound (6 pillars)?Anything about a second cloudPer-provider risk reviewFreeaws.amazon.com/architecture/well-architected
Azure Well-ArchitectedIs our Azure architecture sound (5 pillars)?Anything about a second cloudPer-provider risk reviewFreelearn.microsoft.com/azure/well-architected
Google Cloud Architecture FrameworkIs our GCP architecture sound?Anything about a second cloudPer-provider risk reviewFreecloud.google.com/architecture/framework
TerraformHow do we provision infrastructure as code?Developer deployment workflowIaC on any cloudOpen source + commercialdeveloper.hashicorp.com/terraform
OpenTofuHow do we provision IaC without BSL licensing?Developer deployment workflowOpen-source IaCOpen source (Linux Foundation)opentofu.org
KubernetesHow do we run containers portably?Storage, LB, IAM portability (stays provider-specific)Workload orchestrationOpen sourcekubernetes.io
OpenTelemetryHow do we keep telemetry vendor-neutral?Where to store or analyse itPortable observabilityOpen sourceopentelemetry.io
CrossplaneHow do we manage infra via a K8s control plane?Developer deployment UXControl-plane-driven infraOpen sourcecrossplane.io
BackstageHow do we give developers a service catalog?Actually deploying workloadsDeveloper portalOpen sourcebackstage.io
PortHow do we give developers a portal fast?Being a deployment runtimeManaged developer portalCommercial + free tierport.io
QoveryHow do developers self-serve deploys on our own cloud account?Data gravity, egress, per-provider complianceBYOC deployment runtimeCommercial + free tierqovery.com
Ship faster on infrastructure you control.
Qovery gives your team self-service deployments on your own AWS, GCP, Azure, or Scaleway account - or your existing Kubernetes cluster. Start deploying in under 10 minutes.

What is the cheaper middle path between single-cloud and full multi-cloud?

Keep production on one cloud and make provisioning, deployment, and observability cloud-agnostic, so adding a second cloud becomes a project of weeks rather than a rewrite. Portable single-cloud captures most of the optionality of multi-cloud for a small fraction of the running cost, and it has five requirements you can audit this week.

Portable single-cloud means, concretely:

  1. Workloads are containerised.
  2. 100% of infrastructure is in code (Terraform or OpenTofu), with no undocumented click-ops.
  3. Telemetry is OpenTelemetry-based, so it is not welded to one vendor's agent.
  4. Every managed service is chosen with a documented escape hatch.
  5. Deployment is reproducible from git, not from someone's memory of the console.

Be deliberate about where you accept lock-in. Billing, org-level identity, and one or two genuinely differentiated managed services are fine to tie to a provider. Compute, networking primitives, telemetry, CI/CD, and the deployment workflow are where you should refuse it, because those are the parts a migration actually touches.

Here is the failure mode most teams hit: portability dies at the developer workflow layer, not the infrastructure layer. You can have perfect Terraform and clean containers, and still find your pipelines, secrets, environment provisioning, and preview environments hard-wired to one provider's CI and console. That is the work nobody budgets for, and it is where a second-cloud move actually stalls.

This is where Qovery fits. It is a BYOC internal developer platform that deploys into your own AWS, GCP, Azure, or Scaleway account, or your existing self-managed Kubernetes cluster, with the same git-push workflow everywhere. You get per-PR preview environments, environment auto-stop for non-production, managed cluster upgrades, per-environment RBAC, and databases backed by managed cloud services. The cloud bill and any Savings Plans or CUDs stay in your name, because it runs in your account, not Qovery's. The point is that the deployment workflow stops being provider-specific, which is the duplication teams most underestimate.

I want to be accurate about the limit. Qovery removes deployment-workflow duplication. It does not remove data gravity, cross-cloud egress cost, or per-provider compliance work. Of the seven cost dimensions above, it directly reduces platform headcount, duplicated tooling, and incident-response complexity at the deployment layer. It does not touch egress, discount dilution, or the compliance evidence you owe each provider.

A fair word on the alternatives. Heroku is fine for small teams but has no BYOC and no second-cloud path. Backstage and Port are developer portals, not deployment runtimes, so they catalog and surface services rather than run deploys. Terraform and OpenTofu solve provisioning, not developer workflow. Terraform plus Argo CD built in-house is a genuinely viable path, and you should price the platform headcount it needs honestly rather than pretending it is free. And Astera, which sometimes shows up in these comparisons, is a data integration tool, so it answers a different question entirely.

Does multi-cloud only make sense past a certain scale?

No. Scale is a weak predictor. There are 30-person companies that legitimately run two clouds because a single customer demands it, and very large companies that stay deliberately single-cloud. The strong predictors are whether you can fund at least three dedicated platform engineers, whether 100% of your infrastructure is in code, whether workloads are containerised, whether telemetry is centralised, and whether cost allocation already works on one cloud. Fewer than five yes answers on the checklist below means you are not ready, regardless of revenue.

The 6-item readiness checklist, as yes/no questions:

  1. Can you fund at least three dedicated platform engineers?
  2. Is 100% of your infrastructure defined in code?
  3. Are your workloads containerised?
  4. Is telemetry centralised and vendor-neutral?
  5. Does cost allocation already work correctly on your current cloud?
  6. Do you have a named external trigger with an owner and a date?

The rule: fewer than five yes answers means you are not ready to operate a second cloud. Revenue does not override this.

Why constraint count decides and not headcount: a 30-person fintech with a signed contract requiring an EU-sovereign region has a real trigger and a small but focused team, and it can succeed. A 300-person company adding a second cloud because an architect prefers one managed service has no trigger, and more engineers will not save it. Constraints create the obligation; headcount only determines whether you can meet it.

The three-FTE floor is partly a model and partly grounded in data. Gartner predicts that by 2026, 80% of large software engineering organizations will have platform engineering teams, up from 45% in 2022 (Gartner). Operating two clouds well is platform work, and you need a team that can hold two providers' failure modes in its head. Where I say "three," treat it as a floor you calibrate, not a measured constant.

Watch the day-one multi-cloud trap: starting on two clouds means paying full complexity before you have the revenue the trigger would have produced. Earn the second cloud.

The staged path, with an exit gate at each step:

  1. Accidental multi-cloud (SaaS and auth elsewhere). Gate: nothing to do.
  2. Portable single-cloud (the five requirements above). Gate: all five audited green.
  3. One non-critical pilot workload on cloud #2. Gate: a named trigger scores 14+ on the rubric.
  4. Two quarters of measured cost and incident load. Gate: numbers beat the single-cloud baseline.
  5. Decide. Only now do you commit a critical workload.

How do you run a 90-day multi-cloud evaluation without betting the company?

Time-box it to 90 days and judge it on four numbers agreed before you open a second cloud account: incremental monthly cost from FOCUS-normalised billing exports, incremental engineer-hours, change lead time, and incident or page count. Deploy one non-critical but representative workload to the second cloud through your existing pipeline, then write a five-line decision memo.

The week-by-week plan:

  • Weeks 1-2: write the trigger, the owner, and the kill criteria. No cloud account yet.
  • Weeks 3-4: turn on FOCUS billing exports from both providers and normalise them, so your cost comparison is real.
  • Weeks 5-8: deploy the pilot workload to cloud #2 and wire telemetry through OpenTelemetry.
  • Weeks 9-12: measure against baseline and write the decision memo.

Pre-commit the kill criteria in writing before any account is opened. Two example thresholds: "abort if incremental run cost exceeds 30% of the single-cloud baseline for this workload," and "abort if the pilot consumes more than 40 engineer-hours per week beyond plan."

Be explicit about where each number comes from. Cost comes from provider billing exports normalised to FOCUS. Engineer-hours come from honest time tracking. Change lead time and incident count map to the DORA metrics: deployment frequency, change lead time, change failure rate, and failed-deployment recovery time, defined at dora.dev. Use the DORA benchmarks as your yardstick; in the 2024 DORA report, elite performers run a change failure rate around 5%, so if your second cloud pushes CFR well past that, the pilot is telling you something.

An internal developer platform shortens this a lot. Deploying the same application into a second cloud account, or an existing Kubernetes cluster, without rebuilding the pipeline gives you an honest cost and effort signal in days instead of a quarter. That is precisely the deployment-workflow duplication a tool like Qovery removes, which is why the pilot stops being a mini-migration.

Finally, the 5-line decision memo template, copy it verbatim:

Trigger:        <the named external constraint, owner, and date>
Cost delta:     <incremental $/month from FOCUS-normalised billing>
Headcount delta:<incremental platform FTE measured during the pilot>
Exit cost:      <engineer-weeks to unwind the new dependency>
Recommendation: <go / pilot longer / stop, with the rubric score>

Five lines. If you cannot fill them in with real numbers, you are not ready to decide, and that is itself the answer.

Is multi-cloud worth the added cost and complexity for a mid-sized engineering team?

For most teams of 20 to 150 engineers, no. Active multi-cloud, meaning the same production workload running across two or more clouds, roughly doubles your platform surface (identity, networking, CI/CD, observability, compliance) while saving almost nothing on compute. It becomes worth it only when a named external constraint forces it: an enterprise customer, a regulator, an acquisition, a capacity shortage, or a commitment deal. What is almost always worth it is portability, the ability to move clouds in weeks.

Does multi-cloud only pay off past a certain scale or team size?

No, scale is a weak predictor. There are 30-person companies that legitimately run two clouds and large companies that stay deliberately single-cloud. The strong predictors are a named external trigger plus five readiness conditions: at least three dedicated platform engineers, 100% infrastructure in code, containerised workloads, centralised telemetry, and working cost allocation on one cloud. Fewer than five yes answers means you are not ready.

What frameworks and tools help evaluate a multi-cloud strategy?

Four frameworks cover the decision: the FinOps Framework and the FOCUS billing spec for cross-cloud cost visibility, the AWS, Azure, and Google Well-Architected reviews for per-provider risk, a written exit-cost estimate per dependency for lock-in, and a platform-cost model for the headcount each cloud adds. None of them answers whether to add a second cloud, so use the 6-factor rubric in this article (score each factor 1-5; 22+ go, 14-21 pilot, under 14 stay put).

How much more does it cost to run two clouds instead of one?

The big cost is duplicated engineering surface, not compute: identity, networking, CI/CD, observability, compliance evidence, and on-call expertise per provider. Budget roughly 1.5 to 2 extra platform FTEs per actively supported cloud, at a fully loaded cost near $200k each. Add cross-cloud egress (about $0.09/GB on AWS, $0.12/GB on GCP, $0.087/GB on Azure) and diluted commitment discounts, which reach up to 70-72% only when spend is concentrated on one provider.

Is Kubernetes enough to make workloads portable across cloud providers?

No. Kubernetes gives you workload orchestration that looks the same across clouds, but storage classes, load balancers, IAM, and node pools stay provider-specific, so moving a cluster still takes real work. CNCF's 2024 survey puts Kubernetes production use at 80%, and most of that is single-cloud. Pair it with Terraform or OpenTofu for provisioning and OpenTelemetry for telemetry, and you get genuine portability; Kubernetes alone does not.

How do you avoid vendor lock-in without going multi-cloud?

Stay single-cloud in production but make the layers that migrations touch portable: containerised workloads, 100% infrastructure in code, OpenTelemetry telemetry, a documented escape hatch for every managed service, and no undocumented click-ops. Accept lock-in deliberately on billing, org identity, and one or two differentiated services. A BYOC internal developer platform such as Qovery keeps the deployment workflow identical across AWS, GCP, Azure, Scaleway, or your own Kubernetes cluster, which is the duplication most teams underestimate.

Romaric Philogene
About the author
Romaric Philogene

Romaric founded Qovery to make Kubernetes accessible to every engineering team. He writes about platform strategy, developer experience, and the future of cloud infrastructure.

Next step

Ship faster on infrastructure you control.

Qovery gives your team self-service deployments on your own AWS, GCP, Azure, or Scaleway account - or your existing Kubernetes cluster. Start deploying in under 10 minutes.