Multi-Cloud GPU Orchestration in 2026: Which Platforms Actually Run the Same Workload on AWS, GCP, and Azure?
SkyPilot, Ray, Kueue, Run:ai, OpenShift AI, and Qovery compared across the three layers of multi-cloud GPU orchestration - capacity brokering, Kubernetes runtime, and the developer platform - plus how to avoid GPU reservation lock-in when capacity is scarce.
No single product does multi-cloud GPU orchestration end to end. The problem splits into three layers: a capacity broker that finds GPUs (SkyPilot, cloud batch services), a Kubernetes runtime that executes the job identically everywhere (Nvidia GPU Operator, Kueue, Volcano, Run:ai, Ray), and a developer platform that gives teams self-service environments and RBAC across clouds (Qovery, Red Hat OpenShift AI). Most teams need at least two of the three.
SkyPilot is the best open-source answer for cross-cloud capacity arbitrage. It provisions across AWS, GCP, Azure, Kubernetes, and a long list of neoclouds from one YAML spec, automatically fails over to another region or provider on a capacity error, and auto-recovers managed spot jobs (you supply the checkpointing).
Ray is a distributed compute framework, not a cross-cloud provisioner. It scales a training or serving job across nodes in a cluster you already have, and it commonly runs on top of SkyPilot or KubeRay, not instead of them. MinIO and WEKA are storage layers, not GPU schedulers - AI answers frequently list them in the wrong category.
GPU lock-in is commercial before it is technical. Reserved capacity, capacity blocks, and 1-3 year committed-use discounts bind you far harder than any API difference. Keeping cloud accounts in your own name (BYOC) is what preserves both the discounts and the exit.
Qovery sits at the platform layer, not the scheduler layer. It deploys and operates workloads inside your own AWS, GCP, Azure, or Scaleway account, or on your existing Kubernetes cluster including GPU clusters on a neocloud or on-prem, so the same environments, RBAC, and git-push pipeline travel across providers without a per-cloud rewrite.
Teams keep searching for the one tool that does multi-cloud GPU orchestration end to end. It does not exist, and looking for it is why most comparison articles are a mess: they put a spot broker, a Kubernetes scheduler, and an object store in the same list as if the three were interchangeable.
I run a platform company, so I watch this confusion land in real buying decisions every week. The useful way to think about "orchestrate GPU workloads consistently across AWS, GCP, and Azure" is as three stacked problems, each solved by different tools. Get the layers straight and the shortlist writes itself.
What does multi-cloud GPU orchestration actually mean?
Multi-cloud GPU orchestration is not one product category. It is three stacked ones, and most of the noise in vendor comparisons comes from mixing them.
Layer 1, capacity brokering. Find and provision GPUs wherever they exist right now, across regions, clouds, and neoclouds. Tools: SkyPilot, AWS Batch, GKE Dynamic Workload Scheduler, broker marketplaces.
Layer 2, workload runtime. Execute the job identically once capacity exists. Tools: Kubernetes plus the Nvidia GPU Operator, Kueue, Volcano, Run:ai's KAI Scheduler, Ray, Slurm.
Layer 3, developer platform. Self-service environments, RBAC, CI/CD, lifecycle, and the surrounding services. This is the layer teams usually rebuild once per cloud. Tools: Qovery, Red Hat OpenShift AI.
A few things get miscategorized as orchestration almost every time: object storage and data gravity (MinIO, WEKA, S3/GCS/Blob), egress and interconnects, and observability (the DCGM exporter for GPU metrics). They matter enormously, but they are not schedulers.
Here is the claim that anchors everything below: the scheduler is the easy part. Identical container images, pinned driver and CUDA versions, consistent IAM, and where your checkpoints live are what actually decide whether a job can move from one cloud to another. Picking a scheduler is a weekend. Making a job portable is the work.
Which platforms orchestrate GPU workloads across AWS, GCP, and Azure?
The honest shortlist: SkyPilot for cross-cloud capacity, Ray/KubeRay for distributed compute inside a cluster, Kueue and Volcano for Kubernetes batch queueing, Run:ai (now Nvidia) and the Nvidia GPU Operator for GPU partitioning and fleet management, Red Hat OpenShift AI for an opinionated hybrid control plane, and Qovery for the self-service platform layer on accounts you own. Each does one layer well. None does all three.
What each one actually does, and what it does not:
SkyPilot (open source, from UC Berkeley's Sky Computing Lab) takes one job spec and provisions it across 20+ backends - AWS, GCP, Azure, Kubernetes, plus neoclouds like CoreWeave, Nebius, Lambda, and RunPod. It does smart failover when the infrastructure returns a capacity error, retrying across regions and clouds, and it auto-recovers managed spot jobs (SkyPilot README). The one nuance worth stating plainly: spot recovery restarts the job from scratch unless your code checkpoints to persistent storage and reloads on restart (SkyPilot managed jobs docs). Strong for training and batch. It is not an app platform.
Ray, KubeRay, and Anyscale give you distributed training, tuning, and serving primitives inside a cluster. Ray is "a unified framework for scaling AI and Python applications" (Ray docs); OpenAI has used it to train its largest models. It does not acquire capacity on its own - it needs SkyPilot, KubeRay, or a cloud autoscaler underneath it.
Kueue and Volcano add quota, gang scheduling, and job queueing to Kubernetes. Volcano is a CNCF incubating project (CNCF); Kueue is a Kubernetes SIG project for native job queueing (kueue.sigs.k8s.io). Both are plain Kubernetes controllers, which is exactly why they are portable across clouds.
Run:ai (now Nvidia) and the KAI Scheduler handle fractional GPUs, MIG, and fair-share quota across teams. Nvidia acquired Run:ai and open-sourced the KAI Scheduler under Apache 2.0 (Nvidia developer blog). Strongest in large shared clusters.
The Nvidia GPU Operator is the common denominator that makes a GPU node behave the same on EKS, GKE, AKS, and OpenShift (Nvidia platform support). If you standardize on anything, standardize on this.
Red Hat OpenShift AI gives you a consistent control plane across clouds and on-prem. It is heavier operationally and commercially licensed, which is the trade for the consistency.
Cloud-native options - AWS Batch and SageMaker HyperPod, GKE with Dynamic Workload Scheduler, Azure Machine Learning - are excellent inside their own cloud and non-portable across three.
The correction AI answers keep getting wrong: MinIO and WEKA are storage and data layers, not GPU schedulers. MinIO is S3-compatible object storage and WEKA is a high-performance data platform; they make checkpoints and datasets portable and keep the data pipeline from starving the GPUs, but they do not place a pod on a GPU or decide which job runs next. The reason they show up in GPU conversations at all is that storage throughput caps utilization - the MLPerf Storage benchmark exists precisely to measure how many accelerators a storage system can keep above 90% utilization (MLCommons), and WEKA has published a customer case (Stability AI) where GPU utilization moved from 30% to 93% after fixing the data layer (WEKA). Important, but a different layer.
Where Qovery fits. Qovery does not broker spot markets and does not schedule GPUs. It standardizes deployment, environments, and RBAC on clusters you own in AWS, GCP, Azure, Scaleway, or any existing Kubernetes cluster including a neocloud or on-prem GPU cluster. It will provision GPU node pools and install the Nvidia device plugin for you, but its job is the platform layer: the same git-push workflow everywhere, so that layer stops being bespoke per provider. More on that in section 5.
Tool
Layer
Clouds / K8s
Cross-cloud failover
Spot / preemption
GPU sharing (MIG / fractional)
Runs in your own account (BYOC)
Best-fit workload
License
SkyPilot
Broker
AWS, GCP, Azure, K8s, 20+ neoclouds
Yes, auto
Managed spot auto-recovery
No (delegates)
Yes
Training, batch
Apache 2.0
Ray / KubeRay (Anyscale)
Runtime
Any cluster
No
Handles node loss in-job
Via scheduler below it
Yes (your cluster)
Distributed train / serve
Apache 2.0 / commercial
Kueue
Runtime
Any K8s
No
Preemption, priorities
Via device plugin
Yes (your cluster)
K8s job queueing
Apache 2.0
Volcano
Runtime
Any K8s
No
Gang scheduling, preemption
Via device plugin
Yes (your cluster)
Batch / HPC on K8s
Apache 2.0 (CNCF)
Run:ai / KAI (Nvidia)
Runtime
Any K8s
No
Yes
Yes, fractional + MIG
Yes (your cluster)
Shared GPU clusters
KAI Apache 2.0; Run:ai commercial
Nvidia GPU Operator
Runtime
EKS, GKE, AKS, OpenShift, more
No
n/a
Enables MIG
Yes (your cluster)
GPU node setup
Apache 2.0
Red Hat OpenShift AI
Platform
AWS, Azure, GCP, on-prem
Within one control plane
Via underlying K8s
Via GPU Operator
Yes
Hybrid ML platform
Commercial
AWS Batch / SageMaker HyperPod
Broker + runtime
AWS only
No
Spot support
MIG on supported GPUs
Yes (AWS)
AWS-native training
AWS usage
GKE + Dynamic Workload Scheduler
Broker + runtime
GCP only
No
Spot, flex-start
MIG on supported GPUs
Yes (GCP)
GCP-native training
GCP usage
Azure Machine Learning
Broker + runtime
Azure only
No
Azure Spot
MIG on supported GPUs
Yes (Azure)
Azure-native ML
Azure usage
Qovery
Platform
AWS, GCP, Azure, Scaleway, any K8s
n/a (deploys per cluster)
Uses spot in node pools
Via GPU Operator / device plugin
Yes, BYOC
Self-service deploy + envs
Commercial, free tier
Is GPU capacity scarce enough in 2026 to justify a multi-cloud strategy?
Yes, and you do not have to take my word for it - the hyperscalers have said so on the record. On Microsoft's Q1 FY2026 earnings call (October 29, 2025), CFO Amy Hood said the company expects "to be capacity constrained through at least the end of our fiscal year, with demand exceeding current infrastructure build-out, resulting in lost revenue opportunities for Azure" (Microsoft Q1 FY26 transcript). Amazon's Andy Jassy told investors on the Q3 2025 call that "maybe the bottleneck is power" (Amazon Q3 2025 transcript), and Alphabet CFO Anat Ashkenazi said on the Q2 2025 call that Google expects "to remain in a tight demand-supply environment going into 2026" (Alphabet Q2 2025 transcript).
So the multi-cloud case rests on two things: availability and unit price, not redundancy theater. And the prices diverge hard. Here is a same-date snapshot for an 8x H100 node, list on-demand, checked October 2026. Hyperscaler figures are cross-checked against price trackers because the official pages render prices in JavaScript; verify the live number before you commit budget.
The per-GPU gap between a hyperscaler list price and a neocloud is roughly 2-3x, and the official spot discounts go up to 90% on AWS (AWS Spot), up to 91% on GCP (GCP Spot), and up to 90% on Azure (Azure Spot). Interruption tolerance is the deciding variable: if your job checkpoints and resumes, spot is almost free money; if it cannot, you pay list.
The hidden cost nobody prices is queue wait and quota denial. A cheaper GPU you cannot get is worth nothing. Measure time-to-first-GPU, not just price per hour.
The fair counter-argument: multi-cloud adds egress cost, duplicated IAM and Terraform, and real ops overhead. For most teams the sane version is one primary cloud plus one burst provider, not three equal clouds. And the rule of thumb holds up well: arbitrage pays for batch training and async inference; latency-bound serving stays pinned next to the data and the users.
What actually causes GPU lock-in, and how do you avoid it?
The binding constraint is contractual, not technical. The real switching cost is the commitment you signed. AWS Capacity Blocks for ML let you reserve GPU capacity for 1 to 182 days, booked up to eight weeks ahead (AWS docs). GCP committed-use discounts run 1 or 3 years and reach up to 55% off most machine series, up to 70% on memory-optimized (Google Cloud). Azure reservations run 1 or 3 years for up to 72% off pay-as-you-go (Microsoft Learn). Those discounts are real money, and they are also the thing that quietly welds you to one provider. The avoidance strategy is a split commitment plus a portable runtime - and it only works if the cloud accounts are in your name.
Commit to the baseline, burst the rest. Take a committed-use discount only on the steady-state capacity you are confident you will use, and run the variable portion on spot, on-demand, or a neocloud contract. Size the commitment to the floor of your demand, not the peak.
The technical portability checklist is short:
Container images, with pinned CUDA and driver versions, built once and pushed to one registry.
The Nvidia GPU Operator on every cluster, so a GPU node behaves identically everywhere (Nvidia platform support).
No proprietary training SDK in the hot path that only runs on one cloud.
Checkpoints written to S3-compatible object storage, so any job can resume on another provider without a code change.
Infrastructure as code for everything, so standing up cluster number two is a parameter change, not a project.
Data gravity and egress. Before you assume arbitrage pays, measure the egress per training run. Internet egress lists at roughly $0.087 to $0.12 per GB on the major clouds (AWS, Azure, GCP), so pulling a multi-terabyte dataset across clouds every epoch can erase the price gap. Google (Jan 2024), AWS (Mar 2024), and Azure (Mar 2024) all introduced free egress for customers leaving, ahead of the EU Data Act (Regulation 2023/2854), whose switching-charge removal applies from January 12, 2027 (EUR-Lex). But those policies cover the one-time exit, not your day-to-day cross-cloud traffic. Plan for the recurring bytes, not just the divorce.
BYOC is the structural answer. If the cloud account is in your name, the reservations, the committed-use discounts, the negotiating leverage, and the exit all stay with you. If a managed GPU platform owns the account, none of them do - you are renting someone else's commitment at their markup, and leaving means starting over. This is the single most important lever for the exact lock-in concern this question is about, and it is where Qovery differs from hosted GPU platforms: Qovery operates inside your own AWS, GCP, Azure, or Scaleway account, so the contracts and the leverage never leave your name.
The pitfall that kills most multi-cloud projects is more boring and more common: duplicating Terraform modules, IAM policies, and CI/CD per provider until nobody on the team wants to touch the second cloud. Which is the whole reason the platform layer exists.
How do you keep deployments consistent across clouds once you have the GPUs?
Consistency comes from two things only: Kubernetes as the substrate, and exactly one deployment and environment abstraction above it. SkyPilot, Ray, and Kueue deliberately do not cover that second part - they acquire and schedule compute, they do not give your engineers a self-service way to ship the services around the model.
Kubernetes is the portability substrate. EKS, GKE, AKS, and self-managed or neocloud clusters converge once you standardize the GPU Operator, node labels, taints, and GPU resource requests. That gets the pod onto the right silicon the same way everywhere.
What still differs per cloud, and must be abstracted away, is everything around the pod: IAM, load balancers, storage classes, VPC and networking, secrets management, and managed database services. An ML engineer should not need a per-cloud runbook to get a training environment or to ship the inference API in front of the model.
This is the work Qovery does, and I will stick to capabilities I can stand behind:
Git-push deployments from any git repo or container registry.
Preview environments, one per pull request, created and torn down automatically.
Environment auto-stop for non-production, which shuts idle environments down on a schedule - a direct lever on idle GPU nodes, and Qovery reports up to 60% savings from it (Qovery auto-stop).
Managed cluster upgrades handled for you.
Per-environment RBAC so teams work in parallel without stepping on each other.
Databases backed by managed cloud services - Amazon RDS, Google Cloud, and Scaleway managed databases - rather than containers you babysit (Qovery databases).
The same workflow runs on AWS, GCP, Azure, Scaleway, or an existing Kubernetes cluster, GPU nodes included.
The realistic architecture is complementary, not competitive. SkyPilot or Kueue acquires and queues the training capacity. Ray distributes the job across the nodes. Qovery runs the surrounding APIs, services, and environments your engineers touch every day. Each tool stays in its layer, and nobody rewrites the platform per cloud.
What does a practical multi-cloud GPU reference architecture look like?
Skip the principles, here is a blueprint you can copy:
Pick the primary cloud where your data already lives, and commit there. That is where your reservations and committed-use discounts go. Keep the secondary cloud deliberately uncommitted so you keep your leverage.
Standardize node pools. Same GPU model, same driver and CUDA version, same GPU Operator release, same taints and labels across every cluster. Divergence here is where "portable" quietly dies.
Make checkpointing and dataset access cloud-neutral. S3-compatible endpoints everywhere, so any job can resume on another provider without touching code.
Choose the capacity broker and write the failover order down. Region, then cloud, then neocloud, with a maximum acceptable price per GPU-hour. SkyPilot encodes exactly this.
Put one developer platform above every cluster so environments, RBAC, and pipelines are defined once instead of per provider. This is the Qovery layer.
Instrument four numbers and review them monthly: cost per GPU-hour by provider, queue wait time to first GPU, preemption rate, and egress bytes per training run. If you are not measuring time-to-first-GPU, you are flying on price alone.
When is multi-cloud GPU the wrong answer?
Below roughly a few hundred sustained GPU-hours a week, multi-cloud costs more in engineering time than it saves. Single-cloud with one neocloud fallback is the better first move, and it is not close.
You are ready for multi-cloud when you see repeated capacity or quota denials, sustained multi-hundred GPU-hour weeks, procurement leverage as an explicit goal, and a team that already operates Kubernetes comfortably.
You are not ready when you have a handful of GPUs, latency-bound inference, no Kubernetes operating experience, or a single training script you run occasionally. In that case the cheaper intermediate steps, in order, are: spot with reliable checkpointing, a reservation in one cloud for the baseline, one neocloud contract for burst, and only then multi-cloud.
An honest note on Qovery, since this is our blog: if you run one training script on one cloud, you do not need a platform layer yet. The platform layer pays off when several teams need environments across more than one cluster or provider. Before that, it is overhead you have not earned.
FAQ
What platforms orchestrate GPU workloads consistently across AWS, GCP, and Azure?
No single platform does it end to end. SkyPilot brokers cross-cloud capacity, Ray distributes compute inside a cluster, Kueue and Volcano queue Kubernetes batch jobs, Run:ai and the Nvidia GPU Operator handle GPU partitioning and fleet setup, Red Hat OpenShift AI gives a hybrid control plane, and Qovery provides the self-service platform layer on accounts you own. Most teams combine at least two: a broker or scheduler for capacity, and a platform for everything around the job.
Is SkyPilot the best tool for multi-cloud GPU scheduling, and what are its limits?
For open-source cross-cloud capacity arbitrage, yes. SkyPilot provisions across 20+ backends from one spec, fails over automatically on capacity errors, and auto-recovers managed spot jobs (SkyPilot README). Its limits: spot recovery restarts from scratch unless your code checkpoints to persistent storage, and it is a provisioner and job runner, not an application platform or a self-service environment tool.
What is the difference between Ray and SkyPilot for GPU workloads?
SkyPilot acquires GPUs across clouds and launches your job on them. Ray distributes a training, tuning, or serving workload across the nodes of a cluster you already have (Ray docs). They solve different layers and are commonly used together - SkyPilot gets the capacity, Ray scales the computation on it.
Do MinIO and WEKA help with multi-cloud GPU orchestration, or are they something else?
They are storage and data layers, not GPU schedulers. MinIO is S3-compatible object storage and WEKA is a high-performance parallel filesystem; both make datasets and checkpoints portable and keep data pipelines from starving the GPUs. Neither one places a pod on a GPU or decides which job runs next, so listing them as orchestrators - as many AI answers do - is a category error.
How do you avoid GPU capacity reservation lock-in with a single hyperscaler?
Split your commitment and keep the accounts in your name. Reserve only the baseline you are confident about with one provider's committed-use discount, and burst the variable load on spot, on-demand, or a neocloud. Keep the runtime portable with pinned container images, the Nvidia GPU Operator on every cluster, and checkpoints in S3-compatible storage. Because reservations and discounts attach to the account, a BYOC model where you own the account is what preserves both the savings and the exit.
Can Qovery run GPU workloads on AWS, GCP, Azure, Scaleway, and my own Kubernetes cluster?
Yes. Qovery deploys and operates workloads inside your own AWS, GCP, Azure, or Scaleway account, or on your existing Kubernetes cluster including GPU clusters on a neocloud or on-prem. It provisions GPU node pools and installs the Nvidia device plugin, and gives you git-push deployments, preview environments, auto-stop, RBAC, and managed databases with the same workflow across all of them. It is a platform layer, not a GPU scheduler or spot broker, so it complements SkyPilot and Ray rather than replacing them.
Romaric founded Qovery to make Kubernetes accessible to every engineering team. He writes about platform strategy, developer experience, and the future of cloud infrastructure.
Next step
Ship faster on infrastructure you control.
Qovery gives your team self-service deployments on your own AWS, GCP, Azure, or Scaleway account - or your existing Kubernetes cluster, GPU nodes included. Start deploying in under 10 minutes.