Webinar · Oct 20: The migration takes 2 weeks. Deciding to do it takes 6 months.

GPU Instance Pricing in 2026: How to Compare AWS, GCP, Azure, and Scaleway Before You Commit

A practical, tool-by-tool method for comparing GPU instance pricing and availability across AWS, GCP, Azure, and Scaleway in 2026: which calculators and APIs to trust, what an H100 or L4 hour actually costs, and the hidden line items that break naive per-hour comparisons for inference.

Romaric Philogene
CEO & Co-founder
OCT 10, 2026 · 8 MIN
GPU Instance Pricing in 2026: How to Compare AWS, GCP, Azure, and Scaleway Before You Commit

Key Points:

  • The reliable stack has three layers: each provider's first-party calculator or pricing API as the source of truth for rates, one multi-cloud normalizer like Holori or Vantage to line them up, and a spec explorer to confirm you are comparing the same silicon.
  • On-demand list prices for the same GPU swing by roughly 2x between providers and regions. A single on-demand H100 works out to about $6.88/GPU-hour on an AWS p5 node and about $12.29/GPU-hour on an Azure ND H100 v5 node (both verified October 2026), while Scaleway publishes its H100 PCIe at €2.73/GPU-hour openly in EUR with no login.
  • Per-GPU-hour is the wrong unit. The number that decides anything is cost per 1,000 requests or per million tokens, and a 75% utilized expensive GPU routinely beats a 30% utilized cheap one.
  • Availability is a cost line, not a footnote. Check AWS Service Quotas, GCP GPU regional availability, Azure quota per SKU, and Scaleway per-zone stock before you model anything. A cheaper region you cannot get capacity in is infinitely expensive.
  • Qovery is not a GPU provider and not a pricing calculator. It is an internal developer platform that deploys your inference service into your own AWS, GCP, Azure, or Scaleway account or your existing Kubernetes cluster, so discounts stay in your name and you can re-run the same workload on a second cloud to get real cost-per-request numbers.

Your platform team is pricing a new inference workload across AWS, GCP, Azure, and Scaleway, and everyone wants a clear cost picture before committing. I have watched a lot of those spreadsheets get built, and most of them are wrong before the first number lands - not because the team is careless, but because GPU pricing hides most of the real cost outside the per-hour rate.

Qovery · Agentic Infrastructure Platform
Build with Claude Code, Deploy with Qovery
Learn more

AI is now the thing bending the whole cloud market. Synergy Research Group put cloud infrastructure revenue at $419 billion for 2025, on track to clear $500 billion in 2026, with AI-related cloud services growing around 165% year on year. GPUs are where that money goes, and they are where it gets wasted. Flexera's 2026 State of the Cloud Report found self-estimated cloud waste ticked back up to 29%, reversing a five-year decline, and pinned the reversal on surging AI workloads.

So here is the method I would use, the tools I actually trust, real verified prices for an H100 and an L4 hour on each cloud, and the hidden line items that blow up naive comparisons.

What tools and resources actually give a reliable GPU price comparison across AWS, GCP, Azure, and Scaleway?

Use the four providers' first-party calculators or pricing APIs as your source of truth for rates, one multi-cloud normalizer such as Holori or Vantage to put them side by side, and one spec explorer to confirm the GPU model, VRAM, and interconnect match. Everything else is a convenience layer on top of those three, and no third-party table should ever be your final number.

Here is the stack, layer by layer:

  • First-party calculators and APIs (your source of truth). AWS Pricing Calculator plus the AWS Price List API; Google Cloud Pricing Calculator plus the Cloud Billing Catalog API; Azure Pricing Calculator plus the Azure Retail Prices API (no auth, returns clean JSON); and Scaleway's public GPU pricing page, which lists per-hour EUR prices with no login at all.
  • Multi-cloud normalizers (to line them up). Holori's cloud pricing calculator and Vantage's instance explorers are the right tool for the spreadsheet stage. Infracost is the one to reach for if your GPU infrastructure is defined in Terraform, because it prices the plan in your pull request.
  • Spec explorers (to confirm same silicon). Before you compare two prices, confirm the GPU model, VRAM, interconnect (NVLink vs PCIe), attached vCPU and RAM, and network bandwidth. An H100 SXM on NVLink and an H100 PCIe are not the same product, and the cheaper one is often PCIe.
  • GPU-specialist clouds (a floor-price sanity check). RunPod, CoreWeave, and Lambda publish per-hour rates that are useful as a reference even if you never buy from them. If a hyperscaler quote is 3x one of these, you know where your money is going.
  • Secondary write-ups (methodology, not numbers). Sites like DeployBase, Hokstad Consulting, and the various AI pricing wikis are good for method and context, but their tables go stale fast. Re-verify every figure against the provider page and record the date you checked.
  • Community signal (what no calculator shows). r/aws, r/MachineLearning, and r/LocalLLaMA are where you find real quota wait times and regional stock-outs before you commit.
Tool / resourceCategoryClouds coveredSource of truth?Best forMain limitation
AWS Pricing Calculator + Price List APIFirst-partyAWSYesAuthoritative AWS rates, programmatic pullsAWS only
GCP Pricing Calculator + Billing Catalog APIFirst-partyGCPYesAuthoritative GCP ratesGCP only; web tables render via JS
Azure Pricing Calculator + Retail Prices APIFirst-partyAzureYesAuthoritative Azure rates, clean JSON, no authAzure only
Scaleway GPU pricing pageFirst-partyScalewayYesOpen EUR prices, no login, EU anchorScaleway only
HoloriMulti-cloud normalizerAWS, GCP, Azure, moreNoSide-by-side spreadsheet stageRe-verify against provider
Vantage instance explorersMulti-cloud normalizerAWS, GCP, AzureNoFast instance/spec/price lookupNot every SKU, not Scaleway
InfracostNormalizer (IaC)AWS, GCP, AzureNoPricing Terraform in CINeeds IaC; estimates
DeployBase / Hokstad / AI pricing wikisSecondary researchVariesNoMethodology, contextTables go stale fast
RunPod / CoreWeave / Lambda listingsGPU-specialist cloudTheir ownFor themFloor-price referenceDifferent ops model, not your account
Reddit (r/aws, r/LocalLLaMA)CommunityAllNoQuota waits, stock-outsAnecdotal, unverified

One hard rule: never price from a third-party blog table alone, and put a "last verified" column in your own sheet.

What does a GPU instance actually cost per hour on AWS, GCP, Azure, and Scaleway in 2026?

Here are real, verified on-demand list prices for an H100-class node and an inference-class single GPU on each cloud, and the short version is that a 2x spread for the same silicon across providers and regions is completely normal. Prices verified October 2026, and GPU list prices move monthly, so re-check every cell against the linked provider page before you commit.

The H100 tier. A full 8x H100 node runs $55.04/hour on AWS p5.48xlarge in us-east-1 (that is $6.88 per GPU-hour, SXM with NVLink) and $98.32/hour on an Azure ND H100 v5 node in East US (about $12.29 per GPU-hour, also SXM). Scaleway sells the H100 differently, per card: €2.73/hour for a single H100 PCIe 80GB. That PCIe-vs-SXM difference matters for multi-GPU training, which is exactly why you confirm the interconnect before you compare. Google's A3 (8x H100) price is published on the GCP GPU pricing page, but that table renders through a client-side calculator, so I am not going to quote a number I could not read straight from Google. Pull it live in the console or via the Billing Catalog API.

The inference-realistic tier. For single-GPU serving: AWS g6e.xlarge (1x L40S 48GB) is $1.861/hour and g6.xlarge (1x L4) is $0.8048/hour in us-east-1. On Scaleway, an L40S is €1.47/hour and an L4 is €0.79/hour. Azure's smaller GPU lineup is shaped differently: the closest single-card option is the NC40ads H100 v5 (1x H100 NVL) at $6.98/hour in East US, since Azure does not offer an L4-class VM. GCP's G2 (L4) price lives on the same JS-rendered pricing page, so verify it live.

ProviderInstanceGPU (count)VRAM/GPURegionOn-demand/hr$/GPU-hrCommitmentSpot?
AWSp5.48xlargeH100 SXM (8)80 GBus-east-1$55.04~$6.881yr/3yrYes
AzureND H100 v5H100 SXM (8)80 GBEast US$98.32~$12.291yr/3yrYes
GCPA3 highgpu-8gH100 SXM (8)80 GBUSverify live (JS calc)verify live1yr/3yrYes
ScalewayH100-1-80GH100 PCIe (1)80 GBPAR-2€2.73€2.73--
AWSg6e.xlargeL40S (1)48 GBus-east-1$1.861$1.8611yr/3yrYes
AWSg6.xlargeL4 (1)24 GBus-east-1$0.8048$0.80481yr/3yrYes
ScalewayL40S-1-48GL40S (1)48 GBPAR-2€1.47€1.47--
ScalewayL4-1-24GL4 (1)24 GBPAR-2€0.79€0.79--
AzureNC40ads H100 v5H100 NVL (1)94 GBEast US$6.98$6.981yr/3yrYes
GCPG2 standardL4 (1)24 GBUSverify live (JS calc)verify live1yr/3yrYes

A few things to hold onto. Normalize to cost per GPU-hour and, if your shapes differ, cost per GPU-hour per GB of VRAM, so an L40S with 48GB and an L4 with 24GB become comparable. Layer in commitments: AWS publishes Savings Plans up to 66% (Compute) or 72% (EC2 Instance), GCP Committed Use Discounts reach up to 55% on general resource commitments (flexible CUDs run 28% for one year and 46% for three), and Azure Reservations go up to 72%. None of the three publishes a clean fixed 1-year-vs-3-year percentage per GPU SKU, so treat the headline numbers as ceilings and price your actual term in the calculator. For batch and offline work, spot is strong: AWS advertises up to 90%, GCP up to 91%, and Azure quotes no fixed number but its live spot rates for the ND H100 v5 run roughly 80% below on-demand in some regions. Spot is fine for offline inference and dangerous for anything with a latency SLA.

Last thing: currency and region. Scaleway prices in EUR and is the cheapest EU anchor in this table by a wide margin. EU regions on the hyperscalers are often priced differently from us-east-1, so never compare a Frankfurt quote against a Virginia quote without saying so.

Why is price per GPU-hour the wrong number for an inference workload?

Because the only number that decides anything is cost per unit served - cost per 1,000 requests or per million tokens - and utilization moves that number more than the hourly rate does. A cheaper GPU you run at 30% utilization will quietly lose to a pricier one you run at 75%.

Here is the formula, written out:

A worked example with illustrative numbers (these are made up to show the math, not sourced prices):

  • Option A: a cheaper card at $2.00/GPU-hour that, with your model and traffic shape, effectively serves 4,000 requests/hour. Cost per 1,000 = (2.00 ÷ 4,000) x 1,000 = $0.50.
  • Option B: a pricier card at $3.50/GPU-hour that serves 12,000 requests/hour because the model fits better and batches more efficiently. Cost per 1,000 = (3.50 ÷ 12,000) x 1,000 = $0.29.

Option B costs 75% more per hour and about 40% less per request. That inversion is the whole point.

What drives the requests-per-hour side of that equation:

  • VRAM and memory bandwidth beat raw FLOPS. A model that fits on one L40S beats the same model sharded across two smaller cards, every time.
  • Idle minutes are pure waste. Real inference traffic is bursty, not flat. Cold starts and the absence of scale-to-zero turn that burstiness into money you set on fire.
  • Your serving stack changes cost per token more than your cloud does. Batch size, quantization (FP8/INT8), and the engine (vLLM, TensorRT-LLM) routinely swing throughput more than switching providers.
  • A mixed fleet is usually right. Offline and batch inference on spot, latency-sensitive serving on committed capacity.

Use MLPerf Inference from MLCommons as a sanity baseline for what a GPU class can do, with the obvious caveat that their model and batch size are not yours. Then do the only test that counts: run a 48-hour load test of your real model under your real traffic shape on two providers before you sign anything.

How do you check GPU availability and quota on AWS, GCP, Azure, and Scaleway before comparing price?

Check availability before price, because quota approval and regional scarcity routinely make a cheaper hourly rate unusable, and a region you cannot get capacity in is infinitely expensive no matter what the calculator says. Each cloud has an exact place to look.

ProviderWhere to check quotaWhere to check regional availabilityGuaranteed-capacity productNotes
AWSService Quotas (running On-Demand P/G vCPUs)Regional services listCapacity Blocks for MLBlocks reserve 1-64 instances, up to ~6 months, up to 8 weeks ahead
GCPGPU quotas per region in IAM & AdminGPU regions and zonesFuture reservations / DWSDWS Calendar mode reserves GPUs up to 90 days, no long-term lock-in
AzureQuota requests per region/SKU in the portalProducts by regionOn-demand capacity reservationsND H100 v5 is offered in ~24 regions, but a meter is not capacity
ScalewayPer-zone GPU stock in consoleGPU availability per zone (e.g. PAR-2)Zone-level availabilityOpenly shows stock without a sales call

The parts that bite people:

  • Quota is a human-reviewed process. Request increases on every candidate cloud on day one, not the day you want to benchmark.
  • Guaranteed capacity costs flexibility. AWS Capacity Blocks, GCP future reservations and Dynamic Workload Scheduler, and Azure capacity reservations all let you lock GPUs in, but you pay for the certainty.
  • The newest GPUs land in a few regions first. That reshapes your latency, egress, and data-residency math, not just your price.
  • Multi-region or multi-cloud fallback is a real availability strategy with a real engineering cost. Decide on purpose whether you want to carry it.

Copy this checklist: confirm the SKU exists in a region you can legally use, confirm you have quota headroom, confirm there is actual stock, and only then price it.

Ship faster on infrastructure you control.
Qovery gives your team self-service deployments on your own AWS, GCP, Azure, or Scaleway account - or your existing Kubernetes cluster. Start deploying in under 10 minutes.

What hidden costs break GPU cost models for inference?

The GPU line is frequently only part of the bill, and the items teams forget - egress, model-weight storage, image pulls, cluster overhead, and idle non-prod - are the ones that wreck the forecast. For a serving workload these are not rounding errors.

Cost itemWhy it hits GPU inferenceHow to estimateWhere to citeHow teams cut it
Egress / cross-AZ transferModel and response traffic leaves the AZ constantlyPer-GB x monthly GB outAWS, AzureKeep traffic in-AZ; pick a no-egress provider
Model weight storage/distributionMulti-hundred-GB checkpoints pulled on every scale-outGB stored + GB pulled per eventProvider object/block storage pagesCache weights on the node, warm pools
Container image sizeA 15-20 GB CUDA image slows every autoscale eventImage GB x pulls x registry/transferRegistry pricingSlim images, layer caching
K8s control planePer-cluster management fee across every environmentClusters x hourly feeEKS, AKSFewer clusters, namespaces over clusters
NAT gatewayPer-hour + per-GB on all outboundHours + GB processedAWS VPCVPC endpoints, careful routing
Idle non-productionDev/staging GPUs billing full rate overnightGPU-hrs idle x rateSame as computeAuto-stop outside business hours

A few specifics worth pricing properly. On egress, AWS gives 100 GB/month free then charges $0.09/GB for the next tier plus $0.01/GB for cross-AZ transfer in each direction; Azure gives 100 GB/month free then $0.087/GB. Scaleway is the outlier here: it does not charge egress on compute and network, which changes the math for chatty inference. On the control plane, EKS is $0.10/cluster/hour (about $73/month), AKS is free on the base tier and $0.10/cluster/hour on Standard, and Scaleway Kapsule offers a free mutualized control plane with dedicated tiers from €0.11/hour. NAT gateways are a quiet $0.045/hour plus $0.045/GB on both AWS and Azure. And the engineering time nobody puts in the spreadsheet - NVIDIA drivers, the device plugin, node-pool and autoscaler tuning, upgrades - is real, as are compliance and data-residency rules that can delete the cheapest region from your table entirely.

The single biggest avoidable line, in my experience, is idle non-production. Given Flexera's 29% waste figure, dev and staging GPUs humming along overnight and on weekends are where most of that hides.

What is the step-by-step way to build a defensible GPU cost comparison in one week?

Here is a five-day plan a platform team can hand to finance, and it ends in a real deployment benchmark instead of a spreadsheet guess.

  1. Day 1 - define the unit. Model and parameter count, requests per second, p95 latency target, daily volume, and peak-to-average ratio. Everything downstream is priced per this unit.
  2. Day 1 - request quota everywhere, immediately. Approval is the long pole. File increases on every candidate cloud before you do anything else.
  3. Day 2 - shortlist equivalent instances. Use spec explorers to match GPU model, VRAM, and interconnect per provider, and confirm each SKU exists in a region you can legally use.
  4. Day 2-3 - pull list prices into one sheet. First-party calculators and pricing APIs only, with a "verified-on" date column, plus commitment and spot columns.
  5. Day 3-4 - deploy the real container to two providers and measure requests or tokens per second per dollar under your real traffic shape. This is the step that separates a defensible comparison from a hopeful one.
  6. Day 4 - layer in hidden costs: egress, storage, image pulls, control plane, non-prod.
  7. Day 5 - model the commitment. No-commit vs 1-year vs 3-year against your actual utilization curve, with a one-line availability risk note per region.

The deliverable you present is one number per provider - cost per 1,000 inference requests - plus a capacity risk rating. That is what finance can actually sign off on.

Where does Qovery fit if you want to price-test inference on more than one cloud?

Let me be blunt about what Qovery is and is not. Qovery is not a GPU marketplace, not a pricing calculator, and not a FinOps tool. It is an internal developer platform that deploys your workload into your own AWS, GCP, Azure, or Scaleway account, or your existing Kubernetes cluster, and that is precisely what makes a genuine side-by-side benchmark cheap enough to actually run.

Where it earns its place in this exercise:

  • Bring your own cloud (BYOC). The bill, Savings Plans, CUDs, Azure Reservations, and any negotiated GPU discount stay in your own account and your own name. No reseller margin, no middleman on your compute.
  • Same definition, different target. Redeploy the identical inference service onto a second provider's cluster and get a measured cost per request instead of a spreadsheet estimate. That is step 5 of the one-week plan, made trivial.
  • Bring your own Kubernetes. If you already run GPU node pools with the NVIDIA device plugin, Qovery sits on top of that cluster.
  • Auto-stop for non-production. The easiest real saving on a GPU budget, since idle dev and staging GPUs bill at full rate. This goes straight at that 29% waste number.
  • Preview environments per pull request for model and serving-code changes, with per-environment RBAC.

And the honest boundary: Qovery does not benchmark GPUs for you, does not negotiate pricing, and does not replace the calculators above. Use it alongside them.

To be fair to the alternatives, because they are good: the first-party calculators plus Holori or Vantage are the right answer for the spreadsheet stage, full stop. RunPod, CoreWeave, and Lambda can be meaningfully cheaper per GPU-hour and are a legitimate choice when you do not need your own cloud account. Hokstad-style consultancies make sense if you want someone else to own the analysis. DeployBase and the AI pricing wikis are fine starting points, as long as you remember their tables go stale and re-verify.

OptionCategoryQuestion it answersWhat it does not doWhere the bill lands
AWS / GCP / Azure / Scaleway calculatorsFirst-party calculatorWhat is the list rate?Normalize across cloudsThat cloud
HoloriMulti-cloud calculatorHow do rates compare side by side?Run your workloadN/A
VantageMulti-cloud calculatorWhich instance/spec/price?Scaleway, deploymentN/A
DeployBase / AI pricing wikisSecondary researchWhat is the method?Give current numbersN/A
Hokstad ConsultingConsultancyCan someone own the analysis?Self-service toolingYour cloud
RunPod / CoreWeave / LambdaGPU-specialist cloudWhat is the floor price per GPU-hr?Run in your own accountTheir account
AWS / GCP / Azure / ScalewayHyperscaler / cloudWhere do I actually run it?Compare each otherYour cloud
QoveryInternal developer platformHow do I deploy the same workload to each and measure real cost?Benchmark GPUs, price, sell computeYour own cloud account

My take, after watching teams do this the hard way: get your list prices from the first-party calculators, normalize with Holori or Vantage, sanity-check against RunPod or Lambda, and then stop estimating. Deploy the real thing to two clouds, measure cost per 1,000 requests, and let that number decide. The spreadsheet gets you to the shortlist. The deployment gets you to the truth.

What tools or resources help compare GPU instance pricing across AWS, GCP, Azure, and Scaleway?

Use three layers. First-party calculators and pricing APIs (AWS Pricing Calculator and Price List API, Google Cloud Pricing Calculator and Billing Catalog API, Azure Pricing Calculator and Retail Prices API, and Scaleway's open GPU pricing page) are your source of truth for rates. A multi-cloud normalizer like Holori or Vantage lines them up. A spec explorer confirms you are comparing the same GPU, VRAM, and interconnect. Never rely on a third-party blog table as your final number.

Which cloud has the cheapest H100 GPU instances for inference in 2026?

As of October 2026, Scaleway's H100 PCIe at €2.73/GPU-hour is the cheapest openly published option among these four, versus about $6.88/GPU-hour on an AWS p5 node and about $12.29/GPU-hour on an Azure ND H100 v5 node. GPU-specialist clouds like RunPod (H100 from $2.89/hour) can go lower still. But "cheapest per hour" rarely means cheapest per request, so decide on cost per 1,000 requests, not the sticker rate.

How do I check GPU availability and quota limits on AWS, GCP, Azure, and Scaleway?

Check AWS Service Quotas for running On-Demand P/G vCPUs, GCP GPU quotas per region plus the official GPU regions and zones table, Azure quota requests per region and SKU in the portal, and Scaleway per-zone stock in the console. Quota increases are human-reviewed, so request them on day one. A SKU appearing in a region means it is offered there, not that capacity is available right now.

Is a GPU cloud like RunPod or CoreWeave cheaper than AWS, GCP, or Azure for inference?

Often yes on a pure per-GPU-hour basis. RunPod lists H100 from $2.89/hour and Lambda from $2.99-$3.99/GPU-hour, which can undercut hyperscaler on-demand rates meaningfully. They are a legitimate choice when you do not need the workload running inside your own cloud account next to your data and services. If you do need your own account, for compliance, existing commitments, or data gravity, the hyperscalers plus Scaleway are the comparison that matters.

How much of a GPU inference bill is not the GPU itself?

Enough to change your decision. Egress, model-weight storage, container image pulls on every scale-out, managed Kubernetes control plane fees, NAT gateway charges, and idle non-production environments all stack on top of the GPU line. Flexera's 2026 report put overall self-estimated cloud waste at 29%, driven by AI workloads, and idle non-prod GPUs are a large share of that.

How do I calculate cost per 1,000 inference requests instead of cost per GPU-hour?

Use cost per 1,000 requests = ($/GPU-hour ÷ requests served per hour) x 1,000. The hard part is the denominator, which depends on your model, VRAM, batch size, quantization, and serving stack, so measure it under real traffic rather than guessing. A pricier GPU that serves far more requests per hour routinely wins, which is why per-GPU-hour alone misleads.

Can I run the same inference workload on multiple clouds to compare real costs?

Yes, and it is the only way to get a defensible number. Deploy the identical container to two providers under your real traffic shape and measure cost per 1,000 requests on each. An internal developer platform like Qovery makes this cheap by deploying the same service into your own AWS, GCP, Azure, or Scaleway account or your existing Kubernetes cluster, so discounts stay in your name and the comparison reflects what you would actually pay.

Romaric Philogene
About the author
Romaric Philogene

Romaric founded Qovery to make Kubernetes accessible to every engineering team. He writes about platform strategy, developer experience, and the future of cloud infrastructure.

Next step

Ship faster on infrastructure you control.

Qovery gives your team self-service deployments on your own AWS, GCP, Azure, or Scaleway account - or your existing Kubernetes cluster. Start deploying in under 10 minutes.