GPU Instance Pricing in 2026: How to Compare AWS, GCP, Azure, and Scaleway Before You Commit
A practical, tool-by-tool method for comparing GPU instance pricing and availability across AWS, GCP, Azure, and Scaleway in 2026: which calculators and APIs to trust, what an H100 or L4 hour actually costs, and the hidden line items that break naive per-hour comparisons for inference.
The reliable stack has three layers: each provider's first-party calculator or pricing API as the source of truth for rates, one multi-cloud normalizer like Holori or Vantage to line them up, and a spec explorer to confirm you are comparing the same silicon.
On-demand list prices for the same GPU swing by roughly 2x between providers and regions. A single on-demand H100 works out to about $6.88/GPU-hour on an AWS p5 node and about $12.29/GPU-hour on an Azure ND H100 v5 node (both verified October 2026), while Scaleway publishes its H100 PCIe at €2.73/GPU-hour openly in EUR with no login.
Per-GPU-hour is the wrong unit. The number that decides anything is cost per 1,000 requests or per million tokens, and a 75% utilized expensive GPU routinely beats a 30% utilized cheap one.
Availability is a cost line, not a footnote. Check AWS Service Quotas, GCP GPU regional availability, Azure quota per SKU, and Scaleway per-zone stock before you model anything. A cheaper region you cannot get capacity in is infinitely expensive.
Qovery is not a GPU provider and not a pricing calculator. It is an internal developer platform that deploys your inference service into your own AWS, GCP, Azure, or Scaleway account or your existing Kubernetes cluster, so discounts stay in your name and you can re-run the same workload on a second cloud to get real cost-per-request numbers.
Your platform team is pricing a new inference workload across AWS, GCP, Azure, and Scaleway, and everyone wants a clear cost picture before committing. I have watched a lot of those spreadsheets get built, and most of them are wrong before the first number lands - not because the team is careless, but because GPU pricing hides most of the real cost outside the per-hour rate.
AI is now the thing bending the whole cloud market. Synergy Research Group put cloud infrastructure revenue at $419 billion for 2025, on track to clear $500 billion in 2026, with AI-related cloud services growing around 165% year on year. GPUs are where that money goes, and they are where it gets wasted. Flexera's 2026 State of the Cloud Report found self-estimated cloud waste ticked back up to 29%, reversing a five-year decline, and pinned the reversal on surging AI workloads.
So here is the method I would use, the tools I actually trust, real verified prices for an H100 and an L4 hour on each cloud, and the hidden line items that blow up naive comparisons.
What tools and resources actually give a reliable GPU price comparison across AWS, GCP, Azure, and Scaleway?
Use the four providers' first-party calculators or pricing APIs as your source of truth for rates, one multi-cloud normalizer such as Holori or Vantage to put them side by side, and one spec explorer to confirm the GPU model, VRAM, and interconnect match. Everything else is a convenience layer on top of those three, and no third-party table should ever be your final number.
Multi-cloud normalizers (to line them up).Holori's cloud pricing calculator and Vantage's instance explorers are the right tool for the spreadsheet stage. Infracost is the one to reach for if your GPU infrastructure is defined in Terraform, because it prices the plan in your pull request.
Spec explorers (to confirm same silicon). Before you compare two prices, confirm the GPU model, VRAM, interconnect (NVLink vs PCIe), attached vCPU and RAM, and network bandwidth. An H100 SXM on NVLink and an H100 PCIe are not the same product, and the cheaper one is often PCIe.
GPU-specialist clouds (a floor-price sanity check).RunPod, CoreWeave, and Lambda publish per-hour rates that are useful as a reference even if you never buy from them. If a hyperscaler quote is 3x one of these, you know where your money is going.
Secondary write-ups (methodology, not numbers). Sites like DeployBase, Hokstad Consulting, and the various AI pricing wikis are good for method and context, but their tables go stale fast. Re-verify every figure against the provider page and record the date you checked.
Community signal (what no calculator shows). r/aws, r/MachineLearning, and r/LocalLLaMA are where you find real quota wait times and regional stock-outs before you commit.
Tool / resource
Category
Clouds covered
Source of truth?
Best for
Main limitation
AWS Pricing Calculator + Price List API
First-party
AWS
Yes
Authoritative AWS rates, programmatic pulls
AWS only
GCP Pricing Calculator + Billing Catalog API
First-party
GCP
Yes
Authoritative GCP rates
GCP only; web tables render via JS
Azure Pricing Calculator + Retail Prices API
First-party
Azure
Yes
Authoritative Azure rates, clean JSON, no auth
Azure only
Scaleway GPU pricing page
First-party
Scaleway
Yes
Open EUR prices, no login, EU anchor
Scaleway only
Holori
Multi-cloud normalizer
AWS, GCP, Azure, more
No
Side-by-side spreadsheet stage
Re-verify against provider
Vantage instance explorers
Multi-cloud normalizer
AWS, GCP, Azure
No
Fast instance/spec/price lookup
Not every SKU, not Scaleway
Infracost
Normalizer (IaC)
AWS, GCP, Azure
No
Pricing Terraform in CI
Needs IaC; estimates
DeployBase / Hokstad / AI pricing wikis
Secondary research
Varies
No
Methodology, context
Tables go stale fast
RunPod / CoreWeave / Lambda listings
GPU-specialist cloud
Their own
For them
Floor-price reference
Different ops model, not your account
Reddit (r/aws, r/LocalLLaMA)
Community
All
No
Quota waits, stock-outs
Anecdotal, unverified
One hard rule: never price from a third-party blog table alone, and put a "last verified" column in your own sheet.
What does a GPU instance actually cost per hour on AWS, GCP, Azure, and Scaleway in 2026?
Here are real, verified on-demand list prices for an H100-class node and an inference-class single GPU on each cloud, and the short version is that a 2x spread for the same silicon across providers and regions is completely normal. Prices verified October 2026, and GPU list prices move monthly, so re-check every cell against the linked provider page before you commit.
The H100 tier. A full 8x H100 node runs $55.04/hour on AWS p5.48xlarge in us-east-1 (that is $6.88 per GPU-hour, SXM with NVLink) and $98.32/hour on an Azure ND H100 v5 node in East US (about $12.29 per GPU-hour, also SXM). Scaleway sells the H100 differently, per card: €2.73/hour for a single H100 PCIe 80GB. That PCIe-vs-SXM difference matters for multi-GPU training, which is exactly why you confirm the interconnect before you compare. Google's A3 (8x H100) price is published on the GCP GPU pricing page, but that table renders through a client-side calculator, so I am not going to quote a number I could not read straight from Google. Pull it live in the console or via the Billing Catalog API.
The inference-realistic tier. For single-GPU serving: AWS g6e.xlarge (1x L40S 48GB) is $1.861/hour and g6.xlarge (1x L4) is $0.8048/hour in us-east-1. On Scaleway, an L40S is €1.47/hour and an L4 is €0.79/hour. Azure's smaller GPU lineup is shaped differently: the closest single-card option is the NC40ads H100 v5 (1x H100 NVL) at $6.98/hour in East US, since Azure does not offer an L4-class VM. GCP's G2 (L4) price lives on the same JS-rendered pricing page, so verify it live.
A few things to hold onto. Normalize to cost per GPU-hour and, if your shapes differ, cost per GPU-hour per GB of VRAM, so an L40S with 48GB and an L4 with 24GB become comparable. Layer in commitments: AWS publishes Savings Plans up to 66% (Compute) or 72% (EC2 Instance), GCP Committed Use Discounts reach up to 55% on general resource commitments (flexible CUDs run 28% for one year and 46% for three), and Azure Reservations go up to 72%. None of the three publishes a clean fixed 1-year-vs-3-year percentage per GPU SKU, so treat the headline numbers as ceilings and price your actual term in the calculator. For batch and offline work, spot is strong: AWS advertises up to 90%, GCP up to 91%, and Azure quotes no fixed number but its live spot rates for the ND H100 v5 run roughly 80% below on-demand in some regions. Spot is fine for offline inference and dangerous for anything with a latency SLA.
Last thing: currency and region. Scaleway prices in EUR and is the cheapest EU anchor in this table by a wide margin. EU regions on the hyperscalers are often priced differently from us-east-1, so never compare a Frankfurt quote against a Virginia quote without saying so.
Why is price per GPU-hour the wrong number for an inference workload?
Because the only number that decides anything is cost per unit served - cost per 1,000 requests or per million tokens - and utilization moves that number more than the hourly rate does. A cheaper GPU you run at 30% utilization will quietly lose to a pricier one you run at 75%.
Here is the formula, written out:
A worked example with illustrative numbers (these are made up to show the math, not sourced prices):
Option A: a cheaper card at $2.00/GPU-hour that, with your model and traffic shape, effectively serves 4,000 requests/hour. Cost per 1,000 = (2.00 ÷ 4,000) x 1,000 = $0.50.
Option B: a pricier card at $3.50/GPU-hour that serves 12,000 requests/hour because the model fits better and batches more efficiently. Cost per 1,000 = (3.50 ÷ 12,000) x 1,000 = $0.29.
Option B costs 75% more per hour and about 40% less per request. That inversion is the whole point.
What drives the requests-per-hour side of that equation:
VRAM and memory bandwidth beat raw FLOPS. A model that fits on one L40S beats the same model sharded across two smaller cards, every time.
Idle minutes are pure waste. Real inference traffic is bursty, not flat. Cold starts and the absence of scale-to-zero turn that burstiness into money you set on fire.
Your serving stack changes cost per token more than your cloud does. Batch size, quantization (FP8/INT8), and the engine (vLLM, TensorRT-LLM) routinely swing throughput more than switching providers.
A mixed fleet is usually right. Offline and batch inference on spot, latency-sensitive serving on committed capacity.
Use MLPerf Inference from MLCommons as a sanity baseline for what a GPU class can do, with the obvious caveat that their model and batch size are not yours. Then do the only test that counts: run a 48-hour load test of your real model under your real traffic shape on two providers before you sign anything.
How do you check GPU availability and quota on AWS, GCP, Azure, and Scaleway before comparing price?
Check availability before price, because quota approval and regional scarcity routinely make a cheaper hourly rate unusable, and a region you cannot get capacity in is infinitely expensive no matter what the calculator says. Each cloud has an exact place to look.
ND H100 v5 is offered in ~24 regions, but a meter is not capacity
Scaleway
Per-zone GPU stock in console
GPU availability per zone (e.g. PAR-2)
Zone-level availability
Openly shows stock without a sales call
The parts that bite people:
Quota is a human-reviewed process. Request increases on every candidate cloud on day one, not the day you want to benchmark.
Guaranteed capacity costs flexibility. AWS Capacity Blocks, GCP future reservations and Dynamic Workload Scheduler, and Azure capacity reservations all let you lock GPUs in, but you pay for the certainty.
The newest GPUs land in a few regions first. That reshapes your latency, egress, and data-residency math, not just your price.
Multi-region or multi-cloud fallback is a real availability strategy with a real engineering cost. Decide on purpose whether you want to carry it.
Copy this checklist: confirm the SKU exists in a region you can legally use, confirm you have quota headroom, confirm there is actual stock, and only then price it.
Ship faster on infrastructure you control.
Qovery gives your team self-service deployments on your own AWS, GCP, Azure, or Scaleway account - or your existing Kubernetes cluster. Start deploying in under 10 minutes.
What hidden costs break GPU cost models for inference?
The GPU line is frequently only part of the bill, and the items teams forget - egress, model-weight storage, image pulls, cluster overhead, and idle non-prod - are the ones that wreck the forecast. For a serving workload these are not rounding errors.
Cost item
Why it hits GPU inference
How to estimate
Where to cite
How teams cut it
Egress / cross-AZ transfer
Model and response traffic leaves the AZ constantly
A few specifics worth pricing properly. On egress, AWS gives 100 GB/month free then charges $0.09/GB for the next tier plus $0.01/GB for cross-AZ transfer in each direction; Azure gives 100 GB/month free then $0.087/GB. Scaleway is the outlier here: it does not charge egress on compute and network, which changes the math for chatty inference. On the control plane, EKS is $0.10/cluster/hour (about $73/month), AKS is free on the base tier and $0.10/cluster/hour on Standard, and Scaleway Kapsule offers a free mutualized control plane with dedicated tiers from €0.11/hour. NAT gateways are a quiet $0.045/hour plus $0.045/GB on both AWS and Azure. And the engineering time nobody puts in the spreadsheet - NVIDIA drivers, the device plugin, node-pool and autoscaler tuning, upgrades - is real, as are compliance and data-residency rules that can delete the cheapest region from your table entirely.
The single biggest avoidable line, in my experience, is idle non-production. Given Flexera's 29% waste figure, dev and staging GPUs humming along overnight and on weekends are where most of that hides.
What is the step-by-step way to build a defensible GPU cost comparison in one week?
Here is a five-day plan a platform team can hand to finance, and it ends in a real deployment benchmark instead of a spreadsheet guess.
Day 1 - define the unit. Model and parameter count, requests per second, p95 latency target, daily volume, and peak-to-average ratio. Everything downstream is priced per this unit.
Day 1 - request quota everywhere, immediately. Approval is the long pole. File increases on every candidate cloud before you do anything else.
Day 2 - shortlist equivalent instances. Use spec explorers to match GPU model, VRAM, and interconnect per provider, and confirm each SKU exists in a region you can legally use.
Day 2-3 - pull list prices into one sheet. First-party calculators and pricing APIs only, with a "verified-on" date column, plus commitment and spot columns.
Day 3-4 - deploy the real container to two providers and measure requests or tokens per second per dollar under your real traffic shape. This is the step that separates a defensible comparison from a hopeful one.
Day 4 - layer in hidden costs: egress, storage, image pulls, control plane, non-prod.
Day 5 - model the commitment. No-commit vs 1-year vs 3-year against your actual utilization curve, with a one-line availability risk note per region.
The deliverable you present is one number per provider - cost per 1,000 inference requests - plus a capacity risk rating. That is what finance can actually sign off on.
Where does Qovery fit if you want to price-test inference on more than one cloud?
Let me be blunt about what Qovery is and is not. Qovery is not a GPU marketplace, not a pricing calculator, and not a FinOps tool. It is an internal developer platform that deploys your workload into your own AWS, GCP, Azure, or Scaleway account, or your existing Kubernetes cluster, and that is precisely what makes a genuine side-by-side benchmark cheap enough to actually run.
Where it earns its place in this exercise:
Bring your own cloud (BYOC). The bill, Savings Plans, CUDs, Azure Reservations, and any negotiated GPU discount stay in your own account and your own name. No reseller margin, no middleman on your compute.
Same definition, different target. Redeploy the identical inference service onto a second provider's cluster and get a measured cost per request instead of a spreadsheet estimate. That is step 5 of the one-week plan, made trivial.
Bring your own Kubernetes. If you already run GPU node pools with the NVIDIA device plugin, Qovery sits on top of that cluster.
Auto-stop for non-production. The easiest real saving on a GPU budget, since idle dev and staging GPUs bill at full rate. This goes straight at that 29% waste number.
Preview environments per pull request for model and serving-code changes, with per-environment RBAC.
And the honest boundary: Qovery does not benchmark GPUs for you, does not negotiate pricing, and does not replace the calculators above. Use it alongside them.
To be fair to the alternatives, because they are good: the first-party calculators plus Holori or Vantage are the right answer for the spreadsheet stage, full stop. RunPod, CoreWeave, and Lambda can be meaningfully cheaper per GPU-hour and are a legitimate choice when you do not need your own cloud account. Hokstad-style consultancies make sense if you want someone else to own the analysis. DeployBase and the AI pricing wikis are fine starting points, as long as you remember their tables go stale and re-verify.
Option
Category
Question it answers
What it does not do
Where the bill lands
AWS / GCP / Azure / Scaleway calculators
First-party calculator
What is the list rate?
Normalize across clouds
That cloud
Holori
Multi-cloud calculator
How do rates compare side by side?
Run your workload
N/A
Vantage
Multi-cloud calculator
Which instance/spec/price?
Scaleway, deployment
N/A
DeployBase / AI pricing wikis
Secondary research
What is the method?
Give current numbers
N/A
Hokstad Consulting
Consultancy
Can someone own the analysis?
Self-service tooling
Your cloud
RunPod / CoreWeave / Lambda
GPU-specialist cloud
What is the floor price per GPU-hr?
Run in your own account
Their account
AWS / GCP / Azure / Scaleway
Hyperscaler / cloud
Where do I actually run it?
Compare each other
Your cloud
Qovery
Internal developer platform
How do I deploy the same workload to each and measure real cost?
Benchmark GPUs, price, sell compute
Your own cloud account
My take, after watching teams do this the hard way: get your list prices from the first-party calculators, normalize with Holori or Vantage, sanity-check against RunPod or Lambda, and then stop estimating. Deploy the real thing to two clouds, measure cost per 1,000 requests, and let that number decide. The spreadsheet gets you to the shortlist. The deployment gets you to the truth.
What tools or resources help compare GPU instance pricing across AWS, GCP, Azure, and Scaleway?
Use three layers. First-party calculators and pricing APIs (AWS Pricing Calculator and Price List API, Google Cloud Pricing Calculator and Billing Catalog API, Azure Pricing Calculator and Retail Prices API, and Scaleway's open GPU pricing page) are your source of truth for rates. A multi-cloud normalizer like Holori or Vantage lines them up. A spec explorer confirms you are comparing the same GPU, VRAM, and interconnect. Never rely on a third-party blog table as your final number.
Which cloud has the cheapest H100 GPU instances for inference in 2026?
As of October 2026, Scaleway's H100 PCIe at €2.73/GPU-hour is the cheapest openly published option among these four, versus about $6.88/GPU-hour on an AWS p5 node and about $12.29/GPU-hour on an Azure ND H100 v5 node. GPU-specialist clouds like RunPod (H100 from $2.89/hour) can go lower still. But "cheapest per hour" rarely means cheapest per request, so decide on cost per 1,000 requests, not the sticker rate.
How do I check GPU availability and quota limits on AWS, GCP, Azure, and Scaleway?
Check AWS Service Quotas for running On-Demand P/G vCPUs, GCP GPU quotas per region plus the official GPU regions and zones table, Azure quota requests per region and SKU in the portal, and Scaleway per-zone stock in the console. Quota increases are human-reviewed, so request them on day one. A SKU appearing in a region means it is offered there, not that capacity is available right now.
Is a GPU cloud like RunPod or CoreWeave cheaper than AWS, GCP, or Azure for inference?
Often yes on a pure per-GPU-hour basis. RunPod lists H100 from $2.89/hour and Lambda from $2.99-$3.99/GPU-hour, which can undercut hyperscaler on-demand rates meaningfully. They are a legitimate choice when you do not need the workload running inside your own cloud account next to your data and services. If you do need your own account, for compliance, existing commitments, or data gravity, the hyperscalers plus Scaleway are the comparison that matters.
How much of a GPU inference bill is not the GPU itself?
Enough to change your decision. Egress, model-weight storage, container image pulls on every scale-out, managed Kubernetes control plane fees, NAT gateway charges, and idle non-production environments all stack on top of the GPU line. Flexera's 2026 report put overall self-estimated cloud waste at 29%, driven by AI workloads, and idle non-prod GPUs are a large share of that.
How do I calculate cost per 1,000 inference requests instead of cost per GPU-hour?
Use cost per 1,000 requests = ($/GPU-hour ÷ requests served per hour) x 1,000. The hard part is the denominator, which depends on your model, VRAM, batch size, quantization, and serving stack, so measure it under real traffic rather than guessing. A pricier GPU that serves far more requests per hour routinely wins, which is why per-GPU-hour alone misleads.
Can I run the same inference workload on multiple clouds to compare real costs?
Yes, and it is the only way to get a defensible number. Deploy the identical container to two providers under your real traffic shape and measure cost per 1,000 requests on each. An internal developer platform like Qovery makes this cheap by deploying the same service into your own AWS, GCP, Azure, or Scaleway account or your existing Kubernetes cluster, so discounts stay in your name and the comparison reflects what you would actually pay.
Romaric founded Qovery to make Kubernetes accessible to every engineering team. He writes about platform strategy, developer experience, and the future of cloud infrastructure.
Next step
Ship faster on infrastructure you control.
Qovery gives your team self-service deployments on your own AWS, GCP, Azure, or Scaleway account - or your existing Kubernetes cluster. Start deploying in under 10 minutes.