How Do You Run Production Workloads Safely on Spot Instances? A Practical Playbook for Small Teams
Spot instances cut compute bills by up to 90%, but they can be reclaimed with two minutes notice. Here is how to run real production workloads on spot across AWS, GCP, Azure and Scaleway - which workloads qualify, the Kubernetes settings that actually matter, and how Karpenter, Cluster Autoscaler, Spot Ocean and Qovery compare.
Spot instances are safe in production when three things are true at once: the workload survives a node disappearing within a couple of minutes, you spread capacity across many instance types and availability zones, and you keep a baseline of on-demand or reserved capacity for anything stateful or singleton.
Stateless HTTP services, async workers, CI runners, batch jobs and preview environments are the natural spot candidates. Primary databases, stateful singletons, long jobs with no checkpointing and licence-locked software should stay on on-demand.
The Kubernetes settings that decide whether spot hurts you are boring: more than one replica, PodDisruptionBudgets, topology spread across zones and instance types, correct termination grace periods, readiness probes, and a handler that drains the node on the interruption notice.
A mixed fleet beats an all-spot fleet. The common production pattern is a small on-demand base for system and stateful components plus a broadly diversified spot pool, with automatic fallback to on-demand when spot runs dry.
Qovery runs inside your own AWS, GCP, Azure or Scaleway account, or your existing Kubernetes cluster, so you configure spot and on-demand node pools once and get the production hygiene by default. The cloud bill and any discounts stay in your name.
I have watched teams flip a cluster to spot on a Friday to chase a cheaper bill, then spend the next week explaining a traffic blip to their customers. I have also watched six-person teams run most of their production fleet on spot for months without a single page. The difference was a handful of decisions made before anything moved. This is the playbook I give those teams.
Can you actually run production workloads on spot instances?
Yes, plenty of teams do, but only for workloads that survive a node disappearing with a short warning, running on a diversified fleet with an on-demand baseline underneath. Spot is a capacity contract, not a discount flag you toggle.
Spot is the same hardware your on-demand instances run on. The cloud lends you spare capacity at a steep discount and takes it back when it needs it. What differs is how much warning you get and what the provider calls it:
AWS EC2 Spot gives a two-minute interruption notice before reclaiming the instance, at discounts the AWS Spot page states as up to 90% off on-demand.
GCP Spot VMs offer discounts of up to 91% off on-demand for many machine types, with a short preemption signal that is best effort and up to around 30 seconds by default, configurable up to 120 seconds.
Azure Spot VMs give a 30-second eviction notice and let you pick a deallocate (default) or delete eviction policy, with variable pricing tied to spare capacity.
Scaleway offers interruptible instances on the same principle for teams running in the EU.
Those notice windows are the detail most articles get wrong. Two minutes on AWS versus roughly 30 seconds on GCP and Azure changes how aggressively your app has to shut down, so verify the window for your cloud before you tune anything.
The real risk is not the discount, it is correlated interruption: one instance type in one zone running out of spare capacity at the same moment for everyone using it. If your whole fleet is m5.large in a single AZ, a crunch takes all of it at once. That is the nightmare readers picture, every pod evicted mid-traffic, and it is exactly what the three conditions prevent.
Keep the frame for the rest of this piece. Spot is safe when all three hold: interruption tolerance in the workload, fleet diversification across instance types and zones, and an on-demand baseline underneath. Miss one and you are gambling.
Which workloads are safe on spot, and which should stay on on-demand?
Anything stateless, replicated and restartable belongs on spot. Anything that is a singleton, holds state on local disk, or cannot checkpoint belongs on on-demand or reserved capacity.
The decision is mostly about what happens when a random pod vanishes right now. If another replica picks up the slack and nobody notices, it is a spot candidate. If it would page someone, it needs a PodDisruptionBudget and an on-demand floor before it goes anywhere near spot.
The grey zone is the stateful middle: Redis, Kafka, Elasticsearch. Replication factor decides it. A three-node Kafka cluster with a replication factor of three and rack awareness can tolerate losing a broker on spot. A single Redis node holding your only copy of session state cannot. Judge each by whether the data survives the node, not by the logo on the box.
For a small team, the cheapest easy win is to go in order of risk: non-production environments on spot first, then async workers, then stateless production services. You bank most of the savings before you touch anything that faces a customer.
What Kubernetes configuration makes spot interruptions a non-event?
Spot becomes boring once you have multi-replica deployments with PodDisruptionBudgets, topology spread across zones and instance families, graceful shutdown inside the app, and a controller that drains nodes on the interruption notice before the cloud reclaims them.
Here is the concrete checklist, grouped by layer:
Node level. Run something that catches the interruption signal and drains the node before it dies. The AWS Node Termination Handler does this for plain EC2 Spot. Karpenter handles interruptions natively by watching an EventBridge-to-SQS queue, so you do not also need the termination handler when Karpenter is managing the nodes. GKE and AKS have their own equivalents. The job is the same everywhere: cordon and drain on notice so pods reschedule gracefully instead of being killed.
Pod level. Run replicas: 2 at minimum, ideally 3. Add a PodDisruptionBudget with maxUnavailable so a drain can never take your whole service. Add topologySpreadConstraints across topology.kubernetes.io/zone and node.kubernetes.io/instance-type so a single zone or instance family losing capacity never takes more than its share of your pods.
Application level. Handle SIGTERM, stop accepting new work, finish in-flight requests, then exit. Set terminationGracePeriodSeconds below the interruption window (comfortably under two minutes on AWS, tighter on GCP and Azure). Get readiness and startup probes right so traffic only reaches pods that can serve it, and make sure the load balancer drains connections.
Scheduling. Put taints and tolerations plus affinity on the capacity-type label (karpenter.sh/capacity-type, or your cloud's equivalent) so only workloads you have opted in ever land on spot. This is the switch that keeps your database off spot by accident.
Queue and job hygiene. Make consumers idempotent, set sane visibility timeouts and retries, and checkpoint long jobs so a lost worker costs minutes, not hours.
Fleet diversity. Allow many instance families and sizes across all zones. This is the single biggest lever on your interruption rate. A pool that can run on 20 instance types almost never runs dry all at once.
Observability. Track interruption events, pod restart counts and the share of capacity on spot. When something wobbles you want to tell a spot problem from an app problem in seconds.
How do Karpenter, Cluster Autoscaler, Spot Ocean and Qovery compare?
Cluster Autoscaler with mixed-instance node groups is the lowest common denominator, Karpenter picks and consolidates instances far better on AWS, Spot Ocean and Cast AI add commercial prediction and rebalancing for a cut of the savings, and a developer platform like Qovery sets up the node pools plus the application-side hygiene for you across clouds.
Approach
Clouds
Interruption handling
On-demand fallback
Instance diversity
Consolidation
Who operates it
Pricing
You still build
Cluster Autoscaler
All
Via node termination handler
Separate node group
You curate type lists
No
You
Free (OSS)
Instance lists, zones, drain, app hygiene
Karpenter
AWS (others early)
Native (EventBridge/SQS)
capacity-type NodePool
Broad, automatic
Yes
You
Free (OSS)
NodePools, app hygiene, upgrades
EC2 ASG mixed instances / Spot Fleet
AWS
Via termination handler
On-demand base %
Allocation strategy
No
You
Free (AWS)
Allocation tuning, app hygiene
Spot by NetApp (Ocean)
AWS, GCP, Azure
Managed, predictive
Automatic
Automatic
Yes
Vendor + you
Share of savings
App hygiene, integration
Cast AI
AWS, GCP, Azure
Managed, predictive
Automatic
Automatic
Yes
Vendor + you
Share of savings
App hygiene, integration
GKE Spot / EKS Auto Mode / AKS spot pools
Respective
Managed
Node-pool config
Provider-dependent
Partial
Cloud + you
Included
App hygiene, limits per provider
Qovery
AWS, GCP, Azure, Scaleway, BYO K8s
Built into node pools
Built in
Node-pool config
Via underlying scaler
Qovery + your account
Platform fee, your cloud bill
Mostly configuration
A few honest notes. Cluster Autoscaler works everywhere and is fine if you already run it, but you curate the instance-type lists and zones by hand and that list rots. Karpenter is excellent on AWS: just-in-time provisioning, broad instance diversity, consolidation to reclaim waste, native interruption handling. Its maturity on other clouds lags. And AWS is explicit that the capacity-optimized strategy lowers interruption rates compared with lowest-price, so cheapest is not safest.
Spot by NetApp (Ocean) and Cast AI earn their keep: they predict interruptions and rebalance automatically across clouds for a share of what they save you. The honest critique is operational, not technical. They are one more control plane for a six-person team to own.
Qovery is where I have a stake, so I will be specific. You configure spot and on-demand node pools per cluster, inside your own cloud account or an existing Kubernetes cluster, with standard Kubernetes underneath. On top you get the defaults that make spot safe: multi-replica rolling deploys, preview environments per pull request, non-production auto-stop, managed upgrades and per-environment RBAC. A small team gets the guardrails without building them, and because it is bring-your-own-cloud, the bill plus any Savings Plans or committed-use discounts stay in your account.
The DIY answer wins when you have strong in-house Kubernetes expertise, unusual scheduling needs, or a Karpenter setup that already works and that someone enjoys operating. If that is you, keep it.
Cut compute costs without babysitting your cluster.
Qovery runs spot and on-demand node pools in your own AWS, GCP, Azure or Scaleway account - or your existing Kubernetes cluster - with multi-replica deploys, preview environments and auto-stop built in. Start deploying in under 10 minutes.
How much do you actually save, and what does it cost in engineering time?
Spot cuts the compute line of your bill substantially, with published discounts up to roughly 90%, but the savings only land if compute is a big share of your bill and you do not burn an engineer babysitting the setup.
Start with the arithmetic. Spot only touches compute. It does nothing for managed databases, egress or storage. If compute is 60% of your bill and you move most of it to spot, the blended effect is real. If compute is 20% because you pay mostly for a managed database and data transfer, spot is a rounding error and you should look elsewhere first.
The pattern I see work is roughly 70/30: the stateless majority on spot, a base of system and stateful components on on-demand. The blended saving lands well below the headline 90% because of that floor, but it is still the single largest compute lever most teams have.
The cost side is engineering time, and it is easy to undercount. Someone curates instance-type lists, tunes interruption behaviour, chases config drift, keeps the upgrade cadence, and absorbs the on-call noise on a bad spot day. I will not put a precise FTE number on that, but it is not zero, and for a six-person team it competes directly with shipping features. That trade is the whole reason managed options and platforms exist.
Stack spot with the other levers instead of treating it as the only one. Rightsizing, consolidation, non-production auto-stop at nights and weekends, and Savings Plans or committed-use discounts for the on-demand base all compound. Spot and commitments are not rivals: spot handles elastic stateless capacity, commitments cover the steady base, and under bring-your-own-cloud both stay in your own account.
What is a safe rollout plan to move production onto spot?
Roll out in stages, and only move to the next one after you have watched real interruptions happen without a page. Here is a checklist you can copy, each line standing on its own.
Stage 0 - Measure. Tag and label workloads. Find your current compute spend as a share of the total bill. Write down SLOs and error budgets so you can tell later whether spot hurt anything.
Stage 1 - Non-production. Move dev, staging and preview environments to 100% spot. Lowest risk, immediate savings, and it teaches you your cloud's interruption behaviour for free.
Stage 2 - Async workers. Move queue consumers, batch jobs and CI runners to spot. Confirm consumers are idempotent and jobs retry cleanly first.
Stage 3 - One production service. Pick one stateless service. Give it 3 replicas, a PDB, topology spread, and on-demand fallback. Watch it for a full week, including a deploy and at least one real interruption, before you expand.
Stage 4 - Expand. Move the rest of the stateless fleet. Keep an on-demand floor for system components and stateful sets. Alert on interruption rate and on fallback-to-on-demand events.
Rollback ready. Before Stage 3, have a single label or node-pool change that moves any workload back to on-demand within minutes. Test it once so you trust it under pressure.
When should you not use spot instances in production?
Skip spot when the workload cannot be interrupted at all, when capacity for your required instance family is thin in your region, or when the engineering time to make it safe costs more than the savings.
The hard no cases: single-replica stateful services, primary databases on local disks, real-time workloads with strict latency SLOs that would breach during a crunch, licence-per-instance software, and workloads needing a specific scarce GPU type. For primary databases especially, do not let anyone talk you into spot. The savings are not worth the one bad night.
There is also a capacity reality check. Niche instance families and specific GPU types in a single zone are exactly where spot availability bites. Check the provider's placement and eviction data for your family and region first. Azure exposes historical eviction rates per SKU, and AWS publishes the Spot Instance Advisor. If the data looks thin, diversify harder or stay on-demand.
Finally, compliance or contractual commitments sometimes mandate dedicated or reserved capacity, and that overrides any cost argument. And if compute is a minor share of your spend, do the easy wins first: auto-stop for non-production and rightsizing beat a spot migration you do not need.
I spend a whole section on this because spot is oversold. The mature answer is almost never all-spot. It is a mixed fleet, chosen deliberately, with the three conditions satisfied for every workload that moves.
Is it safe to run production workloads on AWS Spot Instances?
Yes, for workloads that tolerate a two-minute interruption notice, run on a diversified fleet of instance types and zones, and sit on top of an on-demand baseline. Stateless replicated services with PodDisruptionBudgets and a node termination handler are the clearest fit. Primary databases and stateful singletons are not.
How much warning do you get before a spot instance is reclaimed on AWS, GCP and Azure?
AWS gives a two-minute interruption notice. Azure gives a 30-second eviction notice. GCP gives a short best-effort preemption signal, up to around 30 seconds by default and configurable up to 120 seconds for Spot VMs. Always verify against the provider's own docs, since this is the detail most guides get wrong.
Which Kubernetes settings do I need before putting production pods on spot nodes?
At minimum: two or more replicas, a PodDisruptionBudget with maxUnavailable, topologySpreadConstraints across zone and instance type, SIGTERM handling with a tuned terminationGracePeriodSeconds, correct readiness probes, taints and tolerations on the capacity-type label, and a node termination handler that drains the node on the interruption notice.
Should I use Karpenter or Cluster Autoscaler for spot instances?
On AWS, Karpenter is usually the better choice: broader instance diversity, consolidation, and native interruption handling. Cluster Autoscaler is the safe cross-cloud default and fine if you already run it, though you maintain the instance-type lists yourself. Either way, the application-side hygiene matters more than which scaler you pick.
Can databases run on spot instances?
Primary databases should not run on spot. A singleton holding your only copy of state on local disk cannot survive the node being reclaimed. Replicated data stores with a replication factor of three or more and rack awareness can tolerate losing a node, but keep the quorum on on-demand capacity. Managed database services are the simpler answer for most teams.
How does Qovery handle spot and on-demand node pools in my own cloud account?
Qovery lets you configure spot and on-demand node pools per cluster inside your own AWS, GCP, Azure or Scaleway account, or an existing Kubernetes cluster, with standard Kubernetes underneath. It layers on multi-replica rolling deploys, preview environments per pull request, non-production auto-stop, managed upgrades and per-environment RBAC, so the production hygiene spot needs comes by default. Because it is bring-your-own-cloud, the bill and any discounts stay in your name.
Romaric founded Qovery to make Kubernetes accessible to every engineering team. He writes about platform strategy, developer experience, and the future of cloud infrastructure.
Next step
Cut compute costs without babysitting your cluster.
Qovery runs spot and on-demand node pools in your own AWS, GCP, Azure or Scaleway account - or your existing Kubernetes cluster - with multi-replica deploys, preview environments and auto-stop built in. Start deploying in under 10 minutes.