Webinar replay: Heroku to AWS in one command, with an agent doing the work.

DevOps Platforms That Work With Any Kubernetes Cluster: The 2026 Stack for a 3-10 Engineer Team

Which DevOps platforms run on any Kubernetes cluster - managed EKS, GKE, AKS, Kapsule, or your own self-managed cluster - compared layer by layer: GitHub Actions, GitLab CI, Terraform/OpenTofu, Argo CD, Flux, Backstage, Humanitec, Render, Railway and Qovery. Plus the minimum viable pipeline, rollback plan, and real cost math for a team of 3-10 engineers leaving a managed PaaS.

Romaric Philogene
CEO & Co-founder
OCT 2, 2026 · 15 MIN
DevOps Platforms That Work With Any Kubernetes Cluster: The 2026 Stack for a 3-10 Engineer Team

Key Points:

  • A DevOps platform works with any Kubernetes cluster if it talks to the Kubernetes API and installs via Helm or an Operator, instead of depending on one provider's proprietary control plane. By that test: GitHub Actions, GitLab CI, Terraform/OpenTofu, Argo CD, Flux, Argo Rollouts, Backstage, Humanitec, Prometheus/Grafana, Datadog and Qovery are portable. AWS App Runner, ECS/Fargate, Azure Container Apps and Google Cloud Run are not - they are good products, single-cloud by design.
  • Pick one tool per layer, not two. The four layers are CI (GitHub Actions or GitLab CI), infrastructure as code (Terraform or OpenTofu), delivery (Argo CD, Flux, or a platform layer), and the developer-facing interface (Qovery, Backstage, Humanitec). Stacking two tools in the same layer is the most common source of wasted effort I see in small teams.
  • Leaving a managed PaaS hands you seven jobs it was doing silently: build pipelines, environment provisioning, TLS and ingress, secrets, autoscaling, observability, and rollback - plus cluster upgrades. Kubernetes ships a new minor version roughly every 4 months with about 14 months of patch support, so upgrades are recurring work you cannot defer.
  • Qovery is cluster-agnostic. It installs into your own AWS, GCP, Azure or Scaleway account, or onto a Kubernetes cluster you already run (self-managed, on-prem, k3s, RKE2, OpenShift). It adds git-push deploys, preview environments per pull request, environment auto-stop for non-production, one-click redeploy of a previous version, per-environment RBAC and managed cluster upgrades - while the cloud bill and your committed-use discounts stay in your own account.
  • Rollback is two separate problems. Kubernetes reverts a Deployment in seconds with kubectl rollout undo against stored revisions (default revisionHistoryLimit is 10). Nothing reverts a dropped column for you, so expand-contract migrations plus a tested point-in-time restore are the real safety net.

I have watched a lot of small teams leave a managed PaaS. They are rarely asking whether to move. They have already decided, usually because of unit cost at scale, data residency, or the need to sit inside a private network next to resources they already run. What they actually want to know is narrower and more useful: which tools run anywhere, what do we run in what order, and what breaks.

Qovery · Agentic Infrastructure Platform
A control plane for platform teams and their coding agents
Learn more

This is the answer a peer who has done it would give you. AWS shows up as the worked example because it is the most common destination, but everything here works the same on GCP, Azure, Scaleway, and on a Kubernetes cluster you already run.

Which DevOps platforms actually work with any Kubernetes cluster?

A tool works with any Kubernetes cluster if it speaks the Kubernetes API and installs through Helm or an Operator, with no dependency on a single provider's control plane. By that test the portable set is GitHub Actions, GitLab CI, Terraform/OpenTofu, Argo CD, Flux, Argo Rollouts, Backstage, Humanitec, Qovery, Prometheus/Grafana, Datadog and Honeycomb. The not-portable set is AWS App Runner, ECS/Fargate, CodePipeline deploy actions targeting AWS-only runtimes, Azure Container Apps and Google Cloud Run - each is a solid product locked to one provider's runtime.

That line is the whole decision. If a tool installs via Helm or an Operator, or authenticates to kube-apiserver with a kubeconfig, it runs on managed clusters (EKS, GKE, AKS, Scaleway Kapsule, OVHcloud, DigitalOcean, Linode) and on self-managed ones (kubeadm, k3s, RKE2, OpenShift, Talos) without changing anything but the cluster endpoint.

The mistake I see most often is treating tools from different layers as competitors. They are not. There are four layers, and most teams need one tool in each:

  • CI: GitHub Actions or GitLab CI. Runs tests, builds the image.
  • Infrastructure as code: Terraform or OpenTofu. Creates the account-level primitives.
  • Delivery: Argo CD, Flux, or a platform layer. Gets the image running in the cluster.
  • Developer interface: Qovery, Backstage, Humanitec. The thing engineers actually touch.

GitHub Actions is CI. Argo CD is delivery. Qovery is the developer-facing platform layer. They sit in different rows and are routinely used together, not instead of each other. A GitHub Actions pipeline that builds an image and hands off to Qovery or Argo CD is a normal, healthy stack.

The single-cloud options are not bad tools. App Runner, ECS/Fargate, Cloud Run and Azure Container Apps are genuinely good if you are all-in on one provider and intend to stay there. The reason portability is worth paying for is that most organizations are not all-in: the CNCF annual surveys have reported for years that Kubernetes is used in production by the large majority of respondents and that running across multiple clouds or on-prem is common, not exotic. If there is any chance you will add a second cloud, serve a region on a different provider, or acquire a company that runs elsewhere, a proprietary runtime becomes the thing you cannot move.

ToolLayerWorks with any Kubernetes clusterHow it installsSingle-cloud lock-in risk
GitHub ActionsCIYes (cloud-neutral; self-hosted runners can run in-cluster)SaaS control plane + runnersLow
GitLab CICIYesSaaS or self-managed + runnersLow
Terraform / OpenTofuInfrastructure as codeYes (provisions any cluster)CLI / binaryLow (you write the modules)
Argo CDDelivery (GitOps)YesHelm / Operator in-clusterLow
FluxDelivery (GitOps)YesHelm / Operator in-clusterLow
BackstageDeveloper portalYesContainer / Helm, you host itLow
HumanitecPlatform orchestratorYesAgent/Operator in-clusterLow
QoveryDeveloper platform layerYes (BYOC cloud account or bring-your-own existing cluster)Installs into your account or existing clusterLow (bill and discounts stay yours)
AWS App RunnerRuntimeNo - AWS onlyAWS console / APIHigh
AWS ECS / FargateRuntimeNo - AWS onlyAWS console / APIHigh
Azure Container AppsRuntimeNo - Azure onlyAzure console / APIHigh
Google Cloud RunRuntimeNo - GCP onlyGCP console / APIHigh

What do you actually lose when you leave a managed PaaS?

A managed PaaS was silently running at least seven jobs for you. On day one in your own cluster you own every one of them: build and release pipeline, runtime scheduling, TLS and ingress, secrets management, autoscaling, observability, and rollback - plus database provisioning and backups, and cluster upgrades. Audit all of them before you move a single service.

Here is the checklist I hand teams. Copy it and put a name next to each line:

  • Build and release pipeline - tests, image build, image push.
  • Runtime scheduling - what restarts a crashed process, what handles a node dying.
  • TLS and ingress - certificates, renewal, routing, DNS.
  • Secrets management - storage, rotation, injection into pods.
  • Autoscaling - horizontal and vertical, plus cluster node scaling.
  • Observability - logs, metrics, traces, and alerting on all three.
  • Rollback - code and data, which are different problems.
  • Database provisioning and backups - plus a restore you have actually tested.
  • Cluster upgrades and CVE patching - recurring, not one-time.
  • On-call - someone answers the page at 2am.

The triggers for the move are usually some mix of unit cost at scale, compliance and data residency, private networking to resources you already run in a cloud account, committed-use discounts and credits you want to apply, and plain vendor lock-in. Heroku, Render, Railway, Fly.io and Platform.sh are all common starting points - do not assume the exit path is the same for each.

The constraint that frames this entire article: a 3-10 engineer team has no spare headcount. Judge every tool on its ongoing operational load, not on how fast the demo deploys. A slick first deploy means nothing if the tool adds a day a week of upkeep.

And the clock has a price attached. Kubernetes releases about three minor versions a year, roughly every four months, with about 14 months of patch support per minor version. The version skew policy means you upgrade the control plane before the nodes and cannot skip minors at will. On EKS, running past standard support is not just risk - it is money: the EKS control plane is $0.10 per cluster per hour on standard support and $0.60 per cluster per hour on extended support (as of early 2026), a 6x jump that lands automatically once a version ages out. Upgrades are not optional and they are not free.

The honest exception: if your footprint is small and there is no cost, residency, or networking driver, staying on a managed PaaS is the correct engineering decision. Moving to your own cluster to "do DevOps properly" with no concrete driver is how a 5-person team loses a quarter.

What does a minimum viable DevOps setup look like for a team of 3-10 engineers?

The minimum viable stack is five layers: Git-provider CI (GitHub Actions or GitLab CI) for tests and image builds, Terraform or OpenTofu for the landing zone, managed or existing Kubernetes as the runtime, an internal developer platform as the interface engineers actually touch, and Prometheus/Grafana or Datadog for observability. One tool per layer, chosen for low upkeep.

Here is each layer in a few sentences.

Layer 1 - source and CI. GitHub Actions or GitLab CI runs your unit and integration tests, builds the container, and pushes the image to ECR, Artifact Registry, ACR or the GitLab Registry. Authenticate to the cloud with OIDC short-lived credentials on GitHub or the equivalent on GitLab, never long-lived access keys sitting in a secret.

Layer 2 - infrastructure as code. Terraform or OpenTofu creates the account-level primitives: VPC, cluster, IAM, managed database. Keep the module count small and boring. This layer should change monthly, not daily. If it changes daily, you have pushed application concerns into it that belong a layer up.

Layer 3 - runtime. Managed Kubernetes (EKS, GKE, AKS, Scaleway Kapsule) or the cluster you already run. Kubernetes wins despite its complexity for three reasons: portability across every provider, the ecosystem of tooling that assumes it, and standard rollback primitives that exist identically on every distribution. You learn the primitives once and they transfer.

Layer 4 - developer interface. This is the layer teams skip, and skipping it is why migrations stall six weeks in. Your engineers need self-service deploys, per-branch environments, and logs without learning Kubernetes internals. The options are Qovery, Backstage, Humanitec, or a homegrown wrapper around Argo CD. Pick one deliberately - do not let "we'll just use kubectl for now" become the permanent answer.

Layer 5 - observability and guardrails. Logs, metrics, traces, plus per-environment RBAC and a deployment audit trail. "Who deployed what, when" needs an answer at 2am, not a shrug.

Two rules worth bolding:

  • Migrate one non-critical service first. Prove the whole pipeline on something that can break safely.
  • Keep the old platform warm until traffic is fully cut over. It is your rollback of last resort.

How do you know the stack is working? Anchor it to DORA's four keys: deployment frequency, lead time for changes, change failure rate, and failed deployment recovery time. If those four are steady or improving after the move, the stack is doing its job.

LayerThe job it doesRecommended toolsWorks on any Kubernetes clusterWho owns it (10-person team)Ongoing ops load
Layer 1: CITest, build image, push to registryGitHub Actions / GitLab CIYesWhole team (it lives in the repo)Low
Layer 2: Infrastructure as codeVPC, cluster, IAM, managed DBTerraform / OpenTofuYesOne infra-minded engineerLow-medium
Layer 3: RuntimeSchedule and run containersEKS / GKE / AKS / Kapsule / existing clusterYesPlatform layer or managed serviceMedium
Layer 4: Developer interfaceSelf-service deploys, preview envs, logsQovery / Backstage / Humanitec / Argo CDYesPlatform layer ownerLow (Qovery/Humanitec) to high (homegrown)
Layer 5: Observability + guardrailsLogs, metrics, traces, RBAC, auditPrometheus+Grafana / Datadog / HoneycombYesWhole team, one leadMedium

How do you get deployment pipelines and automated testing without a dedicated DevOps engineer?

Keep CI in your Git provider and delegate CD to a platform that watches the branch or the registry. Splitting CI (yours) from CD (delegated) is the one decision that keeps the YAML footprint small enough for a team with no platform engineer.

The pattern in one line: CI runs tests and builds one immutable image per commit SHA, then triggers deployment through a platform API, a webhook, or a GitOps commit. CI proves the artifact is good. CD decides where and how it runs. Keep that boundary clean and neither side grows into a monster.

What belongs in CI:

  • Unit tests and integration tests against ephemeral dependencies spun up in the job.
  • Lint and static analysis.
  • SAST and an image vulnerability scan.
  • Nothing that needs production credentials. If a CI job can reach prod, your blast radius just grew for no reason.

What belongs in CD:

  • Environment-specific config and secrets injection.
  • Progressive rollout with readiness and liveness probes.
  • Automatic halt when readiness fails.
  • An audited approval gate for production.

The highest-leverage practice for a small team is a preview environment per pull request. Reviewers open a real URL instead of reading a YAML diff and imagining the result. It catches the class of bug that only shows up when the service actually boots. The one thing to manage is cost: put auto-stop on non-production environments so idle previews are not billing you overnight.

Watch the CI-minutes trap. GitHub Actions includes 2,000 minutes/month on Free, 3,000 on Team, and 50,000 on GitHub Enterprise Cloud for private repositories, then bills per minute with Linux cheapest and macOS roughly 10x. GitLab includes compute minutes per tier and sells more at a per-minute rate. Heavy test suites burn through the included tier fast. Once overage gets annoying, run self-hosted runners on your own cluster - you already have the capacity, and it removes the per-minute meter entirely.

For security, OIDC federation is the documented replacement for static cloud keys (GitHub, GitLab): the pipeline exchanges a short-lived token for cloud credentials at run time, so there is no long-lived key to leak. And the cheapest production guardrail that exists is GitHub Environments with required reviewers and wait timers - a native approval gate that costs nothing to turn on.

How do you make rollbacks actually safe on your own Kubernetes cluster?

Rollback is two independent problems: reverting code and reverting data. Kubernetes reverts code in seconds with kubectl rollout undo against stored revisions (default revisionHistoryLimit is 10). Nothing reverts a destructive migration for you - which is why expand-contract schema discipline and a tested point-in-time restore are the real safety net.

Code rollback is the easy half, and Kubernetes gives it to you for free. Tag every image with the commit SHA, never :latest, so a rollback points at a known artifact. Rolling updates are gated by readiness probes with default maxSurge and maxUnavailable of 25% - a pod that never passes its readiness probe stalls the rollout instead of taking down the service. When something does get through, kubectl rollout undo returns you to the previous revision in seconds. The reason a platform-level "redeploy previous version" button matters is human: the person on call at 2am may not be a Kubernetes expert, and a button they cannot typo beats a command they might.

Data rollback is the hard half because there is no undo button. Use expand-contract migrations: add the new column, backfill, switch the code, and only drop the old column in a later release - never ship a destructive migration in the same release as the code that needs it. When data really is corrupted, point-in-time recovery is the floor: Amazon RDS restores to any second within the retention window, up to the last 5 minutes, with automated backups retained up to 35 days, and Cloud SQL offers the same PITR capability on GCP. Realistic recovery time is minutes to tens of minutes, and it creates a new instance - so rehearse it before you need it.

Three more layers worth having:

  • Traffic-level safety: health checks, blue/green, canary with Argo Rollouts or Flagger, and feature flags - the cheapest rollback of all for a product change, since it is a config toggle, not a deploy.
  • Pre-production safety: ephemeral environments seeded with anonymized, production-shaped data. RepliByte is one open-source option for the seeding.
  • Operational safety: a deployment audit log, per-environment RBAC so production needs an approval, and a written answer to "who can roll back at 2am."

This is the work that moves two of the DORA keys directly: change failure rate and failed deployment recovery time. Rather than invent numbers, use the published DORA benchmark bands to see where your team sits - the gap between low and elite performers on recovery time is measured in orders of magnitude, and this is the work that closes it.

MechanismWhat it revertsTypical time to recoverBlast radiusSetup effort (small team)Works on any Kubernetes cluster
kubectl rollout undoApplication code / container versionSecondsOne DeploymentLow (built in)Yes
Blue/greenCode, with instant traffic switchSeconds to switchOne service, all-or-nothingMediumYes
Canary (Argo Rollouts / Flagger)Code, gradually with auto-analysisMinutesSmall % of traffic firstMedium-highYes
Feature flagsA product behavior changeInstantScoped to the flagLow-medium (needs a flag system)Yes (app-level)
Managed DB point-in-time restoreData to a prior timestampMinutes to tens of minutesWhole databaseLow (managed), must be rehearsedYes (managed DB, not in-cluster)
Full backup restoreData from a backup snapshotTens of minutes to hoursWhole databaseLow to run, high RTOYes
Ship faster on infrastructure you control.
Qovery gives your team self-service deployments on your own AWS, GCP, Azure or Scaleway account - or on the Kubernetes cluster you already run. Start deploying in under 10 minutes.

Which platform combination should a small team pick, and how do the options compare?

No single tool covers CI, infrastructure, delivery, developer experience and observability, so choose one tool per layer. For most 3-10 engineer teams that is Git-provider CI + Terraform/OpenTofu + an internal developer platform on managed or existing Kubernetes. Argo CD replaces the platform layer only when someone on the team genuinely wants to own Helm and Kustomize manifests forever.

Here is the honest trade-off for each option:

  • GitHub Actions - the best default CI if your code is on GitHub. Weak as a standalone deployment platform; pair it with a CD layer. Keep it.
  • GitLab CI/CD and Auto DevOps - the strongest single-vendor pipeline story, CI through CD in one product. Heavier to self-manage if you run GitLab yourself. Keep it as your CI.
  • Terraform / OpenTofu - unmatched for the landing zone. It is not a developer interface and was never meant to be one. Use it for infrastructure, stop there.
  • Argo CD and Flux - excellent GitOps delivery, both CNCF projects (Argo CD, Flux). The honest trade is that your team writes and owns the Helm/Kustomize manifests, forever. Great when one engineer wants that control; a tax when nobody does.
  • Backstage - a developer portal framework you extend with plugins, not a deployment engine. Powerful catalog and discovery; you build and maintain the plugins. Realistic only if you have time to invest in the portal itself.
  • Humanitec - a platform orchestrator, enterprise-oriented and priced that way. Capable, but heavier than most small teams need.
  • AWS-native (CodePipeline + ECS/Fargate/App Runner) - deep AWS integration and a reasonable path if you are all-in on AWS. The trade is single-cloud lock-in and an uneven developer experience across the pieces.
  • Render and Railway - excellent developer experience. The honest trade is that they are still a managed PaaS running in their account, so a PaaS-to-PaaS move does not fix cost control, data residency, private networking, or committed-use discounts.
  • Datadog and Honeycomb - the observability layer. Complementary to everything above, not an alternative to any of it.
  • Qovery - PaaS-grade developer experience inside your own cloud account or your existing Kubernetes cluster: git-push deploys, preview environments per PR, environment auto-stop, one-click redeploy of a previous version, per-environment RBAC, and managed cluster upgrades.

The distinction that matters most for cost is BYOC. When the platform runs in your own account, the cloud bill, the AWS Compute Savings Plans (up to ~66% off on-demand), EC2 Spot (up to ~90% off), and Google Cloud committed use discounts all stay yours. Those levers only exist when the account is yours - a PaaS running in someone else's account cannot hand them to you.

Three combinations I would actually recommend:

  1. GitHub Actions + Terraform + Qovery on managed Kubernetes - for a 5-person product team with no platform engineer who wants PaaS simplicity without giving up their cloud account.
  2. GitLab CI + Terraform + Argo CD - when one engineer owns infrastructure full time and wants total control of the manifests.
  3. AWS-native CodePipeline + ECS/Fargate - when you are all-in on AWS, already have the expertise, and accept single-cloud lock-in as a deliberate trade.

And plainly: if your footprint is small, you have no compliance driver, no cost pressure, and no need for private networking to existing cloud resources, staying on Render or Railway is the right call. Do not migrate to prove a point.

LayerAny K8s clusterRuns in your own account (BYOC)Git-push deployPreview env per PROne-click rollbackWho writes the YAMLOps burden (3-10 team)Pricing modelBest fit
GitHub ActionsCIYesN/A (SaaS CI)Via CD layerVia CD layerVia CD layerYou (workflow YAML)LowPer-minute + included tierDefault CI on GitHub
GitLab CI/CDCI (+CD)YesN/A / self-managedYes (Auto DevOps)Possible, you wire itVia pipelineYou (.gitlab-ci.yml)Low-mediumPer-minute + tierSingle-vendor pipeline
Terraform / OpenTofuInfra as codeYesYesNoNoNo (infra only)You (HCL)Low-mediumOSS / Cloud tiersThe landing zone
Argo CDDeliveryYesYesVia commitAdd-on toolingYes (git revert / rollback)You (Helm/Kustomize)Medium-highOpen sourceGitOps control
FluxDeliveryYesYesVia commitAdd-on toolingYes (git revert)You (Helm/Kustomize)Medium-highOpen sourceGitOps control
BackstagePortalYesYesVia pluginsVia pluginsVia pluginsYou (plugins + config)HighOpen sourceService catalog / portal
HumanitecOrchestratorYesYesYesYesYesMostly abstractedMediumEnterprise subscriptionLarger orgs standardizing
AWS-nativeCI + runtimeNo (AWS only)Yes (your AWS)PartialHard to do cleanlyVia ECS task revisionsYou (AWS configs)MediumAWS usageAll-in on AWS
RenderPaaSNo (their platform)No (their account)YesYesYesMinimalVery lowManaged PaaSSmall footprint, no cost/residency driver
RailwayPaaSNo (their platform)No (their account)YesYesYesMinimalVery lowManaged PaaSSmall footprint, fast start
QoveryDeveloper platformYesYes (your account or existing cluster)YesYesYes (redeploy previous version)MinimalLowSubscription (cloud bill stays yours)PaaS DX on your own infra

What does running your own Kubernetes platform cost in money and engineering time?

The cloud bill is rarely what breaks a small team - the recurring human cost of operating the platform is. A managed control plane is a predictable, published hourly fee. Cluster upgrades, CVE patching and on-call are the line items with no price tag until you hire for them.

One-time costs are front-loaded and finite: containerizing services, writing the Terraform landing zone, the DNS and TLS cutover, and the data migration with its dual-write window. These hurt once.

Recurring infrastructure costs are predictable and, for a small footprint, usually modest (all figures as of early 2026, check the linked pages since they move):

Recurring human costs are the ones that actually decide this. Cluster minor-version upgrades on a ~14-month support clock, CVE patching, IAM reviews, and an on-call rotation. None of these scale down just because the team is small - a 4-person team upgrading Kubernetes spends nearly the same hours as a 40-person team doing it.

The headcount math is where I would be most careful. Rather than quote a salary I cannot verify for your market, check levels.fyi and the Stack Overflow Developer Survey for current DevOps and platform engineer compensation, and read the Linux Foundation research on the tech talent gap before assuming you can hire quickly - the shortage is real and the hire is slow. Compare that fully loaded cost and hiring timeline against a platform subscription. For many 3-10 person teams the subscription wins on both money and calendar.

Cost levers that only exist in your own account: Compute Savings Plans, committed use discounts, Spot or preemptible nodes, environment auto-stop for non-production, and right-sizing from real usage. A PaaS in someone else's account cannot give you these.

For a 5-15 service migration, a realistic timeline is several weeks to a few months of focused effort depending on how stateful the services are. Treat that as an experience-based estimate from teams we work with, not a sourced statistic - your mileage will vary with the number of databases.

Line itemWho pays it on a managed PaaSWho pays it in your own clusterHow often it recursHow a platform layer like Qovery changes it
Managed control plane feePaaS (bundled in price)You (EKS/GKE/AKS hourly)ContinuousUnchanged - still your cloud bill, in your account
Worker nodesPaaS (bundled)You (direct cloud cost)ContinuousUnchanged, but auto-stop + right-sizing trim it
NAT Gateway + egressPaaS (bundled)You (per-hour + per-GB)ContinuousUnchanged - architecture-driven
CI minutesYou (CI vendor)You (CI vendor or self-hosted)Per buildSelf-hosted runners in-cluster remove the meter
Observability ingestYouYouContinuousUnchanged - pick your vendor
Cluster upgradesPaaS (invisible to you)You (engineering hours)~Every 4-14 monthsManaged by the platform layer
CVE patchingPaaS (invisible)You (engineering hours)OngoingManaged by the platform layer
On-callPartly PaaSYouContinuousReduced - fewer things you operate by hand

What is the step-by-step migration sequence that does not break production?

Prove the full pipeline - build, deploy, preview environment, rehearsed rollback - on one low-risk service before you touch anything customer-facing, and move databases last. Everything below is detail on that order.

  1. Inventory. List every service, dependency, environment variable, cron job, background worker, managed add-on, and domain with a TLS certificate. The thing that breaks a cutover is the dependency nobody wrote down.
  2. Containerize. A clean, reproducible local Docker build per service. Pin base images. No build-time secrets.
  3. Landing zone. Terraform or OpenTofu for VPC, cluster, registry, IAM via OIDC, and the managed database - or skip this entirely and point at the Kubernetes cluster you already run.
  4. CI. Tests and image build with immutable tags per commit SHA, pushed to your registry.
  5. First service. Deploy one non-critical service end to end, including a preview environment per pull request and a rollback drill you actually run, not just document. A rehearsed rollback is the only kind that works under pressure.
  6. Stateful services. Databases last, with replication and a measured cutover window. This is the step with no easy undo, so it goes when everything else is proven.
  7. DNS cutover. Lower TTL days ahead so you can move fast, then cut. Keep the old platform warm as a fallback for at least a week.
  8. After cutover. Per-environment RBAC, audit logging, alerting, a restore test you actually perform, and a written on-call runbook.

For a team with no platform engineer, steps 4-5 and 8 are where Qovery earns its place: it wires up git-push deploys, per-PR preview environments, one-click rollback, RBAC and audit on whatever cluster you bring - including one that already exists. You keep CI in GitHub or GitLab, Terraform for the landing zone, and your observability stack. Qovery is only the developer-facing layer, running in your account.

Frequently Asked Questions

Which DevOps platforms work with any Kubernetes cluster, including self-managed and on-prem?

The portable set is GitHub Actions, GitLab CI, Terraform/OpenTofu, Argo CD, Flux, Argo Rollouts, Backstage, Humanitec, Qovery, Prometheus/Grafana, Datadog and Honeycomb. Each either installs via Helm or an Operator in the cluster or talks to the Kubernetes API with a kubeconfig, so it runs on managed clusters (EKS, GKE, AKS, Kapsule) and self-managed ones (kubeadm, k3s, RKE2, OpenShift, Talos) alike. AWS App Runner, ECS/Fargate, Azure Container Apps and Google Cloud Run are good products but locked to a single cloud.

Can I use an internal developer platform on a Kubernetes cluster I already run, instead of letting it create a new one?

Yes. Qovery, Argo CD, Flux, Backstage and Humanitec all install onto an existing cluster rather than requiring one they created. Qovery specifically supports bring-your-own existing cluster alongside its BYOC mode, so you can point it at a self-managed, on-prem, k3s, RKE2 or OpenShift cluster and get git-push deploys, preview environments and RBAC without rebuilding your infrastructure. This is the common case for teams that already have a cluster and just need a developer-facing layer on top.

Do we need Kubernetes to leave a managed PaaS, or is ECS/Fargate or Cloud Run enough?

If you are certain you will stay on one cloud forever, ECS/Fargate on AWS or Cloud Run on GCP is a legitimate, simpler choice. You trade portability for less operational surface. Choose Kubernetes when you want to avoid single-cloud lock-in, need the broader ecosystem of tooling, or value rollback and scheduling primitives that behave identically on every provider. For most teams leaving a PaaS because of cost or residency, the portability is the whole point.

Should we use GitHub Actions or GitLab CI for deployment pipelines, and do we still need Argo CD?

Use whichever matches where your code lives - GitHub Actions for GitHub, GitLab CI for GitLab. Both are strong CI tools and you keep them regardless of what else you choose. You need Argo CD only if you want GitOps-style delivery and someone on the team is happy to own the Helm and Kustomize manifests. If nobody wants that job, a platform layer like Qovery handles delivery instead, and your CI just builds the image and hands it off.

How do we get preview environments for every pull request after leaving a managed PaaS?

Use a platform layer that creates an environment per pull request automatically - Qovery and Humanitec do this out of the box, and you can build it on Argo CD with extra tooling. The reviewer gets a real, running URL for each PR instead of guessing from a YAML diff. Put auto-stop on these non-production environments so idle previews stop billing you, which is the one cost to watch.

How do you roll back a bad deployment safely, including database changes?

Code rollback is fast: immutable image tags per commit SHA plus kubectl rollout undo returns you to the previous revision in seconds, since Kubernetes keeps the last 10 revisions by default. Data is the hard part - use expand-contract migrations so schema changes stay backward-compatible, and never ship a destructive migration with the code that needs it. Keep managed-database point-in-time recovery (RDS or Cloud SQL) as the floor, and rehearse the restore before you rely on it.

How much DevOps time and money does a 3-10 engineer team need to run its own Kubernetes platform?

The cloud bill for a small footprint is usually modest and predictable - a managed control plane is about $0.10 per cluster per hour on EKS or GKE, plus nodes, NAT and egress. The real cost is human: cluster upgrades every few months, CVE patching, IAM reviews and on-call, none of which scale down with team size. Most teams this size come out ahead using a managed platform layer so they keep their own cloud account and discounts without hiring a dedicated platform engineer, which the current skills shortage makes slow and expensive anyway.

Romaric Philogene
About the author
Romaric Philogene

Romaric founded Qovery to make Kubernetes accessible to every engineering team. He writes about platform strategy, developer experience, and the future of cloud infrastructure.

Next step

Ship faster on infrastructure you control.

Qovery gives your team self-service deployments on your own AWS, GCP, Azure or Scaleway account - or on the Kubernetes cluster you already run. Start deploying in under 10 minutes.