Webinar · Oct 20: The migration takes 2 weeks. Deciding to do it takes 6 months.

DevOps Team Extension: Qovery engineers watching your infrastructure every month

DevOps Team Extension gives growing engineering teams a named Qovery engineer, a continuous infrastructure scan, and a monthly remediation plan, so they can scale reliability, security and compliance without hiring a DevOps team first.

Romaric Philogene
CEO & Co-founder
OCT 10, 2026 · 7 MIN
DevOps Team Extension: Qovery engineers watching your infrastructure every month

Most teams on Qovery start with the same goal: ship product without building a platform team first. That works for a long time. Then the company grows, the number of services doubles, a few customers ask for a SOC 2 report, someone opens a database to an ETL tool "just for now", and the CTO starts asking whether they really need to hire a DevOps engineer now.

Many of the teams asking that question already run on Qovery and are happy with it. What they ask us next is whether we can help them further, with something like professional services. That is why we built DevOps Team Extension. It puts Qovery's own engineers on your infrastructure on a fixed rhythm: we scan it continuously, we tell you what is risky before it breaks, and we help you fix it. You keep the speed of a small team, and you get the judgment of people who run Kubernetes in production every day.

Qovery · Agentic Infrastructure Platform
A control plane for platform teams and their coding agents
Learn more

Key points:

  • DevOps Team Extension is a recurring service, on top of the Qovery platform. A named Qovery engineer reviews your infrastructure every month and works with your team on the fixes.
  • It starts with a free assessment of your Qovery organization: around 150 controls across reliability, security, performance, delivery and cost. It is read-only, and you keep the report whether or not you sign up. We change nothing without your sign-off.
  • Every finding is ranked by risk and tied to your own claims. If your website says SOC 2 or GDPR, we flag where the configuration does not back that up yet.
  • The work is bounded and predictable. A monthly review, working sessions every two weeks, a short list of priorities per quarter, and a score you can track over time.
  • It is built for teams approaching their first DevOps hire, and for teams under reliability, security or compliance pressure that do not have the in-house expertise yet.

Why proactive

Support answers the question you ask. The risky parts of an infrastructure are usually the ones nobody asks about, because nothing has broken yet.

A few examples we see often on growing teams:

  • A production service runs a single replica, and nobody decided that on purpose. It was the default when the service was created two years ago.
  • A database is publicly reachable because an external tool needed access once.
  • The same credential is shared between production, development and every preview environment.
  • Alerting is wired for the cloud provider and the APM, but nothing watches the Kubernetes workload layer, which is where most single-replica failures happen.
  • Critical databases run on burstable instance types that throttle when CPU credits run out, usually at peak traffic.

None of these show up as a support ticket until they cause an incident. DevOps Team Extension finds them while they are still a configuration change.

What you get

A continuous infrastructure scan. We run the Qovery assessment against your organization and keep it current. You can see the results at any time, ranked by severity, with the evidence behind each finding and the recommended fix.

A named engineer. One Qovery engineer owns your account and knows your setup. A second engineer covers absences, so context never sits with one person.

A monthly review (45 minutes). Your engineer walks through what changed, what the scan flagged, and what matters most. We agree on the priorities together. The goal is a handful of concrete wins each quarter that your team can absorb.

Working sessions every two weeks (2 hours). Time with your engineer on the report, or on whatever your team brings: a pre-flight review before a structural change, the production rollout of a new service, a database migration, an architecture question.

Business-hours escalation. Anything the scan flags gets priority, and you have a direct line to engineers who already know your infrastructure.

A quarterly posture review. We re-run the assessment and show how the score moved, pillar by pillar. It doubles as material for your board, your customers' security questionnaires, or your auditor.

What the assessment looks at

The assessment reads your Qovery organization through the API, with GET requests only. It inventories clusters, environments, services, databases, variables and deployment history, then scores five pillars:

PillarExamples of what we check
Reliability and resilienceReplica counts, health probes, availability-zone spread, stateful singletons, backups, disaster recovery readiness
Security and data protectionPublic databases and services, secrets stored as plain variables, credentials shared across environments, SSO, API token scope, encryption at rest
Delivery and operationsAlerting coverage, deployment failures, changes made outside Terraform, preview environment hygiene, drift between running code and branches
Performance and scalabilityAutoscaling headroom, resource requests against real usage
Cost efficiencyOver-reserved CPU and memory, idle environments, spot capacity, scheduling

Findings are mapped to the frameworks you claim publicly (SOC 2, GDPR, ISO 27001, HIPAA, DORA). An open database is ordinary technical debt at a company that claims nothing. At a company that claims SOC 2, it contradicts a published statement, and a customer's security team or an auditor will read it that way.

Some checks need answers only your team has, such as recovery objectives or which endpoints are meant to be public. The report marks those as undetermined, and they go on the agenda for the first review.

How an engagement runs

One engagement, from the first assessment to the day-90 decision
Diagram source (Mermaid)
flowchart LR
    A["Assessment<br/>read-only"] --> B["Joint review<br/>with your team"]
    B --> C["90-day plan<br/>concrete wins"]
    C --> D["Monthly review<br/>+ working sessions"]
    D --> E["Re-run at day 90<br/>you decide"]
    E -. continue .-> D

1. Free kick-off assessment. We run the full assessment and send you the report right away. The first one is free, and the report is yours to keep even if you stop there.

2. Joint review. We go through the findings with your team. This is where the report gets corrected: context the API cannot see, decisions that were deliberate, risks you choose to accept. Every finding ends the session with your team's position attached.

3. First 90 days. We agree on a handful of concrete wins, usually the highest-risk items with the lowest effort, plus one structural topic. A typical list looks like this: move critical databases off burstable instances, close a public endpoint, enable production alerting, isolate credentials per environment, and review production services one by one for replicas, volumes and zone spread.

4. Re-run and decide. At day 90 we re-run the assessment and compare scores. You decide whether to continue.

Get a free assessment of your infrastructure
The first assessment is on us: a read-only review of your Qovery organization, then a session with your team to go through the findings.

What is not included

We want the scope to be clear before you sign anything:

  • No 24/7 on-call. Escalation is during business hours.
  • No application code. We work on infrastructure, configuration and architecture.
  • Nothing outside the perimeter Qovery manages for you.
  • No unlimited hours. Working time is scheduled, and larger projects are quoted separately.
  • No change to your infrastructure without your explicit sign-off.

Why not just ask an AI?

It is a fair question. Models can read a Helm chart, explain a Kubernetes error and draft a Terraform module in seconds. We use them every day: the assessment behind this offer runs with AI, and the Typesense test cluster in our case study was deployed with our own AI skill for Qovery.

The hard part is knowing whether the answer is right for your system. An AI answers the question you ask, with the context you give it, and most infrastructure mistakes come from context nobody thought to give.

The Typesense session shows how this plays out. Ask a model how to make Typesense highly available and it will describe a three-node Raft cluster, which is what we built first. That is a correct answer to the question. With the chart defaults, a rolling restart still lost quorum. Peer hostnames over 64 characters, a bootstrap race with Karpenter and pod IP churn in Raft each needed a fix we only found by breaking the cluster on purpose. And the customer's real constraint turned out to be memory, which a replicated cluster triples. We learned that by asking how their application behaves when search is down, a question no prompt had raised.

Our solutions architects have run production infrastructure for more than ten years, and they use AI heavily. You get both: AI for speed, and an engineer who has seen the failure before to check the answer, ask the missing question, and stand behind the recommendation.

Who it is for

DevOps Team Extension fits best when:

  • Your team is growing, the number of services is going up, and you are starting to ask whether to hire a DevOps or SRE engineer.
  • You have reliability goals for the next quarter but nobody whose job is to own them.
  • Customers or prospects are sending security questionnaires, or you are preparing for SOC 2, ISO 27001 or a similar audit.
  • You failed an audit, or a customer's security review, and need to close the gaps quickly.

It is a poor fit if you already have a platform team that owns these topics day to day. In that case the Qovery platform and our standard support are usually enough.

The math

A mid-level DevOps or SRE engineer costs a growing company somewhere between $90k and $120k a year fully loaded, before recruiting time and ramp-up. DevOps Team Extension costs a fraction of that, and gives you two senior engineers who already know your infrastructure from day one.

The goal is to let you keep growing without that hire until you need it. And when you do hire, that person inherits a clean, documented setup with a clear history of what was fixed and why.

What it looks like in practice

We recently ran DevOps Team Extension with a fast-growing B2B SaaS company running on three EKS clusters. The assessment, the review session with their team, and a Typesense architecture session that changed the plan halfway through are written up in detail here: Inside a DevOps Team Extension engagement: an infrastructure review and a Typesense architecture session.

If your team is at the point where infrastructure questions are starting to pile up, talk with us and we'll run your first assessment for free. You get the report and a review session with your team, and then you decide whether DevOps Team Extension is the right fit.

Romaric Philogene
About the author
Romaric Philogene

Romaric founded Qovery to make Kubernetes accessible to every engineering team. He writes about platform strategy, developer experience, and the future of cloud infrastructure.

Next step

Get a free assessment of your infrastructure

The first assessment is on us: a read-only review of your Qovery organization, then a session with your team to go through the findings.