Webinar · Oct 20: The migration takes 2 weeks. Deciding to do it takes 6 months.

Inside a DevOps Team Extension engagement: an infrastructure review and a Typesense architecture session

How Qovery assessed a growing B2B SaaS company's infrastructure across about 150 controls, what the review with their team changed, and why a Typesense high-availability project turned into a sharding plan.

Romaric Philogene
CEO & Co-founder
OCT 10, 2026 · 12 MIN
Inside a DevOps Team Extension engagement: an infrastructure review and a Typesense architecture session

This is a write-up of the first weeks of a DevOps Team Extension engagement with one of our customers. The company asked us not to name them, so we won't. Everything else is real: the findings, what their team pushed back on, what we got wrong, and the architecture session where the plan changed halfway through.

Key points:

  • The customer is a fast-growing B2B SaaS company running three EKS clusters on AWS with Qovery. Their engineering team was growing, and they wanted to know whether Qovery could help them with more than the platform.
  • We assessed their Qovery organization against about 150 controls. The platform decisions were sound. The gaps concentrated in three places: nothing watched the Kubernetes workload layer, very little was redundant, and the boundary between production and everything else was porous.
  • The joint review corrected the report. Three findings were wrong or incomplete, four new issues came up that no configuration check could have found, and the most useful outcome was a finding the report did not contain at all.
  • The Typesense project changed direction. We built and tested a three-node high-availability cluster, then learned in the working session that their real constraint was memory. The recommendation became logical sharding across smaller instances.
Qovery · Agentic Infrastructure Platform
A control plane for platform teams and their coding agents
Learn more

The context

The customer runs a B2B SaaS product with a growing engineering team. They had been on Qovery for a while, on three EKS clusters in separate AWS accounts, with Karpenter managing nodes. More services were landing in production every month, and more questions were coming to us through support.

Two things made this the right moment for a proactive review.

First, their leadership had started asking a question we hear from many growing teams: do we really need to hire a DevOps engineer right now? They were happy with Qovery and wanted to know if we could help them further, with something like professional services. Their priority for the quarter was already reliability and stability, but nobody on the team owned it full time.

Second, in mid-September a routine node update took down their search service. Typesense ran as a single pod, and when Karpenter replaced the node underneath it, search went down with it. As a stopgap, we froze node rotation on the node pool that hosts it. That stopped the outages, but it came at a cost: AMI and security patches were delayed, Kubernetes upgrades were blocked, and a plain node crash would still take search down. The freeze could only be temporary.

The first weeks of the engagement
Diagram source (Mermaid)
flowchart LR
    A["Mid-September<br/>node update takes search down,<br/>node rotation frozen"] --> B["21 September<br/>read-only assessment"]
    B --> C["22 September<br/>joint review,<br/>90-day plan agreed"]
    C --> D["8 October<br/>Typesense session,<br/>HA replaced by sharding"]

Step 1: the assessment

We started with a read-only assessment of their Qovery organization. It uses GET requests only and changes nothing. It inventoried clusters, environments, services, databases, variables, deployment logs and audit events, then scored about 150 controls across five pillars: reliability, security, performance, delivery and cost.

The overall score came out at 54 out of 100. Reliability and security were measured on a large number of checks, so their scores were solid measurements. Performance and cost rested on very few checks, because most of them need cluster metrics that were not available on every cluster. We reported those two as unmeasured rather than low.

What was already strong

The team had made good platform decisions, and made them on purpose:

  • One AWS account per tier, each with its own role-based credential. No static access keys.
  • IMDSv2 enforced everywhere, every managed database disk encrypted, no SSH keys on any cluster.
  • Every production container image pinned to an explicit version.
  • The public API restricted at the ingress to Cloudflare ranges and a few named IPs.
  • Good configuration hygiene: service-to-service wiring built on Qovery's built-in variables and aliases, so it survives an environment clone, and no credential-shaped value in any plain variable.

Where the gaps were

The gaps concentrated in three places.

Nothing was watching the workload layer. There were no alert receivers and no alert rules on the Kubernetes side. CloudWatch covered AWS infrastructure and an APM tool covered the application, but if a pod crash-looped at night, nobody would know until a customer noticed. One production cluster had no metrics stack at all.

Almost nothing was redundant. Most production services ran a single replica, availability-zone spread was off everywhere, and several singletons held a persistent volume with no backup. Typesense was the most visible case, but not the only one.

The boundary between production and the rest was porous.

  • A production PostgreSQL database, the company's data lake, was publicly reachable. A CIDR allow-list was in place, which is a real mitigation, but the allow-list was the only thing between the internet and the password.
  • One credential was byte-identical across production, development and every preview environment.
  • An environment running in production mode lived on the development cluster.
  • Internal tools, including a BI dashboard, were exposed publicly.
  • All production databases ran on burstable (T-type) instances, which throttle when CPU credits run out, typically at peak.

The report also checked findings against the company's public claims. Their website states SOC 2 and GDPR compliance. Several findings that would be ordinary engineering debt elsewhere contradicted a published statement here, and the report marked them that way.

The recommended order was short: close the public database endpoint, turn on alerting, then isolate credentials per environment (about two hours of work that removes a whole class of exposure).

Step 2: the review session

We then spent 90 minutes going through the report with a member of their leadership team and their lead engineer. On our side were the engineer named on their account, their customer success manager, and me.

We treat the report as a draft until the team has answered it, because the API sees configuration and cannot see intent.

What the team corrected

We got three things wrong. The revised report says so next to each finding, and nothing from the first version was deleted, so the two can still be compared row by row.

  • Preview environments. The report said five preview environments had been broken for about 70 days and nobody had noticed. The team knew. They kept them on purpose, because the only way to delete them ran a Terraform destroy against cloud resources the previews shared. They were avoiding a known hazard.
  • Alerting. The report said "nothing alerts on anything". That was too strong: CloudWatch and the APM tool were wired in. The real gap was alerting on the Kubernetes workload layer, which is still where every single-replica finding lives, so we kept the severity and rewrote the finding.
  • Terraform. The report said an environment deletion could destroy infrastructure. It already had, and the root cause was bigger than we thought: every preview environment provisioned cloud resources under identical names, so previews were not isolated at the cloud layer at all.

What surfaced that no check could find

Four new issues came out of the conversation. In one of them, the team explained that the data lake database had been modified directly in the AWS console. The team needed disk autoscaling, which Qovery-managed RDS did not offer at the time, so they turned it on in AWS. That made sense, but it also detached the instance from Qovery. The CIDR allow-list they had applied at the environment level might never have reached it. That one went on the list for a follow-up check together.

What the team rated higher than we did

Two items. The first was the internal tools exposure: their leadership called an authenticated access portal for internal tools a foundational unlock. The technical change is small. The hard part is giving a non-technical audience access without opening the tools to the internet. We suggested Cloudflare Access in front of the tools, with Google Workspace SSO, so nobody needs a bastion or port-forwarding. The second was moving the burstable databases to non-burstable instances, starting with the data lake and the contacts database, with a read replica first to avoid downtime.

What the team accepted as a risk

No staging environment. The team reads staging as a customer acceptance tier they have no demand for. That is a legitimate decision, and the report now records it as one.

The most useful outcome

The report counted every single-replica production service as a problem. The team's answer was that some of those services are the core of the product and some are internal experiments that are fine on one replica, and that nothing in the platform tells them apart.

The finding underneath the replica count was the absence of a per-service criticality classification. Their lead engineer started walking through production services one by one to set it, and every reliability item in the roadmap now depends on that pass. No automated check would have produced that conclusion. It came from putting the report in front of the people who built the system.

The first 90 days

After the session, we agreed on five concrete wins for the first 90 days:

  1. Typesense: trial a high-availability cluster, with no data migration.
  2. Move the data lake and contacts databases off burstable instance types.
  3. Review production services one by one: replicas, volumes, zone spread, health checks.
  4. Clean up the obsolete preview environments safely (delete the Terraform service with the "don't run Terraform delete" option first, then the environment).
  5. An authenticated portal for internal tools, and production alerting turned on.

The target is a score above 75 on a re-run at day 90.

Want the same review on your infrastructure?
The first assessment is free: a read-only review of your Qovery organization and a session with your team to go through the findings.

Step 3: the Typesense architecture session

Typesense was item one, so we prepared a working session with their lead engineer, our named engineer and our CSM.

What we built before the call

Typesense supports clustering through Raft: one leader and followers, each node with its own disk, writes going through the leader and any node serving reads. With three nodes, losing any one leaves a quorum of two and search keeps working. If search survives the loss of a node, the frozen node pool can rotate nodes again.

We deployed a three-node Typesense cluster on one of our own EKS clusters, on the same stack as the customer (Karpenter, same region), using a Qovery Helm service. The setup:

  • Three replicas with hard pod anti-affinity and zone spread, so each Raft member sits on its own node in its own availability zone.
  • A PodDisruptionBudget of one, so Karpenter can only drain one member at a time.
  • Rolling updates one pod at a time, with a 300-second termination grace period.
  • Peers addressed by stable DNS names through a headless service, with --reset-peers-on-error, so a rescheduled pod with a new IP does not break the cluster.
  • One EBS volume per node, no shared storage.
  • The API key injected from a Qovery secret.

Then we ran a full rolling restart of all three pods, leader included, with a search and a write every 0.5 seconds.

RunConfigurationSearch errors during the restart
1Chart defaultsTwo outages of 23 s and 12 s, quorum lost
2Liveness probe on the peering port12 s
3Plus a longer readiness window6 s, only during the leader change
4Same, with client-side retries (5 retries, 2 s apart)0 failed searches

Follower restarts caused zero errors. The only remaining blip was the leader election when the leader pod itself went away, and standard client retries absorbed it completely for reads. Writes needed a longer retry budget, or a queue that can be replayed.

Three things we learned building it

These apply to anyone running Typesense in cluster mode on Kubernetes:

  1. Peer hostnames must be shorter than 64 characters. Typesense refuses longer names. A peer's FQDN is <release>-N.<release>-headless-svc.<namespace>.svc.cluster.local, and Qovery namespaces are long, so use a very short release name.
  2. There is a bootstrap race with node autoscaling. If peer DNS does not resolve within about 90 seconds, for example because Karpenter is still provisioning nodes, a node gives up on peering for good while its API port stays up. The default TCP liveness probe on the API port never restarts it, so your "three-node" cluster quietly runs on two. Point the liveness probe at the peering port (8107) instead.
  3. Raft tracks pod IPs. A restarted pod comes back with a new IP, and the cluster needs a few seconds to commit the change. Restart the next member too soon and you lose quorum. A readiness probe that requires several consecutive successes (we used 45 seconds) makes rollouts and Karpenter wait long enough.

One more, on our side: Qovery marked the Helm deployment as successful as soon as helm upgrade returned, before the pods were ready. For stateful clusters like this, check health and pod status after the deploy.

What the session changed

We started the session with the HA demo. Then we asked about their setup, and the plan changed.

  • Search was not critical. Their app already fell back to Postgres when Typesense was unavailable, and a dead-letter queue caught failed writes for later replay. When search went down, the product got slower and kept working.
  • Memory was the real constraint. The single node had a 24 GiB memory limit and used about half of it, while using almost no CPU. Storage on disk was small. They could not add search features (more entities, similarity search) without a bigger node.
  • HA does not help with memory. Typesense clustering replicates the full dataset to every node. A three-node cluster means three copies of the index in RAM, roughly tripling memory cost, without adding any capacity. You also need at least three nodes, because two nodes risk split-brain.

For a non-critical feature constrained by memory, paying three times the RAM for redundancy was hard to justify. So we changed the recommendation together.

We drew the options live during the session, screen shared, while we talked. Below is that whiteboard redrawn. It starts from one full 24 GiB instance, shows why a Raft cluster does not add room, then routes each customer to an instance through an assignment table, and ends with many small instances kept around 70% full.

Animated: one full 24 GiB instance, a three-node cluster where every node is full, customers routed by an assignment table to separate instances, then twelve small instances around 70% full with new customers on a new instance and an optional HA cluster
Animated: one full 24 GiB instance, a three-node cluster where every node is full, customers routed by an assignment table to separate instances, then twelve small instances around 70% full with new customers on a new instance and an optional HA cluster

The new plan: logical sharding across small instances

Instead of one large node or a replicated cluster, run several small, independent, single-node Typesense instances, and shard at the application level:

  • Assign each customer to an instance, store that assignment in a reference table, and never recompute it. Hashing a customer ID modulo the instance count works for the first assignment. Recomputing it later would move customers around when you add capacity.
  • Add capacity by adding instances for new customers. Existing assignments stay where they are.
  • Pick the smallest uniform instance size that fits the per-customer data distribution. As an illustration, six instances of about 12 GiB cost less than three of 24 GiB.
  • If some customers need more than degraded search during an outage, give them a small three-node HA cluster of their own and route them there through the same table. Only those customers pay for redundancy.

Their lead engineer suggested a second axis: split by entity type, with one instance for contacts and another for activity data. That is a clean way to unlock new search features one at a time, and the two approaches combine well.

Smaller instances also help with the original problem. A single large node takes minutes to reload its index into memory after a restart. A small instance reloads faster, and losing one affects a slice of customers instead of all of them.

Animated: one 24 GiB node, then a three-node Raft cluster with three full copies of the index, then small sharded instances where a failed instance only sends its own customers to Postgres
Animated: one 24 GiB node, then a three-node Raft cluster with three full copies of the index, then small sharded instances where a failed instance only sends its own customers to Postgres

Migration. Typesense has no native zero-downtime migration path. We had already written a script for the HA test that exports collections, documents, synonym and curation sets, stopwords, presets and aliases through the API, re-imports them with upserts, and compares document counts. It is safe to re-run, and their team will adapt it to the sharding scheme. API keys are never exported by Typesense, so they need to be recreated with the same values.

Warm-up. When a search misses in Typesense, their app reads from Postgres and writes the result back to Typesense. A new, cold instance therefore fills itself through normal traffic, with degraded performance only during warm-up.

How a cold instance fills itself through normal traffic
Diagram source (Mermaid)
flowchart LR
    Q["Search request"] --> T{"Typesense instance"}
    T -- hit --> Res["Results"]
    T -- "miss or unavailable" --> P[("Postgres")]
    P --> Res
    P -. "write back" .-> T

On the Qovery side. Their lead engineer raised a fair concern: managing many separate Typesense services in the Console gets messy. The answer we agreed on is a Helm chart, based on the one we tested, that takes an instance count and deploys N independent instances from a single service, with per-instance resource tuning. We'll build it with them when they are ready to implement.

We also suggested evaluating Meilisearch at some point. It memory-maps its index from disk instead of holding everything in RAM, which addresses the exact constraint they have. They are happy with Typesense's search quality today (they measured it at about ten times better than their previous Postgres GIN indexes in the best case), so this stays an option for later.

What we took away

The report is the start of the conversation. The most valuable finding of the whole engagement (no criticality classification per service) was not in the report. It came from the team explaining why the report's replica count was misleading.

Ask about the application before designing the infrastructure. We prepared a solid HA cluster for a problem that, once we understood how the application handled failures, did not need one. The demo still paid off. It is how we learned what their real constraint was, and the migration script and Kubernetes lessons carried straight over to the new plan.

Detection needs context. A scanner can flag a public database. Only a conversation reveals that it is public on purpose, mitigated, and detached from the platform because of a missing feature. That combination of automated detection and engineers who know the setup is what DevOps Team Extension is built around.

Next steps

  • Their lead engineer is checking whether Typesense offers native shard-key or orchestration options before committing to application-level sharding.
  • We shared the migration script, and we'll help with the migration and build the multi-instance Helm chart when they are ready.
  • The data lake allow-list check and the move off burstable database instances are in progress on their side.
  • Production alerting goes live once a scheduled cluster change lands.
  • At day 90, we re-run the assessment and compare scores.

If your team is growing faster than your infrastructure expertise, DevOps Team Extension starts with exactly this: a read-only assessment and a review session with your team. The first assessment is free. Talk with us to request yours.

Romaric Philogene
About the author
Romaric Philogene

Romaric founded Qovery to make Kubernetes accessible to every engineering team. He writes about platform strategy, developer experience, and the future of cloud infrastructure.

Next step

Want the same review on your infrastructure?

The first assessment is free: a read-only review of your Qovery organization and a session with your team to go through the findings.