Webinar · Oct 20: The migration takes 2 weeks. Deciding to do it takes 6 months.

Replacing NGINX with Envoy Gateway on Hundreds of Clusters: What the Plan Missed

In February we planned to move every Qovery-managed cluster from ingress-nginx to Envoy Gateway by the end of March. It took until October. Here is the rollout model, every class of problem we hit, and what I would do differently.

Benjamin Chastanier
Staff Engineer
OCT 10, 2026 · 12 MIN
Replacing NGINX with Envoy Gateway on Hundreds of Clusters: What the Plan Missed

94% of the clusters Qovery manages have moved from ingress-nginx to Envoy Gateway. NGINX support ended on August 31, the remaining clusters are switching as their owners finish validating, and we plan to remove NGINX, load balancers included, from every cluster by mid-November 2026.

What remains:

Qovery · Agentic Infrastructure Platform
A control plane for platform teams and their coding agents
Learn more
  • Switch the last clusters to Envoy, then delete the NGINX components and their load balancers everywhere.
  • Drop our HTTP 206 compression patch once the upstream fix ships in Envoy Gateway.
  • Find a clean way for customers to layer their own SecurityPolicy, such as OIDC, on top of ours.
  • Remove the GKE workarounds once Google ships the Gateway API features we rely on.

In February I published the plan: move every Qovery-managed cluster from ingress-nginx to Envoy Gateway in four phases, with no downtime, by the end of March 2026. The architecture in that post held up, and the calendar slipped by seven months. This follow-up covers what happened in between across hundreds of clusters, including the incidents we caused ourselves.

Key points

  • The rollout model worked: two edge stacks side by side, a test link per service, and a rollback path at every step.
  • The removal deadline moved twice, mostly because we chose not to force production clusters onto a stack their owners had not validated.
  • Most problems came from three places: Envoy is stricter than NGINX, Gateway API policies replace each other instead of merging, and every Envoy Gateway minor version changed behavior we depended on.
  • The worst incidents came from our own code changes and our own documentation.

Planned versus actual

The February plan against what happened
Diagram source (Mermaid)
gantt
    dateFormat YYYY-MM-DD
    axisFormat %b %y
    section Plan, Feb 2026
    Dual stack on every cluster       :p1, 2026-02-23, 2026-03-09
    Envoy becomes the default         :p2, 2026-03-09, 2026-03-30
    NGINX removed                     :milestone, 2026-04-07, 0d
    section What happened
    Proof of concept                  :a0, 2025-12-24, 2026-02-23
    Dual stack, opt-in then all       :a1, 2026-02-23, 2026-06-15
    Envoy default, cluster by cluster :a2, 2026-05-04, 2026-09-30
    NGINX support ends                :milestone, 2026-08-31, 0d
    Long tail and NGINX removal       :a3, 2026-09-01, 2026-11-15

The dates behind that chart:

  • November 2025: the upstream project announced the ingress-nginx retirement. We compared about ten Gateway API implementations on paper and picked Envoy Gateway.
  • December 24: first proof of concept on EKS, with Envoy Gateway 1.6.1.
  • Late February: new non-production clusters start on Envoy only. New production clusters follow on March 25.
  • March 24: first deadline move. The removal goes from early April to early June.
  • May 4 and 18: Envoy becomes the default for non-production, then production. We did not switch existing production clusters automatically.
  • June 15: every cluster runs the dual stack. Second deadline move: the June removal is not enforced, and support for NGINX ends on August 31 instead.
  • September and October: the long tail. Clusters switch one by one as their owners validate.

How the rollout works in the engine

The whole migration is driven by three cluster settings in the Qovery Engine. Each one is a step, and each step can be undone.

Three cluster settings, four states
Diagram source (Mermaid)
stateDiagram-v2
    direction LR
    NGINX: NGINX only
    Dual: Dual stack
    Envoy: Envoy default
    Removed: NGINX removed
    NGINX --> Dual: deploy_api_gateway
    Dual --> Envoy: use_api_gateway
    Envoy --> Removed: remove_nginx
    Envoy --> Dual: roll back
    Dual --> NGINX: roll back

In the dual stack, two load balancers run in the same cluster, one in front of NGINX and one in front of Envoy, and DNS decides which one gets production traffic.

  • NGINX keeps publishing the cluster's wildcard record through external-dns. Envoy publishes a separate wildcard under new-gateway-api.<cluster domain>.
  • Every route gets a second hostname on that Envoy domain. After one redeploy, each service has a test link that goes through Envoy while production still goes through NGINX.
  • When a cluster switches to Envoy, external-dns stops publishing the NGINX record and starts publishing Envoy's. NGINX stays up so clients with a cached DNS answer still reach the service.
  • The engine can roll back at each step, down to moving cert-manager certificate ownership from Gateway API ListenerSet objects back to Ingress objects.

The three phases: dual stack, default switch, NGINX removed
The three phases: dual stack, default switch, NGINX removed

I would keep this model exactly as it is. Because of it, most of the problems below were found on a test link or rolled back in minutes.

What broke

I grouped the problems by cause, because the same cause came back in different forms for months.

1. Our own defaults

On March 17 we shipped a change that flipped the default value of the three Envoy settings to true in code. Production clusters that had never set them explicitly, which was most of them, inherited the new defaults. For about two hours, new deployments on those clusters failed at the router step, and customers told us before our alerts did.

We rolled back and changed the rule. An advanced setting's value is now written explicitly when it is set, never derived from a default in code, and rollouts go by customer tier. Defaults in code had caused incidents before. This one made the rule stick.

2. Certificates

Certificates generated more support requests than anything else.

  • ACME during dual stack. Our first version created the HTTP-01 challenge as an HTTPRoute on Envoy, but the customer's CNAME still pointed at NGINX, so the challenge never arrived. The fix: during dual stack, NGINX issues certificates; once Envoy is the default, Envoy does, and cert-manager only watches ListenerSet objects from that point.
  • The SSL redirect ate the challenge. With a forced HTTPS redirect, the /.well-known/acme-challenge/ request was redirected too, which would have caused a renewal loop every 60 days or so. We added a dedicated rule ahead of the redirect. That rule first had no backend, so Envoy answered it with a synthesized 500. It now reuses the route's own backends.
  • An expired certificate that "worked". One customer's wildcard certificate expired on April 1. Envoy Gateway before 1.8.3 kept serving the listener anyway. The 1.8.3 upgrade in July correctly rejected it, and every hostname on that certificate started returning 404 at once. Envoy was right, since the certificate had expired four months earlier.
  • Custom domains needed a redeploy. A service that was not redeployed after the dual stack had no ListenerSet and no certificate on the gateway, and served the cluster's default certificate after the switch. "Redeploy every exposed service before you flip" became the first line of every support answer.

3. Envoy is stricter than NGINX

Envoy enforces several things NGINX let through, and each difference showed up in someone's traffic.

  • Timeouts. Envoy's default request timeout is 15 seconds; ours on NGINX was 60. Worse, Envoy's request timeout is a hard total, where NGINX's proxy_read_timeout was idle-based. Long CSV exports, artifact downloads and server-sent event streams were cut at exactly 15 seconds. We added cluster and service settings for the request timeout, the stream idle timeout and the maximum stream duration.
  • Headers with underscores. Envoy rejects the whole request by default. A customer's SAP integration sends them. We now drop the header instead of the request.
  • Encoded slashes. Envoy normalized %2F, which broke a collaborative editor's WebSocket URLs. We added settings for escaped slashes and slash merging.
  • No retries by default. NGINX's proxy_next_upstream quietly retried when Karpenter evicted a pod mid-request. Envoy surfaced those as 503s until we shipped retry settings (two retries by default).
  • Range responses. Once compression worked again (see below), Envoy gzipped HTTP 206 range responses while keeping the original Content-Range. NGINX never compressed 206. A CDN in front of one customer truncated their JavaScript and CSS files. We patched Envoy to skip compression for 206 until the upstream fix lands, and our first patch targeted the wrong filter name.
  • HTTP/2 and multi-domain certificates. With a certificate carrying two names, overlapping listeners made Envoy fall back to HTTP/1.1. We fixed it with an ALPN policy on each route's ListenerSet, which led straight to the next problem.

4. Policies replace each other

This was the most expensive lesson, because it kept coming back. In Gateway API, a policy attached to a route or a ListenerSet replaces the policy attached to the gateway. Coming from NGINX annotations, where settings add up, that is easy to forget.

  • In Envoy Gateway 1.8, a ClientTrafficPolicy could only target a Gateway. Version 1.9 added ListenerSet targets, which is what made the per-route HTTP/2 policy possible. On 1.9, that ListenerSet policy replaced the cluster-wide Gateway policy instead of adding to it, so headers with underscores were rejected again. We caught it on staging and now copy the cluster policy into every ListenerSet policy.
  • To fix compression after the 1.8.3 upgrade, we tried a strategic merge on route policies. As a side effect, a customer's 60-second timeout override fell back to 15 seconds. We reverted the same evening and put compression on each route's own policy instead.
  • Only one SecurityPolicy applies per route, the oldest one, with no merging. That blocks customers who want to layer their own OIDC policy on top of ours, and we don't have a clean answer yet.

5. Client IPs

IP allowlists depend on Envoy seeing the real client address, and that broke in two ways.

  • On Scaleway, PROXY protocol v2 was not enabled on the Envoy load balancer at first, so Envoy saw the load balancer's address and denied every request on allowlisted routes.
  • X-Forwarded-For handling is configured differently in Envoy, through a number of trusted hops. Our own documentation mapped the old NGINX setting to 0 trusted hops when the correct equivalent was "unset". A customer followed it in production and their APIs returned 403 for about 30 minutes. Our documentation caused that outage, and we fixed the mapping the same day.

6. Operating Envoy itself

  • Replicas and draining. A customer saw dropped connections because the data plane had too few replicas and no anti-affinity. We added a replica floor, spread pods across zones, set the disruption budget to one pod at a time, and aligned the load balancer's deregistration delay with Envoy's 60-second drain on every cloud.
  • A data plane at zero. After debugging an upgrade, one Scaleway cluster was left with zero Envoy replicas and nothing to scale it back up. The fix was a minimum of two.
  • Logs. The Envoy Gateway control plane logs at info level by default. On one cluster that added up to 21 GB a day. It now logs warnings only.
  • DNS. Envoy Gateway 1.8 moved us to the standard Gateway API CRD channel, which has no TCPRoute or UDPRoute. external-dns, configured to watch them, crashed on startup, and new clusters published no DNS records until we fixed its sources.
  • GKE. On GKE Autopilot, Google controls the installed Gateway API version, and it currently lags the features we use on Envoy Gateway 1.9, starting with ListenerSet attachment. Its CRDs also conflicted with ours, and Autopilot requires at least 500m of CPU for pods with anti-affinity. GKE clusters attach route certificates to the shared gateway instead, through one ReferenceGrant per secret. Those workarounds are the cost of adopting Gateway API early.

7. Upgrading Envoy Gateway is a migration too

We started on 1.6.1, ran a 1.8.0 development build that we patched ourselves to get ListenerSet support early, then moved to 1.8.2, 1.8.3 and 1.9.1. Almost every step changed something we relied on:

  • 1.8.3 started rejecting expired certificates and changed how compression was configured.
  • 1.9 let a ClientTrafficPolicy target a ListenerSet, and a ListenerSet policy replaces the Gateway one. The engine chart now runs 1.9.1.
  • 1.9 also stopped enabling Lua extensions by default. On October 8 our own control plane, which runs behind Envoy and uses a Lua policy, returned errors for about half an hour after the upgrade. We re-enabled Lua explicitly and moved the next edge change to a weekend window.

I now treat every Envoy Gateway minor version like a small migration, with a changelog review and a staging soak before it reaches a customer cluster.

What maps, and what doesn't

NGINX (Ingress)Envoy Gateway (Gateway API)
IP allowlist and denylistSecurityPolicy authorization on client CIDRs
Basic authSecurityPolicy basic auth (bcrypt is not supported)
CORSSecurityPolicy CORS
Sticky session cookieBackendTrafficPolicy consistent hash on the same cookie
Rate limitingBackendTrafficPolicy local rate limit
Proxy timeoutsRequest timeout, stream idle timeout, max stream duration
rewrite-targetHTTPRoute URL rewrite with a regex
Force SSL redirectA 308 redirect route
Custom HTTP errorsResponse override

No direct equivalent today: proxy-body-size, buffering settings, connection limits and configuration snippets. Envoy handles buffering differently, so most of those settings no longer mean anything, but anyone who relied on a snippet had to rethink it.

Helm charts are their own case. Qovery generates routes for the services it deploys, but customers who ship their own Helm charts with Ingress resources had to add an HTTPRoute, and a ListenerSet for custom domains, themselves.

Why the deadline moved twice

Most of the reasons were organizational.

  • We would not force production. Every exposed service had to be redeployed after the dual stack to get its Gateway API routes. Switching a production cluster whose services had not been redeployed would have broken them, so we let customers switch, and offered to check each cluster with them before the flip.
  • Customers needed time to test. Several asked for more time in the first week. A small team running a product in its busiest season cannot validate an edge proxy change on our calendar.
  • Upstream dependencies. Custom domains needed cert-manager 1.20, and Envoy Gateway minor versions kept changing behavior under us.
  • One owner. I was the main owner of the migration and of most support threads about it for eight months, and it showed in response times.

What went well

The customers who validated early had smooth switches.

  • Alan moved more than a hundred services in about a month and wrote about it in detail.
  • One customer compared every route on Envoy against NGINX before switching, found the responses byte-identical, and measured a time to first byte of 192 ms on Envoy against 228 ms on NGINX.
  • The few incidents we had all had a rollback path that took minutes.

What I would do differently

  • Write every setting explicitly. No behavior should depend on a default in code that can change under a running cluster.
  • Test the documentation. We now review a settings mapping table like code, because ours caused an outage.
  • Plan for policy precedence from day one. Decide which object owns each policy, and write the merge yourself, before the first route-level override.
  • Budget the upgrades. We planned the migration to Envoy Gateway, not the four Envoy Gateway upgrades that came after it.
  • Don't run it with one owner. Release and testing were the bottleneck.
  • Change the edge in quiet hours. We now ship edge changes in a weekend window.

Where we are now

NGINX support on Qovery ended on August 31. The remaining clusters are switching now, and NGINX, including its load balancer, will be removed from every cluster once they have. If you still run a cluster on NGINX, the migration guide has every step, and we will check your cluster with you before the switch if you ask.

If you are planning the same migration on your own clusters: run both stacks, give every route a test hostname, and keep the rollback path until the last client has moved.

Benjamin Chastanier
About the author
Benjamin Chastanier

Benjamin is a staff engineer at Qovery focused on infrastructure automation, Kubernetes internals, and building the deployment engine that powers thousands of clusters.

Next step

Agents ship fast. Guardrails keep them safe.

Qovery ensures every agent action is scoped, audited, and policy-checked. Start deploying in under 10 minutes.