How to Deploy Mistral and Other Open-Weight Models Internally: The 4-Layer 2026 Stack for Coding Agents
A layer-by-layer tool stack for self-hosting Mistral, Llama, Qwen or gpt-oss internally and letting Claude Code and Codex call them - with GPU scale-to-zero, per-key budgets, and scoped deploy access. Honest comparison of vLLM, KubeAI, LiteLLM, SkyPilot, Ray Serve, Ollama, CoreWeave, Bedrock, Vertex AI and Qovery.
No single tool deploys open-weight models internally. The working stack has four layers. An inference server (vLLM, SGLang or TGI), a model orchestrator on Kubernetes (KubeAI or Ray Serve), an OpenAI-compatible gateway for keys, budgets and routing (LiteLLM), and a deployment and governance layer that gives humans and agents scoped, audited deploys (Qovery, or Terraform plus CI you maintain yourself). Buy one and expect the other three and you will rebuild them badly.
Claude Code and Codex both point at a self-hosted model through a custom base URL and key, so the cleanest control point is a gateway in front of vLLM. The agent gets a revocable gateway key with a spend cap, never a cloud credential or a kubeconfig.
Idle GPU hours, not the GPU hourly price, blow up self-hosting budgets. A GPU node left running 730 hours a month to serve an 8x5 workload (about 174 useful hours) wastes roughly 76% of its spend. Scale-to-zero plus auto-stop on non-production beats shopping for a cheaper GPU hour.
Self-hosting only wins above a sustained-usage floor. Compute cost per million tokens as GPU hourly price divided by (measured output tokens/sec x 3600 / 1,000,000), divide by your real utilisation rate, compare to the published hosted per-million price, then add 20-40% for platform engineering. Below a few hundred million tokens a month, a hosted API almost always wins on pure cost.
Agentic coding is where open-weight models are weakest today, so the pragmatic 2026 default is hybrid: self-host the high-volume boring work (autocomplete, embeddings, summarisation, code review) and route hard agentic tasks to a frontier hosted model, with the routing decision made at the gateway so you can change your mind without touching the agents.
Most "how to self-host an LLM" pages pretend one tool does the whole job. It does not. If your platform team wants Claude Code and Codex calling open-weight models you run yourself, while keeping cost and access under control, you are assembling four layers and each one fails in a different way.
I run Kubernetes inside customer cloud accounts for a living, so I will be honest about where the open-source tools win outright, where the hyperscalers are the easy button, and the one layer almost everyone underestimates. Qovery sits at that last layer, and I will tell you plainly what it is not before I tell you what it is.
What tools do you need to deploy open-weight models internally?
You need four layers, not one tool: vLLM (or SGLang or TGI) to serve the weights, KubeAI or Ray Serve to manage model lifecycle and autoscaling on Kubernetes, LiteLLM as the OpenAI-compatible gateway that issues keys and enforces budgets, and a deployment layer (Qovery, or Terraform plus CI you maintain) so people and agents can ship the service without holding cloud admin credentials. Skip any one of them and a specific thing breaks:
No inference server = unusable throughput and cost per token many times too high. vLLM reports up to 24x higher throughput than naive HuggingFace Transformers serving thanks to PagedAttention and continuous batching (vLLM launch post, June 2023; Kwon et al., SOSP 2023).
No orchestrator = manual GPU babysitting and no scale-to-zero. Something has to load weights, autoscale replicas, and send a model back to zero when idle.
No gateway = no per-team cost attribution and no revocable keys. You cannot see who spent what, and you cannot cut off a runaway agent.
No deployment layer = cloud credentials sprayed into CI and agent configs. This is the layer that takes longest to build and the one teams keep forgetting.
Two problems get conflated here, and they have different guardrails. One is an agent calling a self-hosted model. The other is an agent deploying a model service. Most teams need both, and the blast radius of the second is much larger.
The models worth scoping for in 2026 are Mistral (Mistral Small, Devstral, Codestral), Llama, Qwen, DeepSeek and gpt-oss. They all sit behind the same OpenAI-compatible HTTP API once served, which is exactly why the layers above them are interchangeable.
A word on Ollama: it is the right tool on a laptop and the wrong one as shared infrastructure. It has no multi-tenant continuous batching at production concurrency, no budget model, and no RBAC. Great for a developer trying a model locally, not for a team endpoint.
And "just use Bedrock, Vertex AI, SageMaker or GKE directly" does not save you. The managed model endpoint is the easy part. You still write the access, quota and cost-attribution layer yourself, and that is the layer that takes the longest.
Day-one minimum viable stack: one model, vLLM, one gateway, one namespace. Everything else is earned by a metric.
How do you point Claude Code and Codex at a self-hosted model instead of a hosted API?
Both agents accept a custom endpoint, so you serve the open-weight model with vLLM, put LiteLLM in front of it, and give each agent its own gateway key with a budget and a rate limit. Claude Code reads ANTHROPIC_BASE_URL plus a credential (ANTHROPIC_AUTH_TOKEN for a bearer token, ANTHROPIC_API_KEY for an x-api-key gateway), set as environment variables or in a settings file (Claude Code gateway docs, checked October 2026). Codex CLI declares a custom provider in config.toml:
The agent never sees IAM, a kubeconfig, or a provider key.
Two honest caveats, both current as of October 2026. Anthropic supports gateways that speak the Anthropic Messages format and states it does not support routing Claude Code to non-Claude models through any gateway, so pointing Claude Code at an open-weight backend means running a gateway that translates to the Messages format, and expecting rough edges around newer features. Codex is the more natural fit for custom providers, but its wire_api now accepts only responses (the Responses API) as its configuration reference shows, so a vLLM backend that exposes Chat Completions sits behind LiteLLM or a proxy that translates to the Responses protocol. In both cases the gateway is doing real work, which is another reason it is not optional.
What vLLM's OpenAI-compatible server gives you: chat completions, completions, tool and function calling, structured outputs and streaming (vLLM docs). That covers most of what a coding agent asks for.
Now the part the vendor pages skip. Agentic coding leans on tool-calling reliability and long-context behaviour, and that is exactly where open weights still trail. On the Aider polyglot leaderboard (checked October 2026), the top frontier model (GPT-5 high) scores 88.0% while the strongest open-weight entry (DeepSeek-V3.2-Exp) scores 70.2%, roughly an 18-point gap on real multi-language coding tasks. Measure it on your own tasks before you bet an agent workflow on an open-weight model.
The routing pattern that works: the gateway sends bulk and cheap tasks to the self-hosted model, falls back to a hosted model on hard agentic tasks, logs both, and attributes spend per agent. Per-agent gateway keys beat per-developer cloud IAM because you get instant revocation, hard spend caps, and an audit trail of which agent called which model with how many tokens.
Common breakages to expect: context-window mismatch between the model config and what the agent assumes, tool-call schema differences, missing prompt caching (so your token bill runs higher than the hosted equivalent), and streaming timeouts when a load balancer sits between the agent and vLLM.
Client
Custom base URL?
Config surface
Works through a LiteLLM gateway?
Known limitation with open-weight models
Docs (checked Oct 2026)
Claude Code
Yes
ANTHROPIC_BASE_URL + ANTHROPIC_AUTH_TOKEN/ANTHROPIC_API_KEY (env or settings file)
Yes, if the gateway speaks the Anthropic Messages format
Anthropic does not support non-Claude models via a gateway; newer features may not pass through
Which tool does what? vLLM vs KubeAI vs LiteLLM vs SkyPilot vs Ray Serve vs Ollama vs CoreWeave vs Qovery
These tools are not substitutes, they sit at different layers, and the expensive mistake is buying one and expecting the other three. vLLM serves the model, KubeAI and Ray Serve orchestrate it on Kubernetes, LiteLLM governs the calls, SkyPilot sources GPUs across clouds, CoreWeave sells the GPUs, Ollama is local development, and Qovery is the deployment and governance layer running in your own cloud account.
Where each genuinely wins, by name:
vLLM for raw throughput per GPU, behind an OpenAI-compatible server. It is the default serving engine, with 93.3k GitHub stars and over 2,000 contributors as of October 2026 (vllm-project/vllm).
KubeAI for Kubernetes-native serving with scale-from-zero behind an OpenAI-compatible API. It scales models from minReplicas: 0, runs vLLM and Ollama backends, and needs neither Istio nor Knative (kubeai.org).
Ray Serve for distributed or custom multi-model serving when a single engine is not enough (Ray Serve LLM docs).
LiteLLM for virtual keys, budgets and provider fan-out. It exposes an OpenAI-compatible proxy, tracks spend per key, user and team, and enforces max_budget, rpm and tpm limits across 100+ providers (LiteLLM virtual keys).
SkyPilot for multi-cloud GPU sourcing and spot recovery. It documents spot instances at 70-90% cheaper than on-demand and auto-recovers managed jobs after a preemption (SkyPilot managed jobs).
CoreWeave or another neocloud when your hyperscaler has no GPU quota left.
What Qovery is not: not an inference engine, not a GPU scheduler, not a model router. If a page tells you Qovery serves models, close it. What Qovery is: the self-service deploy and governance layer that runs on your own AWS, GCP, Azure or Scaleway account, or your existing self-managed Kubernetes cluster (BYOC, so the cloud bill, committed-use discounts and Savings Plans stay in your name). Verified capabilities: git-push deployments, preview environments per pull request that tear down automatically, environment auto-stop for non-production, managed cluster upgrades and patches, per-environment RBAC, and databases backed by managed cloud services. For agents specifically, Qovery ships an MCP server (read-only by default, organisation-scoped, backed by a Viewer API token) and Qovery Skills that follow the Agent Skills open standard and work with Claude Code, Codex, Cursor and Gemini CLI.
Pick two to start: vLLM plus LiteLLM covers roughly 80% of internal use before you touch orchestration or multi-cloud. Add KubeAI when you need scale-to-zero, SkyPilot when you run out of quota in one cloud, and the governance layer when humans or agents start deploying the service themselves.
Ship faster on infrastructure you control.
Qovery gives your team self-service deployments on your own AWS, GCP, Azure or Scaleway account - or your existing Kubernetes cluster. Start deploying in under 10 minutes.
How do you stop internal GPU costs from exploding once agents can deploy models?
Install four controls before you give anyone a GPU node pool: scale-to-zero on the model server, auto-stop on non-production environments, a TTL on ephemeral environments, and hard per-key spend caps at the gateway. Idle hours dominate the bill, so hunting a cheaper GPU hour is the last lever, not the first.
Here is the idle arithmetic with a real, dated price. An AWS g5.xlarge (one NVIDIA A10G, 24 GB) lists at $1.006/hour on-demand in us-east-1 as of October 2026 (AWS EC2 pricing via instances.vantage.sh). Left running the full 730 hours in a month, that is about $734. If the real workload is 8 hours a day, 5 days a week (about 174 useful hours), you used $175 of it and set fire to the other $559, roughly 76%. No cheaper GPU hour fixes a 76% waste rate. Scale-to-zero does.
The mechanics:
Scale-to-zero: KubeAI scales models from zero, and KEDA drives event- and HTTP-driven autoscaling down to zero replicas. The cost is a cold start while weights load. Mitigate it by caching weights on local NVMe, keeping one warm replica for interactive traffic, and serving first token from a smaller quantised variant.
Spot and preemptible GPUs: SkyPilot documents 70-90% spot savings and auto-recovers managed jobs after a preemption (SkyPilot docs). Use spot for batch and fine-tuning. Do not use it for an interactive coding agent that cannot tolerate an eviction mid-task.
Gateway-level budgets: LiteLLM enforces per-key, per-team and per-model budgets, rate limits and spend tracking across 100+ providers (LiteLLM docs). This is your hard ceiling.
Throughput as the real cost lever: continuous batching, PagedAttention, prefix caching and quantisation move cost per token further than switching GPU vendor. The same PagedAttention work that gives vLLM up to 24x over HuggingFace Transformers (SOSP 2023) is cheaper than negotiating a new GPU contract.
Cloud waste is not a rounding error. Flexera's 2026 State of the Cloud Report puts estimated wasted cloud spend back up to 29%, and GPU is the most expensive line to waste.
At this layer Qovery contributes environment auto-stop for non-production, ephemeral environments per pull request that die with the PR, and BYOC, so GPU spend lands on your own cloud account where your committed-use discounts and Savings Plans apply.
Lever
Tool that provides it
Typical impact
Risk / tradeoff
Effort
Scale-to-zero on the model server
KubeAI, KEDA
Removes idle GPU hours entirely
Cold-start latency
Medium
Auto-stop non-production
Qovery
Kills nights/weekends idle
Restart on first use
Low
Ephemeral env TTL
Qovery, Terraform + CI
Stops zombie environments
Must re-create to resume
Low
Spot / preemptible GPUs
SkyPilot
70-90% off compute
Evictions mid-task
Medium
Per-key budget at gateway
LiteLLM, Portkey
Hard spend ceiling per agent
None material
Low
Continuous batching + quantisation
vLLM, SGLang
Lower cost per token
Quality/latency tuning
Medium
BYOC committed-use discounts
Qovery (BYOC)
Your existing discounts apply
Requires own cloud account
Low
How do you let an AI agent deploy a model service without giving it cloud credentials?
Give the agent a narrow, audited API that requests a deployment of a pre-approved service template, never an IAM role, a kubeconfig or a provider key. The platform team fixes the templates, the environment scope and the blast radius up front, namespace quotas cap the damage, and production stays behind human review.
The threat model is not hypothetical. An agent with cloud admin can spin up GPU fleets, exfiltrate model weights, or delete a cluster, and prompt injection makes that a live path, not a theoretical one. The OWASP Top 10 for LLM Applications 2025 lists LLM01:2025 Prompt Injection and LLM06:2025 Excessive Agency as top risks for exactly this reason: an attacker who can influence the prompt should not be able to reach your cloud account through the agent.
The guardrail pattern, in order:
Pre-approved service templates so the agent picks from a menu, not a blank cloud console.
Per-environment RBAC so a key scoped to staging cannot touch production.
Ephemeral environments with a TTL so anything the agent creates expires by default.
Namespace ResourceQuota and LimitRange so a runaway agent hits a wall before it hits your bill.
Mandatory human review on anything touching production.
This is where Qovery fits, and it maps onto the pattern directly. An approved model service deploys into a scoped environment on your own AWS, GCP, Azure or Scaleway account, or your existing Kubernetes cluster. Per-environment RBAC decides what can touch production, and auto-stop plus ephemeral environments cap both blast radius and bill. For the agent surface specifically, the Qovery MCP server is read-only by default, organisation-scoped, and backed by a Viewer API token, so write access requires an explicit read_write=true parameter plus a separate console setting. The agent gets the same policy-as-code guardrails and audit logging your human engineers use, not raw personal credentials. (If you want the longer version, we wrote a dedicated piece on what an MCP server for infrastructure actually is.)
Honest alternatives, and when each is enough. Terraform in CI with GitHub OIDC and plan-only agent permissions works if you are disciplined about who approves applies. Plain Kubernetes namespaces with quotas and a locked-down service account works for a single team. SkyPilot under a restricted IAM role covers GPU jobs. All of them are things you build and maintain yourself, which is fine until you count the hours.
Whatever you choose, log who or what deployed, which model and weights version, which GPU class, which data it touched, and which gateway key paid for it. EU teams should write this down now rather than later: under the EU AI Act, general-purpose AI model obligations have applied since 2 August 2025, with broader obligations phasing in through 2 August 2026 and beyond. Data residency is a concrete reason to self-host in-region, and an audit trail is a concrete reason to run the deploy through a governed layer instead of a shell script.
When is self-hosting Mistral or Llama actually cheaper than a hosted API?
Self-hosting wins above a sustained-usage floor and loses badly below it. Compute cost per million tokens as GPU hourly price divided by (measured output tokens/sec x 3600 / 1,000,000), divide by your real utilisation rate, compare against the published hosted per-million price, then add 20-40% for platform engineering. Below a few hundred million tokens a month, a hosted API almost always wins on pure cost.
The table below is an illustrative model built from public prices, not a benchmark. Inputs: an always-on g5.xlarge at $1.006/hour (~$734/month) plus about 35% platform overhead, so call it ~$1,000/month fixed for one small quantised model with one warm replica. The hosted comparator is Mistral Large at $0.5/M input and $1.5/M output (checked October 2026), blended to roughly $1/M for a mixed coding workload. The point is not the exact crossover, it is that the self-hosted cost is fixed whether you push 10M or 1B tokens, while the hosted cost scales with usage.
Tokens/month
Self-hosted all-in (illustrative)
Hosted at ~$1/M (illustrative)
Cheaper option
Caveat
10M
~$1,000
~$10
Hosted
Self-host GPU sits idle
100M
~$1,000
~$100
Hosted
Still mostly idle
500M
~$1,000
~$500
Hosted
Approaching the floor
1B
~$1,000
~$1,000
Break-even
Depends on your real utilisation
3B
~$1,000
~$3,000
Self-hosted
Only if utilisation is genuinely high
Replace every number with your own measured throughput and utilisation before you decide anything. Nobody runs a GPU at 100%, and the break-even moves a lot between 30%, 60% and 90% utilisation.
Teams self-host for non-cost reasons too, and those are often the real reason: EU data residency and sovereignty, no-training-on-our-data guarantees, tail latency, air-gapped networks, and model pinning (a hosted endpoint can change under you, your pinned weights cannot). Just do not pretend those are cost savings.
And do not forget the hidden costs: platform-engineer time, GPU driver and node-pool maintenance, Kubernetes cluster upgrades, observability, an on-call rota, and the cost of an idle weekend.
My recommendation is hybrid by default. Self-host autocomplete, embeddings, summarisation and code review, route hard agentic reasoning to a frontier hosted model, and make the routing decision at the gateway so you can change your mind without touching a single agent config.
What does a working reference architecture look like, end to end?
The stack a platform team can copy in 2026: Kubernetes in your own cloud account, a GPU node pool that scales to zero, vLLM serving the open-weight model, KubeAI or Ray Serve managing model lifecycle, LiteLLM as the single OpenAI-compatible endpoint issuing per-agent keys and budgets, and Qovery as the self-service deploy and governance layer on top.
Build it in this order, and cut hard on day one:
One model on vLLM, one namespace, no autoscaling, no multi-cloud, no fine-tuning.
LiteLLM in front, issuing one key per agent with a budget and a rate limit.
Scale-to-zero via KubeAI once you have seen the idle hours on a dashboard.
Ephemeral environments and auto-stop so non-production stops billing overnight.
A four-week rollout that works:
Week 1: one model behind a gateway for one team.
Week 2: budgets, spend attribution and audit logs.
Week 3: scale-to-zero plus ephemeral environments.
Week 4: hand scoped keys and a deploy API (or the Qovery Skills and read-only MCP server) to the agents.
Measure from day one: output tokens/sec, time to first token, GPU utilisation percentage, cost per team per month, prefix-cache hit rate, and the share of requests falling back to the hosted model. That last number tells you whether self-hosting is paying off or whether you are quietly running two bills.
You probably do not need, yet: multi-cluster, custom schedulers, fine-tuning pipelines, or your own router. Add each only when a metric forces it.
The reason I like this shape is boring in the best way. With Qovery, the model service, the gateway and their databases deploy and tear down with the same git-push workflow as the rest of your apps, on your own AWS, GCP, Azure or Scaleway account or your existing Kubernetes cluster. No separate snowflake pipeline for "the AI stuff," and the bill stays in your name.
Frequently asked questions
Can Claude Code and Codex use self-hosted open-source models?
Yes, both point at a custom endpoint. Claude Code reads ANTHROPIC_BASE_URL plus a credential (ANTHROPIC_AUTH_TOKEN or ANTHROPIC_API_KEY), and Codex CLI declares a custom provider in config.toml with base_url, env_key and wire_api. The clean path is to serve the model with vLLM and put a LiteLLM gateway in front, so each agent gets a revocable key with a spend cap. Note that Anthropic does not officially support routing Claude Code to non-Claude models through a gateway (checked October 2026), so Codex is the more natural fit for an open-weight backend.
What is the best tool for serving open-weight models like Mistral, Llama or Qwen internally?
vLLM is the default serving engine for raw throughput per GPU, exposing an OpenAI-compatible API, with 93.3k GitHub stars and 2,000+ contributors as of October 2026. SGLang and TGI are credible alternatives. For scale-to-zero and lifecycle management on Kubernetes, put KubeAI or Ray Serve above vLLM. vLLM serves, the orchestrator manages, and they are not substitutes for each other.
Do I need KubeAI, LiteLLM and SkyPilot, or can one tool do everything?
No single tool does everything, because they sit at different layers. KubeAI orchestrates models on Kubernetes with scale-from-zero, LiteLLM is the gateway that issues keys and enforces budgets, and SkyPilot sources GPUs across clouds and handles spot recovery. Start with vLLM plus LiteLLM, which covers about 80% of internal use, and add the others when a specific problem (idle cost, GPU quota) forces it.
How do I cap how much a coding agent can spend on internal GPU infrastructure?
Put a gateway like LiteLLM in front of the model and give each agent a virtual key with a hard max_budget and rpm/tpm limits, so spend is capped and attributable per agent. Then control the infrastructure cost separately with scale-to-zero on the model server (KubeAI or KEDA) and environment auto-stop on non-production (Qovery), because idle GPU hours, not per-token price, are what actually blow up the bill. A GPU left on all month to serve an 8x5 workload wastes about 76% of its cost.
Is it safe to let an AI agent deploy to my cloud account?
Only if the agent gets a narrow, audited API instead of an IAM role or kubeconfig. Prompt injection and excessive agency are the top two OWASP LLM risks for 2025, so scope the agent to pre-approved templates, per-environment RBAC, ephemeral environments with a TTL, and mandatory human review on production. Qovery's MCP server is read-only by default and organisation-scoped, with write access gated behind an explicit parameter and a console setting, which is the kind of default you want.
Is self-hosting Mistral or Llama cheaper than using a hosted API?
Only above a sustained-usage floor, usually a few hundred million tokens a month. The self-hosted GPU bill is fixed whether you use it or not (an always-on single-GPU node is roughly $1,000/month all-in on AWS as of October 2026), while a hosted API like Mistral Large ($0.5/M in, $1.5/M out) scales with usage, so below the floor hosted wins easily on pure cost. The pragmatic default is hybrid: self-host high-volume routine work, route hard agentic tasks to a frontier hosted model, and decide at the gateway.
Romaric founded Qovery to make Kubernetes accessible to every engineering team. He writes about platform strategy, developer experience, and the future of cloud infrastructure.
Next step
Ship faster on infrastructure you control.
Qovery gives your team self-service deployments on your own AWS, GCP, Azure or Scaleway account - or your existing Kubernetes cluster. Start deploying in under 10 minutes.