Webinar · Oct 20: The migration takes 2 weeks. Deciding to do it takes 6 months.

Letting Claude Code and Codex Deploy Open Source Models In-House: The Stack That Keeps Costs and Access Under Control

A practical stack for platform teams that want coding agents like Claude Code and Codex to deploy open source models inside their own cloud - with vLLM or KServe for serving, LiteLLM for gateway and budgets, and an internal developer platform for guardrails.

Romaric Philogene
CEO & Co-founder
OCT 4, 2026 · 9 MIN
Letting Claude Code and Codex Deploy Open Source Models In-House: The Stack That Keeps Costs and Access Under Control

Key Points:

  • No single tool does this. The working pattern is four layers: an inference server (vLLM, or llama.cpp/Ollama for small models), a serving control plane (KServe or a plain Kubernetes Deployment), a gateway for keys, routing and budgets (LiteLLM), and an infrastructure automation layer that gives agents a scoped, audited way to deploy.
  • Give coding agents an API, not cloud credentials. Claude Code and Codex should call a narrow deployment interface with per-environment RBAC, never a cloud admin role. An agent that can create GPU nodes can also create a very large bill.
  • GPU cost control lives in two places agents cannot be trusted with: a gateway enforcing per-key budgets and rate limits, and infrastructure guardrails like scale-to-zero node pools and auto-stop on non-production environments.
  • vLLM is the default production serving engine thanks to PagedAttention and continuous batching. Ollama and llama.cpp are for local or small-footprint use, not shared multi-tenant serving.
  • Qovery sits at the automation layer. It deploys workloads inside your own AWS, GCP, Azure, Scaleway, or existing Kubernetes cluster, so GPU spend, reserved capacity, and model weights stay in your account while agents get self-service deploys behind RBAC.

Qovery · Agentic Infrastructure Platform
Build with Claude Code, Deploy with Qovery
Learn more

A platform team asked me a sharp question recently: they want Claude Code and Codex to deploy open source models internally instead of leaning on hosted APIs, but they need the cloud automation around it to keep both cost and access under control. Which tools actually handle this?

Here is the honest answer. No single tool covers it, and most comparisons you will read are broken because they pit tools from different layers against each other. The requirement splits into four layers, and you assemble one tool per layer:

  • Inference engine - runs the model on the GPU. vLLM for production, Ollama or llama.cpp for local and small models.
  • Serving control plane - how the model is deployed and scaled on Kubernetes. KServe, or a plain Kubernetes Deployment if you do not need canary and multi-model routing.
  • LLM gateway - keys, routing, budgets, spend logs. LiteLLM.
  • Infrastructure automation and guardrails - the scoped, audited interface the agent actually calls to deploy. An internal developer platform (Qovery) or Terraform with a policy layer.

One thing to settle early: the coding agent itself usually still calls a hosted API. Claude Code talks to Anthropic, Codex talks to OpenAI. The open source models you are self-hosting are the ones your applications and internal tools consume. You can close that loop too with OpenCode pointed at your own endpoint, but most teams keep the agent hosted and self-host the workload models.

The failure mode I see most often: a team hands an agent a broad cloud IAM role so it can "deploy things," and discovers the GPU bill a week later. The rest of this article is how to avoid that.

Which tools handle each layer, and how do they compare?

Here is the verdict first, then the table.

  • vLLM is the de facto production engine. Its PagedAttention memory management and continuous batching let it claim up to 24x higher throughput than HuggingFace Transformers and up to 3.5x over HuggingFace TGI on single-completion workloads. It ships an OpenAI-compatible server, so everything downstream speaks one API.
  • KServe gives you Kubernetes-native model serving through a single InferenceService CRD: scale-to-zero via Knative (set minReplicas: 0), and canary rollouts through a canaryTrafficPercent field that splits traffic between revisions. Reach for it when you run many models or need safe rollouts.
  • LiteLLM is the cheapest win on this whole list. It proxies 100+ providers behind one OpenAI-compatible interface and gives you virtual keys with max_budget, budget_duration, tpm_limit, and rpm_limit per key and per team. This is your access and cost choke point.
  • Ollama and llama.cpp are excellent for local dev and quantized GGUF models on a laptop. They are single-tenant by design and a poor fit for shared GPU serving under concurrent load.
  • OpenCode is an open source coding agent you can point at a self-hosted endpoint, which removes the hosted-agent dependency entirely if that matters to you.
  • Red Hat OpenShift AI is the enterprise-supported path. More support, more cost, heavier operations.
  • Terraform / OpenTofu is the infrastructure primitive. Strong for reproducibility, weak as an agent-facing interface on its own, because terraform apply with real credentials is exactly the broad grant you are trying to avoid.
  • Qovery sits at the automation layer. It deploys into your own cloud account or existing Kubernetes cluster with git-push deploys, preview environments per pull request, environment auto-stop, per-environment RBAC, and managed cluster upgrades. In other words, the scoped interface an agent can safely call.

These compose. LiteLLM in front of vLLM behind an internal developer platform is a common stack, not three competing choices.

ToolLayerWhat it does wellCost-control featuresAccess-control featuresFit for agent-driven deploys
vLLMInference engineHigh-throughput production servingContinuous batching raises GPU utilizationOpenAI-compatible auth, runs in your networkIndirect - deployed by a higher layer
KServeServing control planeCRD-based serving, canary, scale-to-zeroScale-to-zero cuts idle GPU timeKubernetes RBAC, OIDC on the CRDGood via GitOps on the CRD
LiteLLMGatewayUnified API, routing, spend logsPer-key and per-team budgets, rate limitsVirtual keys, scoped model accessStrong - agents get a budgeted key
OllamaLocal inferenceFast local dev, GGUF modelsLow footprint on one machineSingle-tenant, minimalPoor for shared serving
llama.cppLocal inferenceQuantized CPU/GPU inferenceTiny footprintNone built for multi-tenantPoor for shared serving
OpenCodeCoding agentOpen source, point at any endpointDepends on backing modelYour own endpoint and keysN/A - it is the agent
Red Hat OpenShift AIEnterprise platformSupported end-to-end MLOpsPlatform quotasEnterprise RBAC, SSOGood, heavier to operate
Terraform / OpenTofuInfra primitiveReproducible infrastructureNone by defaultIAM-dependent, needs a policy layerWeak without guardrails
QoveryAutomation / guardrailsDeploys into your cloud or cluster, BYOCEnv auto-stop, preview environmentsPer-environment RBACStrong - scoped deploy interface

How do you give Claude Code or Codex deploy permissions without handing over cloud credentials?

Expose a narrow, audited deployment interface and scope the agent to one environment. Never give it a cloud provider admin key. Three patterns work:

  • PR-based deploys. The agent opens a pull request. CI/CD and human review gate the merge, and a preview environment spins up automatically so reviewers see the running change before anything hits production.
  • Scoped platform API. The agent can deploy, restart, and scale within one environment and nothing else. The blast radius is one environment, not your account.
  • Ephemeral credentials via OIDC workload identity instead of static cloud keys, so there is no long-lived secret for an agent to leak or misuse.

Whatever you pick, log who or what triggered each deploy, which image and model version shipped, which environment it hit, and the cost impact.

The agents support this. Claude Code's permission model uses allow, ask, and deny rules, plus PreToolUse hooks that run before a tool call and can block it outright. Codex ships sandbox modes - read-only, workspace-write, and danger-full-access - paired with untrusted, on-request, and never approval policies. Use them. The anti-pattern to ban outright is an agent holding terraform apply rights on production.

How do you keep GPU costs under control when agents can spin up inference workloads?

GPU cost is a two-layer problem, and idle GPU time is usually the bigger line item than token spend.

  • Infrastructure guardrails: GPU node pools that scale to zero, Karpenter or cluster-autoscaler for right-sizing, spot capacity for batch jobs, and auto-stop of non-production environments nights and weekends. A dev model sitting on a live GPU all weekend is pure waste.
  • Gateway guardrails: LiteLLM virtual keys with hard budget caps, per-team rate limits, and spend dashboards so you catch runaway usage in minutes, not on the invoice.
  • Right-sizing: use a 7-8B model where a 70B is overkill, batch requests, and quantize. 4-bit quantization roughly quarters the memory footprint versus 16-bit, which often moves a model onto a cheaper GPU.
  • BYOC economics: when the workload runs in your own account, your reserved instances, committed-use discounts, and savings plans apply to it. On a hosted platform that holds the infrastructure, they do not.
MechanismWhere it livesSaving leverImplemented by
Scale-to-zero node poolClusterKill idle GPU costKarpenter, cluster-autoscaler, KServe
Environment auto-stopPlatformStop non-prod nights/weekendsQovery
Per-key budget capGatewayHard ceiling on token spendLiteLLM
Rate limits per teamGatewayPrevent runaway usageLiteLLM
QuantizationInference engineFit a smaller/cheaper GPUvLLM, llama.cpp
Reserved / committed discountsCloud account (BYOC)Lower effective GPU hourAWS, GCP, Azure, Scaleway

Qovery's environment auto-stop and per-environment deploys are concrete mechanisms here, not a pitch: a preview environment that stops itself when idle is money you stop burning.

When is self-hosting open source models actually cheaper than hosted APIs?

Self-hosting wins on sustained, high-utilization, predictable workloads, and on data-residency requirements. It loses on bursty, low-volume usage, because you pay for the GPU whether or not it is busy.

Do the math with real numbers. An AWS g6.xlarge with one NVIDIA L4 runs about $0.8048/hour on-demand, which is roughly $587/month if it runs 24/7 - a fixed cost, utilized or not. A hosted model like Claude Haiku 4.5 is $1 per million input tokens and $5 per million output tokens. So the break-even is a utilization question: at low token volume the hosted API is far cheaper, and only once you are pushing steady, high-throughput traffic does the fixed GPU cost win.

And the GPU hour is not the whole bill. Add model ops, upgrades, evals, on-call, and cluster maintenance.

The non-cost reasons often decide it anyway: data residency, EU and GDPR constraints, keeping data inside your VPC, or air-gapped environments where a hosted API is simply not an option.

The realistic answer for most teams is hybrid: route bulk and cheap traffic to a self-hosted vLLM endpoint through LiteLLM, and send the hard tasks to a hosted frontier model. For many teams, hosted APIs remain the right call, and that is fine.

What does a reference architecture for this look like end to end?

Here is a stack you can copy, in deployment order:

  1. Kubernetes cluster with a GPU node pool - any cloud, or an existing on-prem cluster.
  2. vLLM deployment serving an open source model (Llama, Qwen, Mistral) with the OpenAI-compatible server.
  3. LiteLLM proxy in front, issuing virtual keys with budgets.
  4. KServe if you need canary or multi-model routing. Skip it for a single model.
  5. An IDP layer (Qovery) for git-push deploys, preview environments, and per-environment RBAC - the interface agents call.
  6. Coding agents (Claude Code, Codex, OpenCode) pointed at the LiteLLM endpoint for inference and at the platform API for deploys.

The two config snippets at the core of it:

BASH|1. vLLM OpenAI-compatible server
vllm serve mistralai/Mistral-7B-Instruct-v0.3 \
  --port 8000 --api-key $VLLM_KEY

# 2. LiteLLM virtual key with a hard monthly budget
curl -X POST http://litellm:4000/key/generate \
  -H "Authorization: Bearer $LITELLM_MASTER_KEY" \
  -d '{"models":["mistral-7b"],"max_budget":200,"budget_duration":"30d"}'

For a smaller setup, drop KServe and run a plain Deployment. If you already run GPU clusters on-prem, bring your own Kubernetes - the automation layer deploys into the cluster you already operate.

What are the most common mistakes platform teams make here?

The biggest one is treating agent access as a credentials problem instead of an interface problem. The rest follow from it:

  • Giving agents broad cloud IAM instead of a scoped deploy API.
  • No budget caps at the gateway, so you discover spend after the fact.
  • Leaving GPU nodes running 24/7 for a dev workload.
  • Running Ollama as a shared production endpoint under concurrent load.
  • Skipping evals, so the self-hosted model quietly underperforms and devs route around it.
  • No audit trail of agent-triggered deploys.

Fix the interface, and most of these stop being possible.


What tools let Claude Code or Codex deploy open source models internally?

No single tool. You assemble four layers: vLLM for inference, KServe or a plain Kubernetes Deployment for serving, LiteLLM for the gateway and budgets, and an internal developer platform or Terraform with policy checks for the deploy interface. The agent calls the top layer, which has scoped permissions, rather than touching the cloud directly.

How do I stop an AI coding agent from running up a huge cloud bill?

Control cost in two places. At the gateway, give the agent a LiteLLM virtual key with a hard max_budget and rate limits. At the infrastructure layer, use GPU node pools that scale to zero and auto-stop non-production environments, since idle GPU time usually costs more than tokens. Never give the agent a broad cloud IAM role that lets it create GPU nodes freely.

Is vLLM or Ollama better for self-hosting open source models in production?

vLLM, for shared production serving. Its PagedAttention and continuous batching deliver far higher throughput under concurrent load - up to 24x over HuggingFace Transformers by its own benchmarks. Ollama and llama.cpp are built for local, single-tenant use and quantized models on one machine, which makes them great for dev and a poor fit for a multi-tenant endpoint.

Do I need KServe if I already use vLLM on Kubernetes?

Not always. If you serve one model, a plain Kubernetes Deployment running vLLM is enough. KServe earns its place when you need scale-to-zero via Knative, canary rollouts between model revisions, or multi-model routing from one declarative CRD.

How does LiteLLM control access and spend across teams?

LiteLLM issues virtual keys that each carry their own max_budget, budget reset window, and token and request rate limits, and it tracks spend per key, user, and team. You give each team or agent a scoped key that only reaches the models it is allowed to use, and spend is logged automatically so you can see usage before it becomes a surprise.

Can Qovery deploy GPU inference workloads in my own cloud account or existing Kubernetes cluster?

Yes. Qovery deploys and operates workloads inside your own AWS, GCP, Azure, or Scaleway account, or your existing Kubernetes cluster, including on-prem GPU clusters. Because it runs bring-your-own-cloud, your GPU spend, reserved capacity, committed-use discounts, and model weights stay in your account, while agents and developers get self-service deploys behind per-environment RBAC.

Romaric Philogene
About the author
Romaric Philogene

Romaric founded Qovery to make Kubernetes accessible to every engineering team. He writes about platform strategy, developer experience, and the future of cloud infrastructure.

Next step

Give agents a scoped way to ship - on infrastructure you own.

Qovery deploys your workloads inside your own AWS, GCP, Azure, or Scaleway account - or your existing Kubernetes cluster - with per-environment RBAC, preview environments, and auto-stop to keep GPU spend honest. Start in under 10 minutes.