Webinar · Oct 20: The migration takes 2 weeks. Deciding to do it takes 6 months.

How to Let Claude Code and Codex Deploy Self-Hosted Models Without Losing Control of Cost and Access

A four-layer stack for platform teams that want AI coding agents like Claude Code and Codex to deploy open-weight models on their own GPU infrastructure, with per-key budgets, RBAC, cost attribution, and multi-cloud portability instead of one-provider lock-in.

Romaric Philogene
CEO & Co-founder
OCT 8, 2026 · 8 MIN
How to Let Claude Code and Codex Deploy Self-Hosted Models Without Losing Control of Cost and Access

I get this question from platform teams almost every week now: we want Claude Code and Codex to deploy open-weight models on our own GPUs instead of leaning on hosted APIs, but we do not want an agent loop to quietly spin up an 8xH100 node and leave it running all weekend. What actually handles this?

Here is the honest answer up front: no single tool does. The working pattern is four layers, and you pick one tool per layer. A serving layer that runs the model (vLLM, Ray Serve, or Ollama for small models). A gateway layer that gives each agent an OpenAI-compatible endpoint with per-key spend limits (LiteLLM). A provisioning layer that hands out GPU capacity with guardrails (OpenTofu/Terraform, or an internal developer platform like Qovery). And a cost attribution layer so you can answer "what did this model cost last week" (OpenCost or Kubecost). Get those four right and the agents become productive without becoming a budget risk.

Qovery · Agentic Infrastructure Platform
Build with Claude Code, Deploy with Qovery
Learn more

Let me walk through each decision, with real numbers.

What does a platform team need to lock down before letting an AI agent deploy models?

Decide five things before you touch a single tool: how the agent gets a deployment interface instead of credentials, where spend ceilings live, how RBAC is scoped per environment, how every action is logged, and whether the path runs on more than one GPU provider.

Agents do not create new failure modes so much as amplify the ones you already have. A human forgets to delete a big GPU node group maybe once a month. An agent stuck in a retry loop can request one every ten minutes, politely, with a reasonable-sounding commit message each time. So the five requirements:

  • The agent gets an API call or a pull request, never an IAM user or a service-account key. A credential in an agent's context is a credential one prompt injection away from being used.
  • Two enforcement points, not one. Token and request budgets at the gateway, plus node-pool maximums and auto-stop at the infrastructure layer. One without the other leaves a hole.
  • RBAC scoped per environment, so a dev-facing agent physically cannot redeploy or stop production inference.
  • An audit trail that ties a cost spike back to a specific agent run, prompt, and commit.
  • Portability, because GPU pricing and availability move faster than any procurement cycle.

Quick sanity check: if the answer to "who pays for this GPU" is "nobody labelled it," stop and fix attribution before you give any agent a deploy button.

Which tools handle each layer, and what does each one actually do?

Self-hosting open-weight models for coding agents breaks into four layers, and no single tool covers all four: serving (vLLM, Ray Serve, Ollama, SGLang), gateway and spend control (LiteLLM), provisioning (OpenTofu/Terraform or a developer platform like Qovery), and cost attribution (OpenCost, Kubecost).

The serving layer is where the model runs. vLLM is the default for high-throughput serving thanks to continuous batching and PagedAttention; its v0.6.0 release reported 2.7x higher throughput and 5x faster time-per-output-token for Llama 3 8B on a single H100 versus the prior version, and the underlying PagedAttention technique is why it became the standard. SGLang is a fast-moving alternative, Ray Serve shines for multi-model composition and autoscaling, Ollama is great for small local models, and Baseten/Modal give you managed serverless GPU if you would rather not run the server yourself.

The gateway layer is the endpoint your agents point at. LiteLLM is an OpenAI-compatible proxy with virtual keys, per-key max_budget, and tpm_limit/rpm_limit rate limits in front of 100+ providers. This is the single most important box to install, because it is where a runaway agent gets stopped.

The provisioning layer gives you GPU capacity. OpenTofu or Terraform for raw control, or an internal developer platform when you want self-service with guardrails baked in.

The cost layer tells you where the money went. OpenCost is a CNCF incubating project (it reached incubation on October 25, 2024) that allocates spend by namespace, controller, label, and pod, including GPU time. Cloud-bill-level FinOps tooling is too coarse to answer per-model questions; OpenCost is not.

Kubernetes sits underneath all of this: the NVIDIA GPU Operator and device plugin expose the cards, dedicated GPU node pools isolate the workload, and KEDA or Knative give you scale-to-zero.

The most common mistake in this stack is buying an inference server and expecting cost governance to come with it. It does not.

LayerRepresentative toolsWhat it solvesWhat it does NOT solveWho owns it
ServingvLLM, SGLang, Ray Serve, OllamaRunning the model fast on a GPUBudgets, access control, attributionML platform / infra
GatewayLiteLLMPer-key budgets, rate limits, routing, the agent's endpointScheduling GPUs, provisioning infraPlatform team
ProvisioningOpenTofu/Terraform, QoveryGPU capacity, environments, RBAC, auto-stopToken-level budgeting, serving modelsPlatform / DevOps
AttributionOpenCost, KubecostCost per model/team/namespaceEnforcing anything (reports only)FinOps / platform

How do Claude Code and Codex safely trigger a deployment without cloud credentials?

The safe pattern is agent-opens-a-pull-request or agent-calls-a-scoped-platform-API, with the platform holding the cloud credentials and the agent holding none. Both keep a review gate and an audit log that a raw terraform apply from an agent can never give you.

Pattern A - the agent opens a PR. CI or the platform spins up an ephemeral preview environment on a GPU node pool, the agent inspects it, and the environment auto-expires. Review history is preserved for free.

Pattern B - the agent calls a scoped platform API (or an MCP server wrapping it) that exposes only a short allowlist: deploy service X to environment Y, redeploy, stop, fetch logs. Nothing else is reachable.

The anti-pattern is handing an agent an IAM user, a kubeconfig with cluster-admin, or shell access to run infra commands. The blast radius is unbounded and there is no review gate.

A few things that make either pattern hold up:

  • Credential hygiene: short-lived OIDC federation instead of static keys, one identity per agent, scoped RBAC per environment.
  • Human-in-the-loop on production promotion only. Gate dev and preview too and you kill the speed benefit that made you want agents in the first place.
  • Logging that ties it together: agent identity, triggering commit, target environment, and the resulting GPU resource request, so a spike is traceable in minutes, not days.

One more reason the interface must be allowlisted actions rather than arbitrary commands: an agent reading an untrusted GitHub issue or a poisoned dependency is a genuine escalation path. If the only thing the agent can do is call four named operations, prompt injection has nowhere to go.

A note on the tools themselves. Claude Code supports routing through an LLM gateway via ANTHROPIC_BASE_URL, which centralizes credentials, usage tracking, cost controls, and audit logging - though Anthropic only supports routing Claude Code to Claude models through it, not arbitrary open-weight ones. Codex is the more flexible client for self-hosted models: you can register a custom provider with any OpenAI-compatible base_url in config.toml, including a LiteLLM gateway sitting in front of your own vLLM endpoint.

Is self-hosting open source models actually cheaper than a hosted API?

Self-hosting wins on sustained, predictable load and loses badly on bursty or low-volume load, because a GPU instance bills by the hour whether it serves one token or a billion. Work out your break-even in tokens per hour before you provision anything.

The formula is simple: compare (GPU $/hour) / (tokens served per hour at realistic utilization) against the hosted $/million tokens. Let me plug in real prices.

As of early 2026, published on-demand pricing puts an NVIDIA H100 at roughly $3.99 per GPU-hour on Lambda, about $6.16 per GPU-hour on CoreWeave (their 8-GPU HGX H100 node lists at $49.24/hour), and around $6.88 per GPU-hour on AWS (the p5.48xlarge lists near $55/hour for eight cards). On the hosted side, Claude Haiku is $0.10 per million input tokens and $0.50 per million output, with Sonnet at $2/$10.

Now the arithmetic. Take one H100 at $4/hour. Assume it sustains 2,000 output tokens/second under real batched load for an 8B model (a deliberately conservative figure - benchmark your own, because vLLM can do better). That is 7.2M tokens/hour, so $4 / 7.2M = about $0.55 per million output tokens at full saturation. Roughly the same as hosted Haiku, and far below Sonnet.

Here is the catch, and it is the whole argument. Run that same $4/hour card at the utilization most clusters actually see - CAST AI measured average GPU utilization at just 5% in 2025 - and your effective cost jumps to around $11 per million tokens, more expensive than Sonnet and twenty times the hosted Haiku price. Idle is the dominant cost. A single 8xH100 node left running from Friday evening to Monday morning is about 63 hours; at CoreWeave's $49.24/hour that is roughly $3,100 of spend for zero work, which can exceed a small team's entire monthly hosted-API bill.

DimensionSelf-hosted open-weight on your GPUsHosted frontier API
Unit cost basis$/GPU-hour$/million tokens
Cost at low utilizationVery high (you pay for idle)Low (pay per token)
Cost at sustained loadLow (~$0.55/M at saturation)Higher per token
Idle cost exposureFull - bills whether used or notNone
Data residencyFull controlProvider region only
Model quality ceilingOpen-weight state of the artFrontier
Ops burdenYou run serving + infraZero
Scaling behaviourYou provision aheadElastic, instant
Price/supply volatilityYou hedge across providersProvider sets price
Time to first token in prodDays (provision + warm)Minutes (API key)

The normal end state is hybrid: route bulk, cheap tasks (commit messages, test generation, summaries) to a local open-weight model, send hard reasoning to a hosted frontier model, and do it all through one gateway so the routing is config, not code.

How do you stop an AI agent from running up a huge GPU bill?

Five controls do almost all of the work: per-key token and request budgets at the gateway, auto-stop or scale-to-zero on every non-production GPU environment, hard node-pool maximums, spot nodes for batch and eval, and per-namespace cost attribution so every model has a named owner.

  • Gateway budgets first. LiteLLM virtual keys with hard monthly caps and RPM/TPM limits stop a runaway loop in seconds, before the infrastructure layer even reacts. This is the cheapest, fastest-payback control you can install.
  • Auto-stop and scale-to-zero on dev and preview. The single biggest lever, because non-prod is idle most of the week and a 5% utilization average means you are paying for a lot of nothing.
  • Hard ceilings: node-pool max size, resource quotas per namespace, GPU limit ranges. The worst case should be bounded, not open-ended.
  • Right-sizing: quantization (AWQ, GPTQ, FP8) and smaller models often remove the need for the expensive GPU class entirely.
  • Spot or preemptible GPUs for batch, evals, and fine-tuning; keep on-demand or reserved for serving only.
  • Attribution: OpenCost or Kubecost labels per model, per team, per environment, so chargeback and anomaly detection are even possible.
  • Alert on cost anomalies daily, not at month-end. An agent can burn a month of budget over a weekend, and a month-end report tells you after the money is gone.

Worth saying plainly: OpenCost reports, it does not enforce. You need the gateway caps and the infrastructure ceilings to actually stop spend. And under a bring-your-own-cloud model, your committed-use discounts, Savings Plans, and reserved capacity stay in your own account and your own name.

Ship faster on infrastructure you control.
Qovery gives your team self-service deployments on your own AWS, GCP, Azure, or Scaleway account - or your existing Kubernetes cluster. Start deploying in under 10 minutes.

How do you avoid locking into one GPU cloud?

Keep the inference workload Kubernetes-native and the deployment path provider-agnostic, so the same container, manifest, and model spec run on AWS, GCP, Azure, Scaleway, a specialist GPU cloud, or your own on-prem cluster without a rewrite.

Portability is a cost strategy as much as a risk one. The same H100 class ranged from about $3.99 to $6.88 per GPU-hour across the providers above, and availability differs sharply by region. A second provider path lets you follow capacity and price instead of begging your account manager for quota.

The portable unit is a container image plus Kubernetes manifests plus an OpenAI-compatible endpoint. Everything above the gateway - your agents, your routing, your budgets - stays unchanged when you move the serving layer. Bring-your-own-Kubernetes should mean an existing EKS, GKE, or AKS cluster, a self-managed cluster, or on-prem hardware are all valid targets.

What genuinely locks you in: proprietary serverless GPU APIs, provider-specific model artifacts and endpoints, deeply coupled IAM, and managed inference services with no exportable spec. The honest trade-off is that fully portable means you own more ops, while managed serverless starts faster and hands control of the account and the bill to someone else. Given ongoing HBM supply constraints and GPU lead times, I would build the second path before I needed it, not during an outage.

Which tools handle this use case best, and how do they compare?

There is no single winner, so pick one tool per layer. Here is my read on each, including what each one does not do.

  • LiteLLM - the practical enforcement point for virtual keys, budgets, rate limits, and multi-provider routing. Not a scheduler, not an infrastructure tool.
  • vLLM - best-in-class open-source throughput for serving. No cost governance, no access control.
  • Ray Serve - excellent multi-model serving, composition, and autoscaling. Not a cost or RBAC layer.
  • OpenCost / Kubecost - the clearest Kubernetes cost attribution. Observability only; it enforces nothing.
  • OpenTofu / Terraform - complete control over GPU infrastructure, but every guardrail, review gate, and self-service path is something you build and then maintain.
  • AWS (EKS, SageMaker, Bedrock) - the deepest GPU ecosystem and capacity. The honest critique is assembly cost across many services and lock-in, not capability.
  • Anthropic and other hosted APIs - best model quality and zero ops, but not self-hosted and not an answer to data residency.
  • Baseten / Modal - legitimate choices for managed serverless GPU with fast cold starts. You trade control over the cloud account and the bill.
  • Heroku - genuinely good developer experience for standard web and API workloads, but not a fit for GPU inference at scale. Reach for it when you want a simple 12-factor app deployed, not when you need to serve a 70B model.
  • Qovery - the provisioning-and-guardrails layer inside your own AWS, GCP, Azure, or Scaleway account or your existing Kubernetes cluster: git-push deploys, preview environments per pull request, per-environment RBAC, auto-stop for non-production, and managed cluster upgrades.
ToolLayerSelf-hosted open-weightBuilt-in cost controlsRBACRuns in your cloud accountMulti-cloudBest for
LiteLLMGatewayYes (proxies any endpoint)Yes (per-key budgets, rate limits)Virtual keysYesYesThe agent's endpoint + spend caps
vLLMServingYesNoNoYesYesFast open-weight serving
Ray ServeServingYesNoNoYesYesMulti-model composition
OpenCost/KubecostAttributionN/ANo (reports only)NoYesYesCost per model/team
OpenTofu/TerraformProvisioningYes (you build it)NoVia cloud IAMYesYesRaw IaC control
AWS (EKS/SageMaker/Bedrock)All, assembledYesPartial (per service)IAMYour accountAWS onlyDeepest GPU ecosystem
Anthropic APIHosted modelNoUsage limitsAPI keysNoN/AFrontier quality, zero ops
BasetenServing (serverless)YesPartialPlatformVendor accountVendor-managedManaged serverless GPU
ModalServing (serverless)YesPartialPlatformVendor accountVendor-managedServerless GPU + fast cold start
HerokuApp PaaSNo (not GPU inference)Dyno-basedPlatform rolesNoNoStandard web apps
QoveryProvisioning + guardrailsVia your serving layerEnv auto-stop, quotasPer-environmentYes (BYOC)AWS/GCP/Azure/Scaleway/BYO-K8sSelf-service deploys with guardrails

A few combinations that work in practice. For a small team testing the waters: LiteLLM in front of a hosted API, no self-hosting yet. For a team with steady internal load: LiteLLM + vLLM + OpenCost on a platform layer that enforces RBAC and auto-stop. For a regulated shop that needs data residency across clouds: the same, Kubernetes-native, with a second provider path ready.

What does a reference setup look like end to end?

A minimal working setup has four moving parts: Claude Code and Codex point at a LiteLLM gateway with per-agent virtual keys and hard budgets; the gateway routes to vLLM services on GPU node pools in your own Kubernetes cluster; those services are deployed through a platform layer that enforces RBAC, preview environments, and auto-stop; OpenCost attributes spend per model and per team.

In words, the wiring is: agent -> gateway (budgets, routing, logging) -> vLLM on a GPU node pool -> cost attribution and alerts, with the platform layer holding every cloud credential and the agent holding none. When someone wants to test a new model version or vLLM config, the agent opens a PR, a preview environment spins up on a real GPU, and it auto-expires - no production exposure.

Day-one settings checklist:

  • Per-key monthly budget and TPM limit on every virtual key.
  • Non-prod auto-stop schedule (evenings and weekends at minimum).
  • Node-pool maximum size and a namespace GPU quota.
  • A daily cost anomaly alert.

Measure weekly: cost per million tokens served, GPU utilization, idle GPU hours, cost per environment, and the share of agent traffic served locally versus by a hosted API.

The rollout order that works: gateway first (cheapest, fastest payback), attribution second, self-hosted serving third, multi-cloud last. And the honest exit criterion - if your sustained load never saturates even one GPU for a few hours a day, skip all of this and stay on a hosted API behind the same gateway. The gateway is worth it either way.

Can Claude Code or Codex deploy infrastructure safely without cloud credentials?

Yes, and they should never hold cloud credentials directly. Give the agent a scoped platform API or let it open a pull request that CI deploys to an ephemeral environment. The platform holds the credentials, the agent holds an allowlist of named actions, and every deploy leaves an audit trail tied to a commit.

What is the cheapest way to self-host an open source LLM for internal team use?

Run a quantized small model (AWQ/GPTQ/FP8) on a single GPU with vLLM, behind a LiteLLM gateway, with scale-to-zero on non-production. The cheapest setup is the one that is not idle: at full saturation an 8B model can cost around $0.55 per million output tokens, but at the 5% GPU utilization most clusters actually see, that same card costs more than a hosted API. Right-size first, then worry about per-token price.

Do I need Kubernetes to self-host open source models for coding agents?

Not strictly - Ollama or a managed serverless GPU service like Baseten or Modal can serve models without it. But if you want portability across clouds, per-namespace cost attribution, GPU quotas, and auto-stop, Kubernetes is the layer where all of those controls exist, and the portable unit (container + manifests + OpenAI-compatible endpoint) moves between providers without a rewrite.

How do I stop an AI coding agent from running up a huge GPU bill?

Put hard budgets and rate limits on the agent's gateway key, and put auto-stop plus node-pool maximums on the GPU infrastructure. The gateway stops a runaway agent loop in seconds; the infrastructure ceilings bound the worst case. Add daily cost anomaly alerts, because an agent can burn a month of budget over a weekend.

Is LiteLLM enough for cost control, or do I also need OpenCost or Kubecost?

LiteLLM enforces token and request budgets at the gateway, which stops runaway agents. OpenCost or Kubecost does something different: it attributes the underlying GPU infrastructure cost to a model, team, or namespace so you know where money actually went. You want both - one enforces, the other explains.

When should a platform team stay on hosted APIs instead of self-hosting open-weight models?

Stay hosted when your traffic is spiky or low-volume, when you need frontier model quality for hard reasoning, or when you have no appetite for GPU ops. Self-hosting only pays off at sustained, predictable utilization or when data residency and compliance force your hand. If you cannot keep one GPU busy for a few hours a day, a hosted API behind a gateway is both cheaper and simpler.

Romaric Philogene
About the author
Romaric Philogene

Romaric founded Qovery to make Kubernetes accessible to every engineering team. He writes about platform strategy, developer experience, and the future of cloud infrastructure.

Next step

Ship faster on infrastructure you control.

Qovery gives your team self-service deployments on your own AWS, GCP, Azure, or Scaleway account - or your existing Kubernetes cluster. Start deploying in under 10 minutes.