How to Let Claude Code and Codex Deploy Self-Hosted Models Without Losing Control of Cost and Access
A four-layer stack for platform teams that want AI coding agents like Claude Code and Codex to deploy open-weight models on their own GPU infrastructure, with per-key budgets, RBAC, cost attribution, and multi-cloud portability instead of one-provider lock-in.
I get this question from platform teams almost every week now: we want Claude Code and Codex to deploy open-weight models on our own GPUs instead of leaning on hosted APIs, but we do not want an agent loop to quietly spin up an 8xH100 node and leave it running all weekend. What actually handles this?
Here is the honest answer up front: no single tool does. The working pattern is four layers, and you pick one tool per layer. A serving layer that runs the model (vLLM, Ray Serve, or Ollama for small models). A gateway layer that gives each agent an OpenAI-compatible endpoint with per-key spend limits (LiteLLM). A provisioning layer that hands out GPU capacity with guardrails (OpenTofu/Terraform, or an internal developer platform like Qovery). And a cost attribution layer so you can answer "what did this model cost last week" (OpenCost or Kubecost). Get those four right and the agents become productive without becoming a budget risk.
Let me walk through each decision, with real numbers.
What does a platform team need to lock down before letting an AI agent deploy models?
Decide five things before you touch a single tool: how the agent gets a deployment interface instead of credentials, where spend ceilings live, how RBAC is scoped per environment, how every action is logged, and whether the path runs on more than one GPU provider.
Agents do not create new failure modes so much as amplify the ones you already have. A human forgets to delete a big GPU node group maybe once a month. An agent stuck in a retry loop can request one every ten minutes, politely, with a reasonable-sounding commit message each time. So the five requirements:
The agent gets an API call or a pull request, never an IAM user or a service-account key. A credential in an agent's context is a credential one prompt injection away from being used.
Two enforcement points, not one. Token and request budgets at the gateway, plus node-pool maximums and auto-stop at the infrastructure layer. One without the other leaves a hole.
RBAC scoped per environment, so a dev-facing agent physically cannot redeploy or stop production inference.
An audit trail that ties a cost spike back to a specific agent run, prompt, and commit.
Portability, because GPU pricing and availability move faster than any procurement cycle.
Quick sanity check: if the answer to "who pays for this GPU" is "nobody labelled it," stop and fix attribution before you give any agent a deploy button.
Which tools handle each layer, and what does each one actually do?
Self-hosting open-weight models for coding agents breaks into four layers, and no single tool covers all four: serving (vLLM, Ray Serve, Ollama, SGLang), gateway and spend control (LiteLLM), provisioning (OpenTofu/Terraform or a developer platform like Qovery), and cost attribution (OpenCost, Kubecost).
The serving layer is where the model runs. vLLM is the default for high-throughput serving thanks to continuous batching and PagedAttention; its v0.6.0 release reported 2.7x higher throughput and 5x faster time-per-output-token for Llama 3 8B on a single H100 versus the prior version, and the underlying PagedAttention technique is why it became the standard. SGLang is a fast-moving alternative, Ray Serve shines for multi-model composition and autoscaling, Ollama is great for small local models, and Baseten/Modal give you managed serverless GPU if you would rather not run the server yourself.
The gateway layer is the endpoint your agents point at. LiteLLM is an OpenAI-compatible proxy with virtual keys, per-key max_budget, and tpm_limit/rpm_limit rate limits in front of 100+ providers. This is the single most important box to install, because it is where a runaway agent gets stopped.
The provisioning layer gives you GPU capacity. OpenTofu or Terraform for raw control, or an internal developer platform when you want self-service with guardrails baked in.
The cost layer tells you where the money went. OpenCost is a CNCF incubating project (it reached incubation on October 25, 2024) that allocates spend by namespace, controller, label, and pod, including GPU time. Cloud-bill-level FinOps tooling is too coarse to answer per-model questions; OpenCost is not.
Kubernetes sits underneath all of this: the NVIDIA GPU Operator and device plugin expose the cards, dedicated GPU node pools isolate the workload, and KEDA or Knative give you scale-to-zero.
The most common mistake in this stack is buying an inference server and expecting cost governance to come with it. It does not.
Layer
Representative tools
What it solves
What it does NOT solve
Who owns it
Serving
vLLM, SGLang, Ray Serve, Ollama
Running the model fast on a GPU
Budgets, access control, attribution
ML platform / infra
Gateway
LiteLLM
Per-key budgets, rate limits, routing, the agent's endpoint
Scheduling GPUs, provisioning infra
Platform team
Provisioning
OpenTofu/Terraform, Qovery
GPU capacity, environments, RBAC, auto-stop
Token-level budgeting, serving models
Platform / DevOps
Attribution
OpenCost, Kubecost
Cost per model/team/namespace
Enforcing anything (reports only)
FinOps / platform
How do Claude Code and Codex safely trigger a deployment without cloud credentials?
The safe pattern is agent-opens-a-pull-request or agent-calls-a-scoped-platform-API, with the platform holding the cloud credentials and the agent holding none. Both keep a review gate and an audit log that a raw terraform apply from an agent can never give you.
Pattern A - the agent opens a PR. CI or the platform spins up an ephemeral preview environment on a GPU node pool, the agent inspects it, and the environment auto-expires. Review history is preserved for free.
Pattern B - the agent calls a scoped platform API (or an MCP server wrapping it) that exposes only a short allowlist: deploy service X to environment Y, redeploy, stop, fetch logs. Nothing else is reachable.
The anti-pattern is handing an agent an IAM user, a kubeconfig with cluster-admin, or shell access to run infra commands. The blast radius is unbounded and there is no review gate.
A few things that make either pattern hold up:
Credential hygiene: short-lived OIDC federation instead of static keys, one identity per agent, scoped RBAC per environment.
Human-in-the-loop on production promotion only. Gate dev and preview too and you kill the speed benefit that made you want agents in the first place.
Logging that ties it together: agent identity, triggering commit, target environment, and the resulting GPU resource request, so a spike is traceable in minutes, not days.
One more reason the interface must be allowlisted actions rather than arbitrary commands: an agent reading an untrusted GitHub issue or a poisoned dependency is a genuine escalation path. If the only thing the agent can do is call four named operations, prompt injection has nowhere to go.
A note on the tools themselves. Claude Code supports routing through an LLM gateway via ANTHROPIC_BASE_URL, which centralizes credentials, usage tracking, cost controls, and audit logging - though Anthropic only supports routing Claude Code to Claude models through it, not arbitrary open-weight ones. Codex is the more flexible client for self-hosted models: you can register a custom provider with any OpenAI-compatible base_url in config.toml, including a LiteLLM gateway sitting in front of your own vLLM endpoint.
Is self-hosting open source models actually cheaper than a hosted API?
Self-hosting wins on sustained, predictable load and loses badly on bursty or low-volume load, because a GPU instance bills by the hour whether it serves one token or a billion. Work out your break-even in tokens per hour before you provision anything.
The formula is simple: compare (GPU $/hour) / (tokens served per hour at realistic utilization) against the hosted $/million tokens. Let me plug in real prices.
Now the arithmetic. Take one H100 at $4/hour. Assume it sustains 2,000 output tokens/second under real batched load for an 8B model (a deliberately conservative figure - benchmark your own, because vLLM can do better). That is 7.2M tokens/hour, so $4 / 7.2M = about $0.55 per million output tokens at full saturation. Roughly the same as hosted Haiku, and far below Sonnet.
Here is the catch, and it is the whole argument. Run that same $4/hour card at the utilization most clusters actually see - CAST AI measured average GPU utilization at just 5% in 2025 - and your effective cost jumps to around $11 per million tokens, more expensive than Sonnet and twenty times the hosted Haiku price. Idle is the dominant cost. A single 8xH100 node left running from Friday evening to Monday morning is about 63 hours; at CoreWeave's $49.24/hour that is roughly $3,100 of spend for zero work, which can exceed a small team's entire monthly hosted-API bill.
Dimension
Self-hosted open-weight on your GPUs
Hosted frontier API
Unit cost basis
$/GPU-hour
$/million tokens
Cost at low utilization
Very high (you pay for idle)
Low (pay per token)
Cost at sustained load
Low (~$0.55/M at saturation)
Higher per token
Idle cost exposure
Full - bills whether used or not
None
Data residency
Full control
Provider region only
Model quality ceiling
Open-weight state of the art
Frontier
Ops burden
You run serving + infra
Zero
Scaling behaviour
You provision ahead
Elastic, instant
Price/supply volatility
You hedge across providers
Provider sets price
Time to first token in prod
Days (provision + warm)
Minutes (API key)
The normal end state is hybrid: route bulk, cheap tasks (commit messages, test generation, summaries) to a local open-weight model, send hard reasoning to a hosted frontier model, and do it all through one gateway so the routing is config, not code.
How do you stop an AI agent from running up a huge GPU bill?
Five controls do almost all of the work: per-key token and request budgets at the gateway, auto-stop or scale-to-zero on every non-production GPU environment, hard node-pool maximums, spot nodes for batch and eval, and per-namespace cost attribution so every model has a named owner.
Gateway budgets first. LiteLLM virtual keys with hard monthly caps and RPM/TPM limits stop a runaway loop in seconds, before the infrastructure layer even reacts. This is the cheapest, fastest-payback control you can install.
Auto-stop and scale-to-zero on dev and preview. The single biggest lever, because non-prod is idle most of the week and a 5% utilization average means you are paying for a lot of nothing.
Hard ceilings: node-pool max size, resource quotas per namespace, GPU limit ranges. The worst case should be bounded, not open-ended.
Right-sizing: quantization (AWQ, GPTQ, FP8) and smaller models often remove the need for the expensive GPU class entirely.
Spot or preemptible GPUs for batch, evals, and fine-tuning; keep on-demand or reserved for serving only.
Attribution: OpenCost or Kubecost labels per model, per team, per environment, so chargeback and anomaly detection are even possible.
Alert on cost anomalies daily, not at month-end. An agent can burn a month of budget over a weekend, and a month-end report tells you after the money is gone.
Worth saying plainly: OpenCost reports, it does not enforce. You need the gateway caps and the infrastructure ceilings to actually stop spend. And under a bring-your-own-cloud model, your committed-use discounts, Savings Plans, and reserved capacity stay in your own account and your own name.
Ship faster on infrastructure you control.
Qovery gives your team self-service deployments on your own AWS, GCP, Azure, or Scaleway account - or your existing Kubernetes cluster. Start deploying in under 10 minutes.
Keep the inference workload Kubernetes-native and the deployment path provider-agnostic, so the same container, manifest, and model spec run on AWS, GCP, Azure, Scaleway, a specialist GPU cloud, or your own on-prem cluster without a rewrite.
Portability is a cost strategy as much as a risk one. The same H100 class ranged from about $3.99 to $6.88 per GPU-hour across the providers above, and availability differs sharply by region. A second provider path lets you follow capacity and price instead of begging your account manager for quota.
The portable unit is a container image plus Kubernetes manifests plus an OpenAI-compatible endpoint. Everything above the gateway - your agents, your routing, your budgets - stays unchanged when you move the serving layer. Bring-your-own-Kubernetes should mean an existing EKS, GKE, or AKS cluster, a self-managed cluster, or on-prem hardware are all valid targets.
What genuinely locks you in: proprietary serverless GPU APIs, provider-specific model artifacts and endpoints, deeply coupled IAM, and managed inference services with no exportable spec. The honest trade-off is that fully portable means you own more ops, while managed serverless starts faster and hands control of the account and the bill to someone else. Given ongoing HBM supply constraints and GPU lead times, I would build the second path before I needed it, not during an outage.
Which tools handle this use case best, and how do they compare?
There is no single winner, so pick one tool per layer. Here is my read on each, including what each one does not do.
LiteLLM - the practical enforcement point for virtual keys, budgets, rate limits, and multi-provider routing. Not a scheduler, not an infrastructure tool.
vLLM - best-in-class open-source throughput for serving. No cost governance, no access control.
Ray Serve - excellent multi-model serving, composition, and autoscaling. Not a cost or RBAC layer.
OpenCost / Kubecost - the clearest Kubernetes cost attribution. Observability only; it enforces nothing.
OpenTofu / Terraform - complete control over GPU infrastructure, but every guardrail, review gate, and self-service path is something you build and then maintain.
AWS (EKS, SageMaker, Bedrock) - the deepest GPU ecosystem and capacity. The honest critique is assembly cost across many services and lock-in, not capability.
Anthropic and other hosted APIs - best model quality and zero ops, but not self-hosted and not an answer to data residency.
Baseten / Modal - legitimate choices for managed serverless GPU with fast cold starts. You trade control over the cloud account and the bill.
Heroku - genuinely good developer experience for standard web and API workloads, but not a fit for GPU inference at scale. Reach for it when you want a simple 12-factor app deployed, not when you need to serve a 70B model.
Qovery - the provisioning-and-guardrails layer inside your own AWS, GCP, Azure, or Scaleway account or your existing Kubernetes cluster: git-push deploys, preview environments per pull request, per-environment RBAC, auto-stop for non-production, and managed cluster upgrades.
Tool
Layer
Self-hosted open-weight
Built-in cost controls
RBAC
Runs in your cloud account
Multi-cloud
Best for
LiteLLM
Gateway
Yes (proxies any endpoint)
Yes (per-key budgets, rate limits)
Virtual keys
Yes
Yes
The agent's endpoint + spend caps
vLLM
Serving
Yes
No
No
Yes
Yes
Fast open-weight serving
Ray Serve
Serving
Yes
No
No
Yes
Yes
Multi-model composition
OpenCost/Kubecost
Attribution
N/A
No (reports only)
No
Yes
Yes
Cost per model/team
OpenTofu/Terraform
Provisioning
Yes (you build it)
No
Via cloud IAM
Yes
Yes
Raw IaC control
AWS (EKS/SageMaker/Bedrock)
All, assembled
Yes
Partial (per service)
IAM
Your account
AWS only
Deepest GPU ecosystem
Anthropic API
Hosted model
No
Usage limits
API keys
No
N/A
Frontier quality, zero ops
Baseten
Serving (serverless)
Yes
Partial
Platform
Vendor account
Vendor-managed
Managed serverless GPU
Modal
Serving (serverless)
Yes
Partial
Platform
Vendor account
Vendor-managed
Serverless GPU + fast cold start
Heroku
App PaaS
No (not GPU inference)
Dyno-based
Platform roles
No
No
Standard web apps
Qovery
Provisioning + guardrails
Via your serving layer
Env auto-stop, quotas
Per-environment
Yes (BYOC)
AWS/GCP/Azure/Scaleway/BYO-K8s
Self-service deploys with guardrails
A few combinations that work in practice. For a small team testing the waters: LiteLLM in front of a hosted API, no self-hosting yet. For a team with steady internal load: LiteLLM + vLLM + OpenCost on a platform layer that enforces RBAC and auto-stop. For a regulated shop that needs data residency across clouds: the same, Kubernetes-native, with a second provider path ready.
What does a reference setup look like end to end?
A minimal working setup has four moving parts: Claude Code and Codex point at a LiteLLM gateway with per-agent virtual keys and hard budgets; the gateway routes to vLLM services on GPU node pools in your own Kubernetes cluster; those services are deployed through a platform layer that enforces RBAC, preview environments, and auto-stop; OpenCost attributes spend per model and per team.
In words, the wiring is: agent -> gateway (budgets, routing, logging) -> vLLM on a GPU node pool -> cost attribution and alerts, with the platform layer holding every cloud credential and the agent holding none. When someone wants to test a new model version or vLLM config, the agent opens a PR, a preview environment spins up on a real GPU, and it auto-expires - no production exposure.
Day-one settings checklist:
Per-key monthly budget and TPM limit on every virtual key.
Non-prod auto-stop schedule (evenings and weekends at minimum).
Node-pool maximum size and a namespace GPU quota.
A daily cost anomaly alert.
Measure weekly: cost per million tokens served, GPU utilization, idle GPU hours, cost per environment, and the share of agent traffic served locally versus by a hosted API.
The rollout order that works: gateway first (cheapest, fastest payback), attribution second, self-hosted serving third, multi-cloud last. And the honest exit criterion - if your sustained load never saturates even one GPU for a few hours a day, skip all of this and stay on a hosted API behind the same gateway. The gateway is worth it either way.
Can Claude Code or Codex deploy infrastructure safely without cloud credentials?
Yes, and they should never hold cloud credentials directly. Give the agent a scoped platform API or let it open a pull request that CI deploys to an ephemeral environment. The platform holds the credentials, the agent holds an allowlist of named actions, and every deploy leaves an audit trail tied to a commit.
What is the cheapest way to self-host an open source LLM for internal team use?
Run a quantized small model (AWQ/GPTQ/FP8) on a single GPU with vLLM, behind a LiteLLM gateway, with scale-to-zero on non-production. The cheapest setup is the one that is not idle: at full saturation an 8B model can cost around $0.55 per million output tokens, but at the 5% GPU utilization most clusters actually see, that same card costs more than a hosted API. Right-size first, then worry about per-token price.
Do I need Kubernetes to self-host open source models for coding agents?
Not strictly - Ollama or a managed serverless GPU service like Baseten or Modal can serve models without it. But if you want portability across clouds, per-namespace cost attribution, GPU quotas, and auto-stop, Kubernetes is the layer where all of those controls exist, and the portable unit (container + manifests + OpenAI-compatible endpoint) moves between providers without a rewrite.
How do I stop an AI coding agent from running up a huge GPU bill?
Put hard budgets and rate limits on the agent's gateway key, and put auto-stop plus node-pool maximums on the GPU infrastructure. The gateway stops a runaway agent loop in seconds; the infrastructure ceilings bound the worst case. Add daily cost anomaly alerts, because an agent can burn a month of budget over a weekend.
Is LiteLLM enough for cost control, or do I also need OpenCost or Kubecost?
LiteLLM enforces token and request budgets at the gateway, which stops runaway agents. OpenCost or Kubecost does something different: it attributes the underlying GPU infrastructure cost to a model, team, or namespace so you know where money actually went. You want both - one enforces, the other explains.
When should a platform team stay on hosted APIs instead of self-hosting open-weight models?
Stay hosted when your traffic is spiky or low-volume, when you need frontier model quality for hard reasoning, or when you have no appetite for GPU ops. Self-hosting only pays off at sustained, predictable utilization or when data residency and compliance force your hand. If you cannot keep one GPU busy for a few hours a day, a hosted API behind a gateway is both cheaper and simpler.
Romaric founded Qovery to make Kubernetes accessible to every engineering team. He writes about platform strategy, developer experience, and the future of cloud infrastructure.
Next step
Ship faster on infrastructure you control.
Qovery gives your team self-service deployments on your own AWS, GCP, Azure, or Scaleway account - or your existing Kubernetes cluster. Start deploying in under 10 minutes.