vLLM on OpenShift
Serve models on OpenShift with vLLM — an OpenAI-compatible endpoint that any agent (claude-code, codex, opencode) or cap-evolve adapter can call. No external API keys needed.
Architecture
┌─────────────────────────────────────────────────────────────┐
│ OpenShift cluster (namespace: cap-evolve) │
│ │
│ vLLM Deployment (GPU node) │
│ └─ Serves: http://vllm-svc.cap-evolve.svc:8000/v1 │
│ │
│ Consumers: │
│ cap-evolve runner → MODEL=openai/<model-name> │
│ harbor task pods → --ae ANTHROPIC_BASE_URL=http://... │
│ any OpenAI client → curl http://vllm-svc:8000/v1/chat │
└─────────────────────────────────────────────────────────────┘
Deploy
# 1. Namespace
oc new-project cap-evolve # or use an existing namespace
# 2. HuggingFace token (for gated models)
oc create secret generic hf-token \
--from-literal=HF_TOKEN=<YOUR_HF_TOKEN>
# 3. Deploy vLLM
oc apply -f openshift/manifests/vllm-serving.yaml
# 4. Wait for the model to load (5-15 min first time)
oc get pods -l component=vllm -w
# 5. Verify
curl -s http://vllm-svc:8000/v1/models
Model sizing
| Size | GPUs | --tensor-parallel-size | Memory |
|---|---|---|---|
| 7B | 1 | 1 (default) | 8 Gi |
| 14B | 2 | 2 | 32 Gi |
| 32B+ | 4 | 4 | 64 Gi |
Edit vllm-serving.yaml to change the model, GPU count, and
--tensor-parallel-size.
Connecting agents to vLLM
vLLM exposes an OpenAI-compatible API. Agents connect via environment variables — the exact vars depend on the agent.
| Agent | Env vars |
|---|---|
| cap-evolve (litellm) | MODEL=openai/<model> OPENAI_API_BASE=http://vllm-svc:8000/v1 OPENAI_API_KEY=dummy |
| claude-code | ANTHROPIC_BASE_URL=http://vllm-svc:8000 ANTHROPIC_API_KEY=dummy ANTHROPIC_MODEL=<model> |
| codex / opencode | OPENAI_BASE_URL=http://vllm-svc:8000/v1 OPENAI_API_KEY=dummy |
For Harbor task pods, pass these via --ae flags. Use the full
Service DNS when crossing namespaces:
http://vllm-svc.<namespace>.svc:8000.
Manifests
All manifests live under
openshift/:
| File | Purpose |
|---|---|
namespace.yaml | Namespace |
service-account.yaml | ServiceAccount + SCC for Harbor task pods |
vllm-serving.yaml | vLLM Deployment + Service + Route (agent model) |
vllm-optimizer.yaml | Optional second vLLM for the optimizer (larger model) |
pvc.yaml | Persistent storage for run artifacts |
secrets.yaml | Template for HuggingFace token |
Troubleshooting
| Problem | Fix |
|---|---|
| Pod pending (no GPU) | Check oc describe pod — node selector may not match your GPU labels |
| OOMKilled | Model too large for requested memory — increase resources or use fewer GPUs with a smaller model |
| Context length exceeded | Set --max-model-len in the vLLM args, or use a model with larger native context |
| Connection refused from task pod | Use full DNS: http://<svc>.<namespace>.svc:8000 |
References
- Harbor integration — using Harbor with cap-evolve (includes OpenShift orchestrator setup)
- vLLM documentation
- OpenShift manifests