vLLM on OpenShift

Serve models on OpenShift with vLLM — an OpenAI-compatible endpoint that any agent (claude-code, codex, opencode) or cap-evolve adapter can call. No external API keys needed.

Architecture

┌─────────────────────────────────────────────────────────────┐
│  OpenShift cluster  (namespace: cap-evolve)                  │
│                                                              │
│  vLLM Deployment (GPU node)                                  │
│    └─ Serves: http://vllm-svc.cap-evolve.svc:8000/v1        │
│                                                              │
│  Consumers:                                                  │
│    cap-evolve runner   → MODEL=openai/<model-name>         │
│    harbor task pods    → --ae ANTHROPIC_BASE_URL=http://...   │
│    any OpenAI client   → curl http://vllm-svc:8000/v1/chat   │
└─────────────────────────────────────────────────────────────┘

Deploy

# 1. Namespace
oc new-project cap-evolve    # or use an existing namespace

# 2. HuggingFace token (for gated models)
oc create secret generic hf-token \
  --from-literal=HF_TOKEN=<YOUR_HF_TOKEN>

# 3. Deploy vLLM
oc apply -f openshift/manifests/vllm-serving.yaml

# 4. Wait for the model to load (5-15 min first time)
oc get pods -l component=vllm -w

# 5. Verify
curl -s http://vllm-svc:8000/v1/models

Model sizing

SizeGPUs--tensor-parallel-sizeMemory
7B11 (default)8 Gi
14B2232 Gi
32B+4464 Gi

Edit vllm-serving.yaml to change the model, GPU count, and --tensor-parallel-size.

Connecting agents to vLLM

vLLM exposes an OpenAI-compatible API. Agents connect via environment variables — the exact vars depend on the agent.

AgentEnv vars
cap-evolve (litellm)MODEL=openai/<model> OPENAI_API_BASE=http://vllm-svc:8000/v1 OPENAI_API_KEY=dummy
claude-codeANTHROPIC_BASE_URL=http://vllm-svc:8000 ANTHROPIC_API_KEY=dummy ANTHROPIC_MODEL=<model>
codex / opencodeOPENAI_BASE_URL=http://vllm-svc:8000/v1 OPENAI_API_KEY=dummy

For Harbor task pods, pass these via --ae flags. Use the full Service DNS when crossing namespaces: http://vllm-svc.<namespace>.svc:8000.

Manifests

All manifests live under openshift/:

FilePurpose
namespace.yamlNamespace
service-account.yamlServiceAccount + SCC for Harbor task pods
vllm-serving.yamlvLLM Deployment + Service + Route (agent model)
vllm-optimizer.yamlOptional second vLLM for the optimizer (larger model)
pvc.yamlPersistent storage for run artifacts
secrets.yamlTemplate for HuggingFace token

Troubleshooting

ProblemFix
Pod pending (no GPU)Check oc describe pod — node selector may not match your GPU labels
OOMKilledModel too large for requested memory — increase resources or use fewer GPUs with a smaller model
Context length exceededSet --max-model-len in the vLLM args, or use a model with larger native context
Connection refused from task podUse full DNS: http://<svc>.<namespace>.svc:8000

References