OpenShift + vLLM
AI Cost Firewall v0.8.2 includes OpenShift-specific deployment assets under deploy/openshift/. The AIF image itself remains a generic OCI image and the normal Docker/Compose deployment remains supported.
A common topology is:
Open WebUI / OpenAI-compatible client
-> AIF
-> vLLM chat service -> chat model
-> vLLM embedding service -> embedding model -> Qdrant
The chat and embedding endpoints may be separate services. This is useful when the chat model and embedding model have different GPU/runtime requirements.
OpenShift security posture
The provided base Deployment is intended to run under OpenShift restricted-v2 without requesting anyuid or a custom SCC. The example:
- does not set a fixed
runAsUserorrunAsGroup; - drops all Linux capabilities;
- disables privilege escalation;
- uses a read-only root filesystem;
- uses
RuntimeDefaultseccomp; - does not mount a service-account token.
Outside OpenShift, the image retains its normal numeric non-root fallback user.
Configure secrets and services
Create the Secret before applying the base Kustomization:
cp deploy/openshift/examples/secret.example.yaml /tmp/ai-firewall-secret.yaml
# edit /tmp/ai-firewall-secret.yaml
oc apply -f /tmp/ai-firewall-secret.yaml
Review deploy/openshift/base/configmap.yaml and replace the example chat and embedding service names with the cluster-local vLLM Services. Then apply:
oc apply -k deploy/openshift/base
Start with chat/inference first
For an initial integration, start AIF in Observe mode and keep semantic cache disabled until the embedding service is verified.
AIF_ENFORCEMENT_MODE=observe
AIF_SEMANTIC_CACHE_ENABLED=false
Validate the path:
Open WebUI -> AIF -> vLLM chat
Check AIF probes and model discovery:
curl -i http://<aif-service>/healthz
curl -i http://<aif-service>/startupz
curl -i http://<aif-service>/readyz
curl -s http://<aif-service>/v1/models
GET /v1/models is proxied to the configured chat/inference vLLM endpoint. It does not expose the independently configured embedding service.
Verify the embedding service before semantic cache
Query the embedding vLLM service directly to obtain the exact served model ID:
curl -s http://vllm-embeddings:8000/v1/models | jq
Then send one OpenAI-compatible embedding request and count the returned vector elements. Configure:
AIF_EMBEDDING_MODEL=<exact-served-model-id>
AIF_QDRANT_VECTOR_SIZE=<actual-returned-vector-length>
Do not assume a Nomic or other embedding-model dimension from the family name alone. The exact model/version and serving configuration determine the output expected by Qdrant.
After verification, enable semantic cache in Observe mode and validate evaluation hit/miss/error and prospective savings metrics before changing AIF to enforce.
Probe behavior
Use:
/healthzfor process liveness;/startupzfor strict startup initialization of enabled Redis/Qdrant caches marked readiness-required;/readyzfor normal traffic readiness.
If AIF starts before a readiness-required Redis/Qdrant backend and fail-open/Observe startup creates a no-op cache, /startupz returns 503. The orchestrator should restart the pod after the dependency becomes available so AIF initializes the real cache client.
Content arrays and guards
AIF v0.8.2 can preserve OpenAI-style non-string message content for upstream forwarding and exact cache. Such requests bypass semantic cache. Current Security, Privacy, and Usage Guard integrations inspect string message content only, so nested text inside content arrays is not yet independently processed by those modules.
Optional NetworkPolicy and ServiceMonitor
Example assets under deploy/openshift/examples/ are not applied automatically. Adapt selectors and egress rules to the target cluster. Do not enable a default-deny egress policy until DNS, Redis, Qdrant, chat vLLM, embedding vLLM, Audit, and any enabled guard-service paths have been enumerated.