What is AI Cost Firewall?
These docs describe AI Cost Firewall v0.8.2 — Deployment Hardening. v0.8.2 keeps the Evaluation / Observe Mode introduced in v0.8.0 and hardens the same gateway for container and orchestrator deployments.
AI Cost Firewall is a self-hosted, OpenAI-compatible control plane for production AI traffic.
It sits between your application and an upstream LLM provider and gives operators a single enforcement point for request routing, exact and semantic caching, cost visibility, optional security and privacy controls, and structured execution evidence.
Applications send requests to AI Cost Firewall instead of calling the provider directly. In normal enforce mode, the firewall can reuse eligible cached responses and forward only misses upstream. In observe mode, it evaluates the same cache decisions using isolated shadow state while continuing to call the live upstream provider for every eligible request. v0.8.2 also adds stricter startup probing, safer cache identity for OpenAI-compatible extension fields, model discovery proxying, and deployment-specific OpenShift assets.
AI Cost Firewall supports practical OpenAI-compatible upstream and embedding endpoints, including cloud APIs, local gateways, and self-hosted model servers.
Why it exists
Production AI systems need more than a direct SDK call to a model provider. Teams need to control unnecessary model spend, keep sensitive data and malicious instructions away from upstream models, understand how requests were processed, and retain evidence for later investigation or compliance workflows.
AI Cost Firewall provides that control point while preserving the OpenAI-compatible application integration pattern.
Two cache layers reduce avoidable inference:
- Exact cache — reuses responses for identical normalized requests.
- Semantic cache — reuses responses for similar prompts when similarity is high enough.
Optional VCAL modules extend the same request path with security, privacy, organizational usage-policy, audit, and compliance capabilities.
Core capabilities
- OpenAI-compatible
POST /v1/chat/completionsendpoint - OpenAI-compatible
GET /v1/modelsproxy to the configured chat upstream - controlled OpenAI-compatible chat-completion streaming with full-response controls before client commit
- Redis / Valkey exact caching
- Qdrant semantic caching
- AIF-level
enforceand non-disruptiveobservemodes - isolated shadow exact/semantic cache state for production evaluations
- per-model cost, gross-savings, net-savings, and token-savings metrics
- optional VCAL Security Guard request/response inspection
- optional VCAL Privacy Guard anonymization, redaction, and restore processing
- optional VCAL Usage Guard organizational usage-policy evaluation
- structured
vcal.evidence.eventschema v1.1 events with stable trace correlation - optional buffered HTTP evidence delivery to VCAL Audit
- optional downstream VCAL Compliance evidence processing
- Prometheus metrics and Grafana dashboards
- strict configuration validation and model allowlist behavior through
model_price - request size limits and explicit timeout controls
- liveness, startup, and readiness endpoints (
/healthz,/startupz,/readyz) - graceful shutdown and validated hot reload via
SIGHUP - configurable fail-open / fail-closed behavior for cache and guard dependency failures
- semantic cache lifecycle controls
- exact-cache identity that preserves OpenAI-compatible request/message extension fields
- safe pass-through of non-string message content with semantic-cache bypass
- generic non-root OCI image hardening
- Docker Compose deployment examples
- OpenShift
restricted-v2deployment assets underdeploy/openshift/ /versionendpoint for release and compatibility introspection
Deployment model
AI Cost Firewall can run by itself as an OpenAI-compatible caching and cost-control gateway or as the request-processing core of the broader VCAL stack.
Application
-> AI Cost Firewall
-> Security Guard (optional)
-> Privacy Guard (optional)
-> Usage Guard (optional)
-> exact / semantic cache or upstream LLM
-> Audit evidence delivery (optional)
-> Compliance (optional)
Prometheus
-> Grafana for deep metrics and diagnostics
The control plane is self-hosted and customer-controlled. Model providers can change without requiring the application to integrate directly with every provider-specific control layer.
Deployment examples
Ready-to-run deployment patterns are available under:
deploy/examples/
These include OpenAI cloud, local Ollama, hybrid OpenAI + local embeddings, OpenRouter, and a full local stack with dashboards. OpenShift-specific manifests are kept separately under deploy/openshift/ so the runtime image and Docker Compose path remain platform-neutral.
Recommended reading path
- What it does
- Quick Start with Docker
- OpenShift + vLLM (when deploying on OpenShift)
- Evaluation Mode
- Request flow
- Architecture overview
- Configuration overview
- Runtime overview
- Metrics
- Dashboards