Skip to main content

What is AI Cost Firewall?

Current release

These docs describe AI Cost Firewall v0.8.2 — Deployment Hardening. v0.8.2 keeps the Evaluation / Observe Mode introduced in v0.8.0 and hardens the same gateway for container and orchestrator deployments.

AI Cost Firewall is a self-hosted, OpenAI-compatible control plane for production AI traffic.

It sits between your application and an upstream LLM provider and gives operators a single enforcement point for request routing, exact and semantic caching, cost visibility, optional security and privacy controls, and structured execution evidence.

Applications send requests to AI Cost Firewall instead of calling the provider directly. In normal enforce mode, the firewall can reuse eligible cached responses and forward only misses upstream. In observe mode, it evaluates the same cache decisions using isolated shadow state while continuing to call the live upstream provider for every eligible request. v0.8.2 also adds stricter startup probing, safer cache identity for OpenAI-compatible extension fields, model discovery proxying, and deployment-specific OpenShift assets.

AI Cost Firewall supports practical OpenAI-compatible upstream and embedding endpoints, including cloud APIs, local gateways, and self-hosted model servers.

Why it exists​

Production AI systems need more than a direct SDK call to a model provider. Teams need to control unnecessary model spend, keep sensitive data and malicious instructions away from upstream models, understand how requests were processed, and retain evidence for later investigation or compliance workflows.

AI Cost Firewall provides that control point while preserving the OpenAI-compatible application integration pattern.

Two cache layers reduce avoidable inference:

  1. Exact cache — reuses responses for identical normalized requests.
  2. Semantic cache — reuses responses for similar prompts when similarity is high enough.

Optional VCAL modules extend the same request path with security, privacy, organizational usage-policy, audit, and compliance capabilities.

Core capabilities​

  • OpenAI-compatible POST /v1/chat/completions endpoint
  • OpenAI-compatible GET /v1/models proxy to the configured chat upstream
  • controlled OpenAI-compatible chat-completion streaming with full-response controls before client commit
  • Redis / Valkey exact caching
  • Qdrant semantic caching
  • AIF-level enforce and non-disruptive observe modes
  • isolated shadow exact/semantic cache state for production evaluations
  • per-model cost, gross-savings, net-savings, and token-savings metrics
  • optional VCAL Security Guard request/response inspection
  • optional VCAL Privacy Guard anonymization, redaction, and restore processing
  • optional VCAL Usage Guard organizational usage-policy evaluation
  • structured vcal.evidence.event schema v1.1 events with stable trace correlation
  • optional buffered HTTP evidence delivery to VCAL Audit
  • optional downstream VCAL Compliance evidence processing
  • Prometheus metrics and Grafana dashboards
  • strict configuration validation and model allowlist behavior through model_price
  • request size limits and explicit timeout controls
  • liveness, startup, and readiness endpoints (/healthz, /startupz, /readyz)
  • graceful shutdown and validated hot reload via SIGHUP
  • configurable fail-open / fail-closed behavior for cache and guard dependency failures
  • semantic cache lifecycle controls
  • exact-cache identity that preserves OpenAI-compatible request/message extension fields
  • safe pass-through of non-string message content with semantic-cache bypass
  • generic non-root OCI image hardening
  • Docker Compose deployment examples
  • OpenShift restricted-v2 deployment assets under deploy/openshift/
  • /version endpoint for release and compatibility introspection

Deployment model​

AI Cost Firewall can run by itself as an OpenAI-compatible caching and cost-control gateway or as the request-processing core of the broader VCAL stack.

Application
-> AI Cost Firewall
-> Security Guard (optional)
-> Privacy Guard (optional)
-> Usage Guard (optional)
-> exact / semantic cache or upstream LLM
-> Audit evidence delivery (optional)
-> Compliance (optional)

Prometheus
-> Grafana for deep metrics and diagnostics

The control plane is self-hosted and customer-controlled. Model providers can change without requiring the application to integrate directly with every provider-specific control layer.

Deployment examples​

Ready-to-run deployment patterns are available under:

deploy/examples/

These include OpenAI cloud, local Ollama, hybrid OpenAI + local embeddings, OpenRouter, and a full local stack with dashboards. OpenShift-specific manifests are kept separately under deploy/openshift/ so the runtime image and Docker Compose path remain platform-neutral.

  1. What it does
  2. Quick Start with Docker
  3. OpenShift + vLLM (when deploying on OpenShift)
  4. Evaluation Mode
  5. Request flow
  6. Architecture overview
  7. Configuration overview
  8. Runtime overview
  9. Metrics
  10. Dashboards