What AI Cost Firewall Does
AI Cost Firewall is the OpenAI-compatible request-processing layer of the VCAL control plane.
Instead of sending every request directly to an LLM provider, applications send chat-completion requests to AI Cost Firewall. The firewall validates the request, applies the enabled controls, checks whether an exact or semantically similar response can be reused, and forwards only the remaining requests upstream.
Responsibilities
- validate and normalize requests while preserving OpenAI-compatible extension fields
- proxy chat-side model discovery through
GET /v1/models - enforce configured model and request constraints
- inspect requests and responses through optional VCAL Security Guard
- anonymize or redact sensitive request text through optional VCAL Privacy Guard
- evaluate organizational acceptable-use policy through optional VCAL Usage Guard
- check exact cache
- check semantic cache
- serve eligible cache hits in
enforcemode - evaluate would-have cache decisions using isolated shadow state in
observemode - forward misses upstream in
enforce, and continue to the live upstream for evaluated requests inobserve - store production or shadow cache entries according to the active mode
- estimate actual production savings separately from prospective observe-mode cost and token avoidance
- expose Prometheus metrics for operational and cost analysis
- emit schema-versioned evidence events with stable trace correlation
- optionally deliver evidence asynchronously to VCAL Audit through a bounded, batched HTTP sink
- support downstream VCAL Compliance workflows through retained Audit evidence
- handle liveness, startup/readiness dependency policy, graceful shutdown, and validated reload behavior
- support controlled OpenAI-compatible streaming while preserving the full response-control path
OpenAI-compatible gateway
Supported application-facing endpoints:
POST /v1/chat/completions
GET /v1/models
/v1/models is proxied to the configured chat/inference upstream. It is intended for OpenAI-compatible client model discovery and does not expose the separately configured embedding endpoint.
Existing applications can normally point their OpenAI-compatible client to AI Cost Firewall by changing the base URL rather than changing the request format.
Two cache layers
| Layer | Backend | Purpose |
|---|---|---|
| Exact cache | Redis / Valkey | Reuse identical normalized requests |
| Semantic cache | Qdrant | Reuse semantically similar prompts |
Exact cache hits are fastest and do not require an embedding lookup. v0.8.2 includes preserved top-level and message-level OpenAI-compatible extension fields in exact-cache identity. Semantic cache hits require embedding work but can reuse answers for differently worded requests with sufficiently similar meaning. Requests containing tools, response_format, or any non-string message.content are deliberately excluded from semantic caching in v0.8.2.
Control-plane extensions
AI Cost Firewall can orchestrate optional VCAL modules without requiring the application to call those services directly.
VCAL Security Guard
Security Guard provides a deterministic, auditable first control layer for text-based risks such as prompt injection, jailbreak attempts, system-prompt extraction, unsafe tool-use instructions, data-exfiltration attempts, and common cyber-abuse patterns.
VCAL Privacy Guard
Privacy Guard can detect, redact, anonymize, and restore supported sensitive text values before cache or upstream processing. Current detectors include payment-card patterns, email addresses, phone numbers, IPv4 addresses, US SSNs, IBANs, API keys and secrets, bearer tokens, JWTs, and private cryptographic keys.
VCAL Usage Guard
Usage Guard evaluates request text against an organizational usage policy after Privacy Guard and before cache lookup or upstream forwarding. It can return explicit allow, warn, block, or escalate decisions. Deterministic rules do not require a local LLM; optional semantic matching can use an embeddings path when configured by the Usage Guard policy.
VCAL Audit
Audit receives structured execution evidence asynchronously, persists it independently of the AI Firewall process, reconstructs traces, applies retention, supports export, and verifies its tamper-evident record chain.
VCAL Compliance
Compliance works downstream of Audit. It synchronizes retained evidence and uses it for activity summaries, control assessment, framework mappings, reports, evidence packages, and trace-oriented compliance workflows.
Deployment patterns
AI Cost Firewall supports practical OpenAI-compatible deployment patterns including:
- OpenAI cloud deployments
- fully local Ollama deployments
- hybrid OpenAI + local embedding deployments
- OpenRouter routing deployments
- self-hosted vLLM deployments
- OpenShift deployments using the generic AIF OCI image plus deployment-specific
restricted-v2manifests
Runnable examples are available under:
deploy/examples/
The AI Cost Firewall core can still be deployed without the optional VCAL modules when only gateway, caching, cost control, and Prometheus observability are required.