Skip to main content

What AI Cost Firewall Does

AI Cost Firewall is the OpenAI-compatible request-processing layer of the VCAL control plane.

Instead of sending every request directly to an LLM provider, applications send chat-completion requests to AI Cost Firewall. The firewall validates the request, applies the enabled controls, checks whether an exact or semantically similar response can be reused, and forwards only the remaining requests upstream.

Responsibilities​

  • validate and normalize requests while preserving OpenAI-compatible extension fields
  • proxy chat-side model discovery through GET /v1/models
  • enforce configured model and request constraints
  • inspect requests and responses through optional VCAL Security Guard
  • anonymize or redact sensitive request text through optional VCAL Privacy Guard
  • evaluate organizational acceptable-use policy through optional VCAL Usage Guard
  • check exact cache
  • check semantic cache
  • serve eligible cache hits in enforce mode
  • evaluate would-have cache decisions using isolated shadow state in observe mode
  • forward misses upstream in enforce, and continue to the live upstream for evaluated requests in observe
  • store production or shadow cache entries according to the active mode
  • estimate actual production savings separately from prospective observe-mode cost and token avoidance
  • expose Prometheus metrics for operational and cost analysis
  • emit schema-versioned evidence events with stable trace correlation
  • optionally deliver evidence asynchronously to VCAL Audit through a bounded, batched HTTP sink
  • support downstream VCAL Compliance workflows through retained Audit evidence
  • handle liveness, startup/readiness dependency policy, graceful shutdown, and validated reload behavior
  • support controlled OpenAI-compatible streaming while preserving the full response-control path

OpenAI-compatible gateway​

Supported application-facing endpoints:

POST /v1/chat/completions
GET /v1/models

/v1/models is proxied to the configured chat/inference upstream. It is intended for OpenAI-compatible client model discovery and does not expose the separately configured embedding endpoint.

Existing applications can normally point their OpenAI-compatible client to AI Cost Firewall by changing the base URL rather than changing the request format.

Two cache layers​

LayerBackendPurpose
Exact cacheRedis / ValkeyReuse identical normalized requests
Semantic cacheQdrantReuse semantically similar prompts

Exact cache hits are fastest and do not require an embedding lookup. v0.8.2 includes preserved top-level and message-level OpenAI-compatible extension fields in exact-cache identity. Semantic cache hits require embedding work but can reuse answers for differently worded requests with sufficiently similar meaning. Requests containing tools, response_format, or any non-string message.content are deliberately excluded from semantic caching in v0.8.2.

Control-plane extensions​

AI Cost Firewall can orchestrate optional VCAL modules without requiring the application to call those services directly.

VCAL Security Guard​

Security Guard provides a deterministic, auditable first control layer for text-based risks such as prompt injection, jailbreak attempts, system-prompt extraction, unsafe tool-use instructions, data-exfiltration attempts, and common cyber-abuse patterns.

VCAL Privacy Guard​

Privacy Guard can detect, redact, anonymize, and restore supported sensitive text values before cache or upstream processing. Current detectors include payment-card patterns, email addresses, phone numbers, IPv4 addresses, US SSNs, IBANs, API keys and secrets, bearer tokens, JWTs, and private cryptographic keys.

VCAL Usage Guard​

Usage Guard evaluates request text against an organizational usage policy after Privacy Guard and before cache lookup or upstream forwarding. It can return explicit allow, warn, block, or escalate decisions. Deterministic rules do not require a local LLM; optional semantic matching can use an embeddings path when configured by the Usage Guard policy.

VCAL Audit​

Audit receives structured execution evidence asynchronously, persists it independently of the AI Firewall process, reconstructs traces, applies retention, supports export, and verifies its tamper-evident record chain.

VCAL Compliance​

Compliance works downstream of Audit. It synchronizes retained evidence and uses it for activity summaries, control assessment, framework mappings, reports, evidence packages, and trace-oriented compliance workflows.

Deployment patterns​

AI Cost Firewall supports practical OpenAI-compatible deployment patterns including:

  • OpenAI cloud deployments
  • fully local Ollama deployments
  • hybrid OpenAI + local embedding deployments
  • OpenRouter routing deployments
  • self-hosted vLLM deployments
  • OpenShift deployments using the generic AIF OCI image plus deployment-specific restricted-v2 manifests

Runnable examples are available under:

deploy/examples/

The AI Cost Firewall core can still be deployed without the optional VCAL modules when only gateway, caching, cost control, and Prometheus observability are required.