Architecture Overview
AI Cost Firewall is the request-processing core of a self-hosted AI control plane. It keeps the application-facing API OpenAI-compatible while centralizing caching, upstream routing, optional guard orchestration, evidence delivery, and operational telemetry.
Components
AI Cost Firewall
The Rust + Axum gateway that validates requests, applies configured guard orchestration, checks caches, runs normal enforced cache reuse or non-disruptive evaluation according to aif_enforcement_mode, exposes metrics, and emits structured evidence events.
Redis / Valkey
Stores exact cache entries.
Qdrant
Stores semantic cache entries and performs vector search. AI Cost Firewall uses Qdrant gRPC on port 6334.
OpenAI-compatible chat upstream
Receives requests that cannot be served from cache or that explicitly bypass cache.
OpenAI-compatible embedding provider
Semantic caching uses embeddings generated by the configured embedding provider. The embedding provider may be the same service as the chat upstream or a separate OpenAI-compatible endpoint.
VCAL Security Guard
Optional synchronous request/response inspection. Request-side blocks stop processing before Privacy Guard, cache lookup, or upstream forwarding.
VCAL Privacy Guard
Optional synchronous sensitive-data processing. Request text can be anonymized or redacted before Usage Guard, cache/upstream processing, and restored on the response path when mappings are available.
VCAL Usage Guard
Optional synchronous organizational usage-policy evaluation. Usage Guard runs after Privacy Guard on the request path and before cache lookup or upstream forwarding. It can return allow, warn, block, or escalate decisions.
VCAL Audit
Optional asynchronous evidence consumer. Audit persists evidence outside the gateway process, reconstructs traces, exports retained evidence, applies retention, and verifies its authoritative tamper-evident chain.
VCAL Compliance
Optional downstream evidence processor. Compliance synchronizes retained Audit events and uses them for control assessment, reporting, framework mappings, evidence packages, and related governance workflows.
Prometheus and Grafana
Prometheus scrapes runtime metrics from AI Cost Firewall and enabled VCAL services. Grafana remains the deep metrics and diagnostics interface for infrastructure and PromQL-oriented analysis.
OpenAI-compatible provider flexibility
AI Cost Firewall supports practical OpenAI-compatible providers including:
- OpenAI
- Ollama
- LM Studio
- vLLM
- LiteLLM
- OpenRouter
without requiring provider-specific configuration blocks. AIF v0.8.2 also proxies GET /v1/models to the configured chat upstream for client-side model discovery.
Runnable deployment patterns are available under:
deploy/examples/
Deployment hardening in v0.8.2
The AIF runtime image remains a generic OCI image. v0.8.2 hardens it for non-root, read-only container execution and adds OpenShift-specific restricted-v2 manifests under deploy/openshift/ without making the normal Docker/Compose path OpenShift-dependent.
For orchestrators, /startupz verifies that enabled Redis/Qdrant cache backends marked readiness-required actually initialized in the current process. /readyz remains the normal traffic-readiness endpoint and /healthz remains liveness-only.
A reference OpenShift topology can keep chat and embeddings on separate OpenAI-compatible vLLM services:
Open WebUI / client -> AIF -> vLLM chat
\-> vLLM embeddings -> Qdrant
Canonical response path, Evaluation Mode, and controlled streaming
AI Cost Firewall v0.8.2 preserves the controlled stream=true path introduced in v0.7.0 and the AIF-level observe mode introduced in v0.8.0. Provider SSE is consumed internally and assembled into the same canonical ChatCompletionResponse used by ordinary JSON upstream responses and eligible cache hits. Response controls run on that canonical response before the final transport is selected.
┌── exact / semantic cache hit ───────────┐
│ │
stream=false cache miss ──┼─► ordinary upstream JSON ───────────────┤
│ │
stream=true cache miss ───┴─► provider SSE ─► parser / assembler ──┤
│
▼
canonical ChatCompletionResponse
│
▼
response Security Guard
eligible cache store
(cache miss only, pre-restore)
Privacy Guard restore
usage / cost accounting
metrics / evidence
│
┌────────────┴────────────┐
▼ ▼
JSON serializer SSE encoder / replay
stream=false stream=true
→ [DONE]
In enforce mode, an eligible exact or semantic cache hit can enter the canonical response path without an upstream chat call. In observe mode, the same cache decision is evaluated against isolated shadow state, recorded as a would-have action, and the live upstream path still supplies the canonical response delivered to the client.
The commit barrier is deliberate: no generated response content leaves AI Cost Firewall until the complete response has passed the configured response controls. Malformed, truncated, oversized, stalled, or otherwise failed provider streams return a normal HTTP error before downstream SSE is committed. Stream intake is bounded by max_stream_upstream_bytes, upstream_timeout_seconds limits the idle gap between provider SSE chunks, and a separate 15-minute absolute generation ceiling prevents indefinite drip-feed streams.
Enterprise guard orchestration
AI Cost Firewall can orchestrate optional VCAL guard modules while keeping the core gateway usable as a standalone caching and cost-control layer.
Guard modules can be enabled independently or in combination:
AI Firewall only
AI Firewall + VCAL Security Guard
AI Firewall + VCAL Privacy Guard
AI Firewall + VCAL Usage Guard
AI Firewall + Security Guard + Privacy Guard + Usage Guard
Full guarded flow:
Client
-> AI Cost Firewall
-> VCAL Security Guard request scan
-> VCAL Privacy Guard anonymize/redact
-> VCAL Usage Guard policy evaluation
-> exact/semantic cache lookup or upstream LLM
-> VCAL Security Guard response scan
-> VCAL Privacy Guard restore
-> Client
Ordering matters:
- Security Guard evaluates the raw request before privacy mapping or usage-policy evaluation.
- Privacy Guard transforms sensitive text before Usage Guard, Redis, Qdrant, semantic cache payloads, or upstream providers.
- Usage Guard evaluates the post-privacy request before cache lookup or upstream forwarding.
- Security Guard scans assistant output before Privacy Guard restore.
- Privacy Guard restore is the final response transformation.
Evidence and audit boundary
AI Cost Firewall emits vcal.evidence.event schema v1.1 records with stable trace_id correlation. In Evaluation Mode, additive attributes such as enforcement_mode, decision, would_action, and applied_action keep hypothetical cache behavior separate from actions actually applied to live traffic. Each trace that emits request.received ends with exactly one request.completed or request.failed terminal event.
Operational metrics and retained evidence have different purposes:
- Prometheus answers aggregate operational questions over time.
- VCAL Audit preserves request-level execution evidence and trace history.
- VCAL Compliance consumes retained evidence for governance and reporting workflows.
Buffered Audit delivery
AI Cost Firewall can route structured evidence to a buffered HTTP sink:
request processing
-> evidence event
-> bounded in-memory queue
-> batch by size or flush interval
-> POST /v1/events/batch
-> VCAL Audit
The Audit path is asynchronous. Temporary Audit failures do not normally fail the client request. Retry exhaustion can result in evidence loss because the current sender queue is memory-backed.