Request Flow
AI Cost Firewall validates each request, applies request-side controls, checks cache, and uses either ordinary upstream JSON (stream=false or omitted) or controlled provider SSE (stream=true). In enforce mode, eligible cache hits can satisfy the request directly. In v0.8.0 observe mode, cache decisions are evaluated against isolated shadow state while the live upstream call still runs. All delivered responses converge on the same canonical ChatCompletionResponse before response controls and final delivery.
receive request
→ enforce request body and prompt-size limits
→ normalize request
→ Security Guard request scan
→ Privacy Guard scan/anonymize/redact
→ Usage Guard policy evaluation
→ check per-request cache bypass
→ exact cache lookup, if enabled
→ semantic cache lookup, if enabled, request is semantically eligible, and exact did not decide the shadow/enforced path
→ enforce mode: cache hit OR upstream request on miss/bypass
→ observe mode: record would-have cache decision, then call live upstream
→ stream=false/omitted: receive ordinary JSON
→ stream=true: consume provider SSE and assemble complete response
→ canonical ChatCompletionResponse
→ Security Guard response scan
→ eligible production or shadow cache store according to the active mode (pre-Privacy-restore)
→ Privacy Guard restore, if mapping exists
→ actual/evaluation accounting, metrics, and evidence
→ JSON serializer for stream=false OR approved SSE replay for stream=true
→ emit request.completed
Security Guard and Usage Guard enforcement decisions stop the request before cache lookup or upstream processing. Client-facing block behavior remains configurable:
completionreturns HTTP 200 with a safe OpenAI-compatible assistant message;error_jsonreturns the structured HTTP 403 guard error.
For response-side Security Guard blocks, blocked assistant content is not returned and Privacy Guard restore is skipped. Guard failures remain subject to guard_fail_open. Failure paths emit request.failed; successful paths emit request.completed.
Canonical response convergence
ordinary upstream JSON ───────┐
├─► canonical ChatCompletionResponse
assembled upstream stream ────┤
│
eligible cache hit ───────────┘
│
▼
response controls
eligible cache store
(cache miss only, pre-restore)
Privacy restore
accounting / evidence
│
┌────────────┴────────────┐
▼ ▼
JSON serializer SSE encoder
stream=false stream=true
For controlled streaming, AIF does not expose provider chunks directly. It consumes and validates the complete provider stream first, so response Security Guard scanning and Privacy Guard restoration are not skipped.
Exact cache
The firewall checks Redis / Valkey for an identical normalized request. v0.8.2 includes preserved top-level and message-level OpenAI-compatible extension fields in normalized exact-cache identity, while transport-only streaming fields are removed before the cache key is computed. In observe mode, the exact lookup uses the isolated evaluation namespace and a hit is recorded as a would-have decision rather than being served to the client.
Semantic cache
If exact cache misses, semantic cache can search Qdrant for similar prompts only when the request is semantically eligible. v0.8.2 skips semantic lookup/store for requests containing tools, response_format, or any non-string message.content. The request can still use exact cache and can still be forwarded upstream.
Semantic cache entries include:
- inserted_at
- expires_at
Expired entries are skipped during lookup and never reused.
A candidate is reusable only if:
similarity_score >= semantic_similarity_threshold
AND
expires_at > now
AND
cached response payload is valid
Expired entries are filtered before similarity ranking.
When:
semantic_cache_fail_open true;
runtime semantic lookup failures behave like cache misses and requests continue upstream normally in enforce mode.
In observe mode, semantic evaluation failures are always non-blocking for live application traffic. They are recorded with evaluation error telemetry and the live upstream request continues.
Upstream request
In enforce mode, the request is forwarded upstream when no valid cache hit exists or when cache is bypassed. In observe mode, the live upstream request still runs even when shadow cache evaluation reports a would-have hit.
The chat provider and embedding provider may use separate OpenAI-compatible endpoints.
Cache storage
Eligible cache-miss responses can be stored in Redis and Qdrant after response Security Guard scanning and before Privacy Guard restoration. In enforce mode this updates production cache state; in observe mode only isolated shadow state is updated. The same pre-restore cache semantics apply whether the provider returned ordinary JSON or AIF assembled a controlled provider stream.
Evidence lifecycle
AI Firewall emits structured vcal.evidence.event schema v1.1 records through application logs.
Every trace that emits request.received ends with exactly one terminal event:
request.completed
or:
request.failed
The same trace_id correlates request validation, Security Guard, Privacy Guard, and Usage Guard evidence, cache activity, upstream activity, and the terminal outcome. Security Guard rule_id values are preserved when available.
Buffered Audit delivery
AI Firewall can route structured evidence to a buffered HTTP sink:
request processing
-> evidence event
-> bounded in-memory queue
-> batch by size or flush interval
-> POST /v1/events/batch
-> VCAL Audit
The Audit path is asynchronous. Temporary Audit failures do not normally fail the client request. Retry exhaustion can result in evidence loss because the current sender queue is memory-backed.