Caching Strategy
AI Cost Firewall evaluates cache reuse in stages:
- exact cache (Redis)
- semantic cache (Qdrant)
- upstream request
In enforce mode, an eligible cache hit can satisfy the request without an upstream chat call. In v0.8.0 observe mode, the same decision is evaluated against isolated shadow state, but the live upstream request still runs and supplies the client response.
Exact cache
Backend:
Redis / Valkey
Exact cache stores responses for identical normalized requests. v0.8.2 includes preserved OpenAI-compatible request and message extension fields in exact-cache identity, so requests that differ in tool definitions, tool-call history, structured-output settings, reasoning/provider extensions, or non-string content do not collapse onto the same exact key. Delivery-only streaming fields are removed before cache identity is calculated.
Benefits:
- very low latency
- no embedding lookup cost
- predictable matching
Semantic cache
Backend:
Qdrant
Semantic cache stores embeddings and response payloads for semantically similar prompts.
Example similar prompts:
"Explain Redis briefly"
"What is Redis used for?"
Semantic cache lookup requires an embedding request. When embedding_price is configured, this cost is included in net savings calculations.
Semantic cache may introduce embedding overhead.
Semantic-cache eligibility in v0.8.2
Semantic lookup and store are skipped when the request contains tools, contains response_format, or has any non-string message.content. Such requests can still use exact cache when exact caching is enabled and otherwise pass through to the upstream normally.
This prevents structured/multimodal-shaped payloads from entering the text embedding path and avoids unsafe semantic reuse across different non-text inputs.
AI Cost Firewall therefore distinguishes:
- gross savings
- embedding overhead
- net savings
Evaluation / shadow cache behavior
aif_enforcement_mode observe; keeps evaluation cache state separate from production state.
- exact-cache evaluation uses an isolated Redis key namespace;
- semantic-cache evaluation uses an isolated Qdrant collection;
- an exact shadow hit skips shadow semantic lookup/store because normal enforcement would already have returned;
- a semantic shadow hit can warm the shadow exact cache;
- an exact + semantic shadow miss stores the approved live upstream response in shadow cache when eligible;
- cache-bypass requests perform no shadow lookup and no shadow store.
Would-have hits do not increment normal production cache-hit or savings counters. See Evaluation Mode and Metrics.
Similarity threshold
semantic_similarity_threshold 0.92;
Typical values:
| Value | Behavior |
|---|---|
0.85 | aggressive reuse |
0.92 | balanced default |
0.97 | strict reuse |
Freshness
Semantic entries include inserted_at and expires_at. Expired entries are not reused.
Runnable deployment examples are available under:
deploy/examples/
Caching with Privacy Guard
When VCAL Privacy Guard is enabled in anonymize mode, AI Firewall calls Privacy Guard before cache lookup. This helps keep raw sensitive values out of Redis, Qdrant, semantic cache payloads, and upstream LLM calls.
Example:
Original:
Analyze login from 185.23.10.5 by john@example.com
Cache/upstream path:
Analyze login from [IP_1] by [EMAIL_1]
Final response after restore:
john@example.com logged in from 185.23.10.5
Cache bypass requests still pass through enabled guard orchestration. Bypass skips cache lookup and cache storage; it does not bypass Security Guard, Privacy Guard, or Usage Guard. In observe mode, bypass also skips shadow lookup/store so explicitly non-cacheable traffic is not counted as an optimization opportunity.