Caching Strategy
AI Cost Firewall evaluates cache reuse in stages:
- exact cache (Redis)
- semantic cache (Qdrant)
- upstream request
Exact cache
Backend:
Redis / Valkey
Exact cache stores responses for identical normalized requests.
Benefits:
- very low latency
- no embedding lookup cost
- predictable matching
Semantic cache
Backend:
Qdrant
Semantic cache stores embeddings and response payloads for semantically similar prompts.
Example similar prompts:
"Explain Redis briefly"
"What is Redis used for?"
Semantic cache lookup requires an embedding request. When embedding_price is configured, this cost is included in net savings calculations.
Semantic cache may introduce embedding overhead.
AI Cost Firewall therefore distinguishes:
- gross savings
- embedding overhead
- net savings
Similarity threshold
semantic_similarity_threshold 0.92;
Typical values:
| Value | Behavior |
|---|---|
0.85 | aggressive reuse |
0.92 | balanced default |
0.97 | strict reuse |
Freshness
Semantic entries include inserted_at and expires_at. Expired entries are not reused.
Runnable deployment examples are available under:
deploy/examples/
Caching with Privacy Guard
When VCAL Privacy Guard is enabled in anonymize mode, AI Firewall calls Privacy Guard before cache lookup. This helps keep raw sensitive values out of Redis, Qdrant, semantic cache payloads, and upstream LLM calls.
Example:
Original:
Analyze login from 185.23.10.5 by john@example.com
Cache/upstream path:
Analyze login from [IP_1] by [EMAIL_1]
Final response after restore:
john@example.com logged in from 185.23.10.5
Cache bypass requests still pass through enabled guard orchestration. Bypass only skips cache lookup and cache storage; it does not bypass Security Guard or Privacy Guard.