Skip to main content

AI Cost Firewall v0.8.0 — Evaluation Mode

AI Cost Firewall v0.8.0 adds a non-disruptive Evaluation / Observe Mode for production pilots. AIF can now remain in the live request/response path, evaluate what exact and semantic caching would have done, and collect prospective optimization evidence without serving the shadow cached response to the application.

The normal production behavior remains the default through:

aif_enforcement_mode enforce;

For a pilot assessment:

aif_enforcement_mode observe;

Highlights​

  • adds AIF enforce and observe runtime modes;
  • uses isolated Redis shadow state for exact-cache evaluation;
  • uses an isolated Qdrant collection for semantic-cache evaluation;
  • records would-have exact/semantic decisions while still calling the live upstream provider;
  • prevents shadow evaluation state from contaminating production cache state;
  • preserves cache-bypass semantics by skipping shadow lookup/store for bypassed requests;
  • adds dedicated evaluation metrics for would-have hits, potentially avoidable upstream calls/tokens, estimated cost avoidance, shadow stores, and evaluation errors;
  • adds enforcement_mode, decision, would_action, and applied_action evidence attributes;
  • keeps Redis, Qdrant, and embedding evaluation failures non-blocking for live traffic;
  • keeps AIF ready when optional evaluation dependencies are unavailable;
  • recovers Redis exact-cache and Qdrant semantic evaluation after dependency restart;
  • preserves transport-independent JSON/SSE cache identity from v0.7.0 controlled streaming.

Evaluation semantics​

A shadow cache hit does not replace the upstream response:

request
→ shadow cache evaluation
→ would-have decision
→ live upstream call
→ response controls
→ shadow state update when eligible
→ client receives live upstream response

Would-have hits do not increment normal production hit or savings counters.

Evidence​

Example shadow exact hit:

{
"enforcement_mode": "observe",
"decision": "exact_hit",
"would_action": "serve_exact_cache",
"applied_action": "call_upstream"
}

The evidence schema remains vcal.evidence.event v1.1 because the new fields are additive attributes.

Reliability and recovery​

Evaluation Mode is designed so optional evaluation infrastructure cannot interrupt the application being assessed.

Validated behavior includes:

  • Redis outage while observing: live request continues;
  • Qdrant outage while observing: live request continues;
  • readiness remains healthy during evaluation dependency loss;
  • evaluation error telemetry records the failure;
  • Redis exact-cache evaluation resumes after Redis returns;
  • Qdrant semantic evaluation resumes after Qdrant returns.

Controlled streaming​

Evaluation Mode works with the controlled streaming architecture introduced in v0.7.0. Stream and JSON delivery remain transport-independent for cache identity, but a would-have hit in observe still results in a live upstream request.

VCAL module scope​

This release adds Evaluation Mode to AI Cost Firewall caching and cost optimization only.

It does not add independent off / observe / enforce modes to VCAL Security Guard, Privacy Guard, or Usage Guard. Existing guard configuration and enforcement behavior remain unchanged.

No VCAL module upgrade is required solely for AIF v0.8.0 Evaluation Mode.

Upgrade notes​

Existing deployments that do not set aif_enforcement_mode retain normal production behavior because enforce is the default.

Evaluation state is intentionally isolated from production state and is not promoted automatically when switching from observe to enforce.