Evaluation Mode
AI Cost Firewall v0.8.0 adds an AIF-level Evaluation Mode for low-risk production pilots.
The runtime mode is selected with:
aif_enforcement_mode enforce;
or:
aif_enforcement_mode observe;
enforce is the default and preserves normal production behavior.
In observe mode, AI Cost Firewall remains in the live request/response path and evaluates what exact and semantic caching would have done, but a cache hit does not replace the live upstream response.
Core guarantee
Evaluation Mode measures what AIF would have done without serving the shadow cached response to the application.
For an eligible request in observe mode:
request
→ request-side controls
→ cache bypass check
→ shadow exact lookup
→ shadow semantic lookup if exact misses
→ record would-have decision
→ live upstream call
→ response controls
→ shadow store when enforcement would have stored
→ accounting / evidence
→ return live upstream response
Isolated shadow state
Evaluation state is isolated from production cache state.
- exact-cache evaluation uses a separate Redis key namespace;
- semantic-cache evaluation uses a separate Qdrant collection;
- switching from
observetoenforcedoes not automatically promote shadow entries into production cache state.
This prevents a pilot from warming or mutating the cache that will later serve production responses.
Shadow decision semantics
| Shadow result | Live behavior | Shadow behavior |
|---|---|---|
| exact hit | call upstream | do not perform semantic lookup/store because enforcement would already have returned |
| exact miss + semantic hit | call upstream | warm shadow exact cache from the semantic result |
| exact + semantic miss | call upstream | store the approved live upstream response in shadow exact and semantic cache when eligible |
| cache bypass | call upstream | perform no shadow lookup or shadow store |
The bypass behavior is intentional: traffic explicitly marked non-cacheable should not inflate an evaluation's prospective savings.
Failure behavior
Evaluation infrastructure must not interrupt the application being assessed.
In observe mode:
- Redis lookup/store failures are recorded as evaluation errors and the request continues;
- Qdrant lookup/store failures are recorded as evaluation errors and the request continues;
- embedding failures used only for semantic evaluation are recorded and the request continues;
- optional evaluation dependencies do not make AIF unready;
- Redis and Qdrant evaluation paths recover after the dependency returns.
Upstream failures remain visible to the client because the live upstream response is still the response source.
Metrics and accounting
Would-have cache hits do not increment normal production hit/savings counters.
Evaluation-specific metrics report:
- requests evaluated;
- would-have exact and semantic outcomes;
- potentially avoidable upstream calls;
- potentially avoidable tokens;
- estimated gross and net cost avoidance;
- shadow store outcomes;
- evaluation errors.
For a shadow hit, prospective token/cost avoidance is based on the actual live upstream response observed for the current request. This keeps the estimate tied to the call that enforcement would have avoided.
See Metrics.
Evidence
Evaluation evidence separates the hypothetical cache action from what AIF actually applied to live traffic.
Example exact shadow hit:
{
"enforcement_mode": "observe",
"decision": "exact_hit",
"would_action": "serve_exact_cache",
"applied_action": "call_upstream"
}
The evidence schema remains vcal.evidence.event v1.1; these values are additive event attributes.
See Evidence Events.
Controlled streaming
Evaluation Mode preserves the transport-independent cache identity introduced with controlled streaming.
A request seeded with stream=true can still produce a would-have exact hit for a later JSON request, and a JSON seed can produce a would-have exact hit for a later controlled-stream request. In observe mode, the live upstream is still called in both cases.
VCAL Guard scope
aif_enforcement_mode controls AIF caching/cost optimization only.
It does not change the configured behavior of VCAL Security Guard, Privacy Guard, or Usage Guard. Independent guard off / observe / enforce modes are not part of AIF v0.8.0.