Skip to main content

Evaluation Mode

AI Cost Firewall v0.8.0 adds an AIF-level Evaluation Mode for low-risk production pilots.

The runtime mode is selected with:

aif_enforcement_mode enforce;

or:

aif_enforcement_mode observe;

enforce is the default and preserves normal production behavior.

In observe mode, AI Cost Firewall remains in the live request/response path and evaluates what exact and semantic caching would have done, but a cache hit does not replace the live upstream response.

Core guarantee​

Evaluation Mode measures what AIF would have done without serving the shadow cached response to the application.

For an eligible request in observe mode:

request
→ request-side controls
→ cache bypass check
→ shadow exact lookup
→ shadow semantic lookup if exact misses
→ record would-have decision
→ live upstream call
→ response controls
→ shadow store when enforcement would have stored
→ accounting / evidence
→ return live upstream response

Isolated shadow state​

Evaluation state is isolated from production cache state.

  • exact-cache evaluation uses a separate Redis key namespace;
  • semantic-cache evaluation uses a separate Qdrant collection;
  • switching from observe to enforce does not automatically promote shadow entries into production cache state.

This prevents a pilot from warming or mutating the cache that will later serve production responses.

Shadow decision semantics​

Shadow resultLive behaviorShadow behavior
exact hitcall upstreamdo not perform semantic lookup/store because enforcement would already have returned
exact miss + semantic hitcall upstreamwarm shadow exact cache from the semantic result
exact + semantic misscall upstreamstore the approved live upstream response in shadow exact and semantic cache when eligible
cache bypasscall upstreamperform no shadow lookup or shadow store

The bypass behavior is intentional: traffic explicitly marked non-cacheable should not inflate an evaluation's prospective savings.

Failure behavior​

Evaluation infrastructure must not interrupt the application being assessed.

In observe mode:

  • Redis lookup/store failures are recorded as evaluation errors and the request continues;
  • Qdrant lookup/store failures are recorded as evaluation errors and the request continues;
  • embedding failures used only for semantic evaluation are recorded and the request continues;
  • optional evaluation dependencies do not make AIF unready;
  • Redis and Qdrant evaluation paths recover after the dependency returns.

Upstream failures remain visible to the client because the live upstream response is still the response source.

Metrics and accounting​

Would-have cache hits do not increment normal production hit/savings counters.

Evaluation-specific metrics report:

  • requests evaluated;
  • would-have exact and semantic outcomes;
  • potentially avoidable upstream calls;
  • potentially avoidable tokens;
  • estimated gross and net cost avoidance;
  • shadow store outcomes;
  • evaluation errors.

For a shadow hit, prospective token/cost avoidance is based on the actual live upstream response observed for the current request. This keeps the estimate tied to the call that enforcement would have avoided.

See Metrics.

Evidence​

Evaluation evidence separates the hypothetical cache action from what AIF actually applied to live traffic.

Example exact shadow hit:

{
"enforcement_mode": "observe",
"decision": "exact_hit",
"would_action": "serve_exact_cache",
"applied_action": "call_upstream"
}

The evidence schema remains vcal.evidence.event v1.1; these values are additive event attributes.

See Evidence Events.

Controlled streaming​

Evaluation Mode preserves the transport-independent cache identity introduced with controlled streaming.

A request seeded with stream=true can still produce a would-have exact hit for a later JSON request, and a JSON seed can produce a would-have exact hit for a later controlled-stream request. In observe mode, the live upstream is still called in both cases.

VCAL Guard scope​

aif_enforcement_mode controls AIF caching/cost optimization only.

It does not change the configured behavior of VCAL Security Guard, Privacy Guard, or Usage Guard. Independent guard off / observe / enforce modes are not part of AIF v0.8.0.