Runtime Overview
AI Cost Firewall supports:
- liveness, startup, and readiness endpoints
- graceful shutdown
- request draining
- upstream and embedding timeout handling
- hot reload through
SIGHUP - runtime metrics
- AIF
enforce/observeruntime modes - isolated shadow exact/semantic cache state in Evaluation Mode
- semantic cache fail-open behavior
- OpenAI-compatible provider diagnostics
- embedding provider timeout visibility
- release and compatibility introspection through
/version - OpenAI-compatible chat-side model discovery through
/v1/models - optional Security Guard, Privacy Guard, and Usage Guard orchestration
- configurable guard fail-open/fail-closed behavior
Startup dependencies
Redis is required for exact caching.
In normal enforce operation, readiness dependency policy can be tuned independently of request-path fail-open behavior:
readiness_requires_redis true;
readiness_requires_qdrant false;
readiness_requires_upstream false;
Qdrant is required for normal enforced semantic-cache operation when:
semantic_cache_enabled true;
semantic_cache_fail_open applies to normal enforce runtime semantic lookup failures.
In observe mode, Redis, Qdrant, and embedding failures that affect only evaluation are non-blocking. AIF can keep serving the application, remains ready, records evaluation error telemetry, and resumes evaluation when the dependency recovers.
AI Cost Firewall validates runtime dependencies during startup and reload according to the active enforcement mode. v0.8.2 also exposes /startupz for orchestrators. If an enabled Redis or Qdrant cache is marked readiness-required but could not initialize in the current process, /startupz returns 503 so Kubernetes/OpenShift can restart the pod rather than leave it on a startup no-op cache.
This includes:
- loaded configuration summary
- Redis connectivity
- Qdrant connectivity
- semantic cache configuration completeness
- vector-size compatibility
- OpenAI-compatible upstream and embedding provider configuration
Use /healthz for liveness, /startupz for startup-probe gating, and /readyz for normal traffic readiness.
During graceful shutdown:
- readiness becomes unavailable
- new requests are rejected
- in-flight requests continue
AI Cost Firewall supports nginx-style configuration reload using:
SIGHUP
semantic_cache_fail_open affects normal enforce runtime semantic lookup behavior. Evaluation dependencies are treated differently in observe mode because they must not become a live-traffic dependency.
Version endpoint
The /version endpoint reports the running release, compatibility model, active AIF enforcement mode, effective cache scope, and effective semantic-cache collection.
Guard runtime behavior
AI Firewall can orchestrate VCAL Security Guard, VCAL Privacy Guard, and VCAL Usage Guard.
Recommended full enterprise order:
Security Guard request scan
→ Privacy Guard scan/anonymize/redact
→ Usage Guard request policy evaluation
→ exact/semantic cache lookup or upstream LLM
→ Security Guard response scan
→ Privacy Guard restore
guard_fail_open controls what happens when an enabled guard is unavailable, times out, or returns an invalid response contract.
Recommended enterprise setting:
guard_fail_open false;
This fails closed instead of sending unscanned or unanonymized traffic to cache, upstream providers, or clients.
Controlled streaming requests use the same guard ordering. On an upstream cache miss, AIF consumes provider SSE internally, assembles the complete response, applies response Security Guard and Privacy restoration, records accounting/evidence, and only then replays OpenAI-compatible SSE to the client.
upstream_timeout_seconds bounds the maximum idle gap between provider SSE chunks. A separate 15-minute absolute generation ceiling prevents indefinite upstream permit retention.
Evidence lifecycle
AI Firewall emits structured evidence events through application logs using schema vcal.evidence.event version 1.1.
For every trace that emits request.received, runtime processing emits exactly one terminal event:
request.completedfor successful delivery;request.failedfor validation, guard, cache-fatal, or upstream failures.
Security Guard and Usage Guard block decisions are retained as guard evidence. Security Guard rule_id values are preserved when available.