AI Cost Firewall v0.8.2 — Deployment Hardening
AI Cost Firewall v0.8.2 hardens the v0.8 Evaluation Mode baseline for containerized and orchestrated deployments. It does not replace the Observe -> Enforce workflow introduced in v0.8.0; it improves cache correctness, OpenAI-compatible gateway behavior, startup semantics, container security, and deployment readiness.
Exact-cache identity hardening
Exact-cache identity now preserves flattened OpenAI-compatible request and message extension fields. This prevents requests that differ in tool definitions, tool-call history, structured-output settings, reasoning/provider extensions, or other preserved fields from incorrectly sharing an exact-cache entry.
Transport-only streaming settings are still removed before cache identity is calculated so JSON and controlled-SSE delivery can reuse the same eligible completion.
Content arrays and semantic-cache safety
OpenAI-style non-string message.content values, including content arrays, remain accepted and are preserved for upstream forwarding and exact-cache identity.
Semantic cache is skipped when a request contains:
tools;response_format; or- any non-string
message.content.
This prevents structured/multimodal-shaped payloads from entering the text embedding path. Current Security, Privacy, and Usage Guard integrations inspect string message content only; nested text parts inside arrays are not yet independently processed.
Strict startup probe
AIF now exposes:
GET /startupz
When exact or semantic cache is enabled and the corresponding Redis/Qdrant backend is marked readiness-required, /startupz verifies that the backend actually initialized in the current process. This is intended for Kubernetes/OpenShift startup probes and prevents a pod that started before a cache dependency from remaining permanently on a fail-open no-op cache.
/healthz remains liveness-only and /readyz remains normal traffic readiness.
Model discovery
AIF now proxies:
GET /v1/models
to the configured chat/inference upstream. This improves compatibility with OpenAI-compatible clients such as Open WebUI and self-hosted backends such as vLLM. Model-list requests are not counted as chat inference calls for cache-savings accounting.
Container and OpenShift hardening
The AIF image remains a generic OCI image suitable for Docker, Docker Compose, Kubernetes, and OpenShift. v0.8.2 adds/retains:
- numeric non-root runtime identity;
- explicit SIGTERM handling;
- read-only-root-filesystem compatibility;
- no privilege-escalation requirement;
- no additional Linux capabilities requirement.
OpenShift-specific restricted-v2 manifests are provided separately under deploy/openshift/. They do not change the generic runtime or Docker Compose deployment model.
The reference OpenShift layout supports separate OpenAI-compatible vLLM services for chat/inference and embeddings. The embedding model ID and vector dimension must be verified in the target cluster before semantic cache is enabled.
Validation
The v0.8.2 source passed the project Cargo checks and test suite. The final release container was scanned with:
- Docker Scout: no detected vulnerabilities;
- Trivy: 0 Critical / 0 High findings.
Upgrade notes
Existing Docker/Compose deployments can continue to use the same deployment model. Review the new /startupz behavior if you use readiness-required Redis/Qdrant dependencies, and test requests using tools, response_format, or non-string content because semantic-cache eligibility is intentionally stricter in v0.8.2.