Skip to main content

AI Cost Firewall v0.8.2 — Deployment Hardening

AI Cost Firewall v0.8.2 hardens the v0.8 Evaluation Mode baseline for containerized and orchestrated deployments. It does not replace the Observe -> Enforce workflow introduced in v0.8.0; it improves cache correctness, OpenAI-compatible gateway behavior, startup semantics, container security, and deployment readiness.

Exact-cache identity hardening​

Exact-cache identity now preserves flattened OpenAI-compatible request and message extension fields. This prevents requests that differ in tool definitions, tool-call history, structured-output settings, reasoning/provider extensions, or other preserved fields from incorrectly sharing an exact-cache entry.

Transport-only streaming settings are still removed before cache identity is calculated so JSON and controlled-SSE delivery can reuse the same eligible completion.

Content arrays and semantic-cache safety​

OpenAI-style non-string message.content values, including content arrays, remain accepted and are preserved for upstream forwarding and exact-cache identity.

Semantic cache is skipped when a request contains:

  • tools;
  • response_format; or
  • any non-string message.content.

This prevents structured/multimodal-shaped payloads from entering the text embedding path. Current Security, Privacy, and Usage Guard integrations inspect string message content only; nested text parts inside arrays are not yet independently processed.

Strict startup probe​

AIF now exposes:

GET /startupz

When exact or semantic cache is enabled and the corresponding Redis/Qdrant backend is marked readiness-required, /startupz verifies that the backend actually initialized in the current process. This is intended for Kubernetes/OpenShift startup probes and prevents a pod that started before a cache dependency from remaining permanently on a fail-open no-op cache.

/healthz remains liveness-only and /readyz remains normal traffic readiness.

Model discovery​

AIF now proxies:

GET /v1/models

to the configured chat/inference upstream. This improves compatibility with OpenAI-compatible clients such as Open WebUI and self-hosted backends such as vLLM. Model-list requests are not counted as chat inference calls for cache-savings accounting.

Container and OpenShift hardening​

The AIF image remains a generic OCI image suitable for Docker, Docker Compose, Kubernetes, and OpenShift. v0.8.2 adds/retains:

  • numeric non-root runtime identity;
  • explicit SIGTERM handling;
  • read-only-root-filesystem compatibility;
  • no privilege-escalation requirement;
  • no additional Linux capabilities requirement.

OpenShift-specific restricted-v2 manifests are provided separately under deploy/openshift/. They do not change the generic runtime or Docker Compose deployment model.

The reference OpenShift layout supports separate OpenAI-compatible vLLM services for chat/inference and embeddings. The embedding model ID and vector dimension must be verified in the target cluster before semantic cache is enabled.

Validation​

The v0.8.2 source passed the project Cargo checks and test suite. The final release container was scanned with:

  • Docker Scout: no detected vulnerabilities;
  • Trivy: 0 Critical / 0 High findings.

Upgrade notes​

Existing Docker/Compose deployments can continue to use the same deployment model. Review the new /startupz behavior if you use readiness-required Redis/Qdrant dependencies, and test requests using tools, response_format, or non-string content because semantic-cache eligibility is intentionally stricter in v0.8.2.