Skip to main content

AI Cost Firewall v0.6.0

AI Cost Firewall v0.6.0 extends the production control path with VCAL Usage Guard orchestration and adds clearer client-facing handling for Security Guard and Usage Guard policy blocks.

Highlights​

  • adds optional VCAL Usage Guard request-policy evaluation;
  • establishes the full guard order as Security Guard -> Privacy Guard -> Usage Guard -> cache/upstream;
  • adds configurable safe-completion or structured-error responses for Security Guard blocks;
  • adds configurable safe-completion or structured-error responses for Usage Guard block and escalate outcomes;
  • adds Usage Guard orchestration metrics, including aif_usage_blocks_total;
  • extends Grafana diagnostics with Usage Guard policy-block and guard-operational signals;
  • documents readiness dependency controls independently from request-path fail-open behavior;
  • documents optional /metrics authentication controls;
  • aligns Audit delivery directives with the current audit_timeout_seconds and audit_retry_max_backoff_ms configuration model.

Usage Guard orchestration​

Usage Guard runs after Privacy Guard on the request path:

Security Guard request scan
-> Privacy Guard anonymize/redact
-> Usage Guard policy evaluation
-> exact/semantic cache lookup or upstream LLM
-> Security Guard response scan
-> Privacy Guard restore

Example configuration:

usage_guard_enabled true;
usage_guard_url http://vcal-usage-guard:8095;
usage_guard_api_key your-usage-guard-key;
usage_guard_mode enforce;
usage_guard_tenant_id your-tenant-id;
usage_guard_policy_id business-use-only;
usage_guard_timeout_seconds 3;
usage_guard_block_response completion;

Deterministic Usage Guard rules do not require a local LLM. Optional semantic matching can use an embeddings path when enabled by the Usage Guard policy.

Client-facing guard block behavior​

Security Guard and Usage Guard policy decisions are separated from the response presented to the application.

security_guard_block_response completion;
usage_guard_block_response completion;

completion returns HTTP 200 with a safe OpenAI-compatible assistant message. error_json returns the structured HTTP 403 guard error.

Security Guard response-side blocks continue to suppress blocked assistant content and skip Privacy Guard restoration.

Observability​

AI Cost Firewall v0.6.0 adds Usage Guard to the common guard observability model:

aif_guard_requests_total
aif_guard_latency_seconds
aif_security_blocks_total
aif_privacy_restore_skipped_total
aif_usage_blocks_total

Usage Guard block metrics can include policy category and rule metadata when available.

Readiness and metrics access​

v0.6.0 configuration includes explicit readiness dependency policy:

readiness_requires_redis true;
readiness_requires_qdrant false;
readiness_requires_upstream false;

and optional metrics endpoint authentication:

metrics_auth_required false;
# metrics_auth_token replace-with-prometheus-token;

Readiness requirements are separate from request-path cache fail-open behavior.

Audit delivery configuration​

The current Audit delivery configuration uses:

audit_timeout_seconds 5;
audit_retry_initial_backoff_ms 250;
audit_retry_max_backoff_ms 5000;

Delivery remains asynchronous and memory-backed. Queue exhaustion or retry exhaustion can still result in evidence loss; disk-backed local replay is not provided by this release.

Compatibility notes​

  • The client and upstream API style remains OpenAI-compatible.
  • stream=true remains globally rejected with HTTP 422 before cache, guard, or upstream processing.
  • Security Guard, Privacy Guard, Usage Guard, Audit, and Compliance are optional commercial modules; the core gateway remains deployable without them.
  • Current AI Firewall guard modules inspect text content. Raw image, audio, video, and binary payloads are not directly scanned or classified.
  • The default Security Guard and Usage Guard client block-response mode in the example configuration is completion. Deployments that depend on HTTP 403 guard errors should set error_json explicitly.

Upgrade check​

After deployment:

curl -s http://localhost:8080/version

Confirm that the running service reports 0.6.0.