Configuration Overview
AI Cost Firewall v0.8.2 retains aif_enforcement_mode enforce|observe from v0.8.0. enforce remains the default. observe evaluates exact/semantic cache opportunities using isolated shadow state while the live upstream response remains authoritative.
Controlled stream=true requests remain fully supported. Provider SSE is assembled internally, response controls run on the complete canonical response, and approved content is replayed as OpenAI-compatible SSE.
v0.8.2 additionally provides /startupz, /v1/models, safer exact-cache identity for OpenAI-compatible extension fields, and semantic-cache bypass for non-string message content.
Evidence events are emitted through application logs rather than config directives. Enable them with:
RUST_LOG=info,vcal_evidence=info
AI Cost Firewall uses nginx-style configuration.
directive value;
Example:
listen_addr 0.0.0.0:8080;
Directives are case-sensitive and must end with a semicolon.
Configuration model
AI Cost Firewall keeps configuration intentionally flat. OpenAI-compatible cloud services, local model servers, and proxy gateways are configured with the same directive style instead of provider-specific configuration blocks.
Use:
upstream_*directives for chat-completion trafficembedding_*directives for semantic-cache embeddings- cache directives for exact-cache and semantic-cache behavior
- optional guard directives for VCAL Guard integrations
upstream_base_url and embedding_base_url may use either a provider root URL or its /v1 base path. Do not configure full endpoint paths such as /v1/chat/completions, /v1/embeddings, or /v1/models.
Example configuration
The example below reflects the v0.8.2 configuration model while keeping optional integrations disabled by default.
config_version 1;
listen_addr 0.0.0.0:8080;
aif_enforcement_mode enforce;
redis_url redis://redis:6379;
redis_timeout_seconds 2;
upstream_provider openai_compatible;
upstream_base_url https://api.openai.com;
upstream_api_key replace-with-upstream-api-key;
streaming_enabled true;
max_stream_upstream_bytes 8M;
embedding_provider openai_compatible;
embedding_base_url https://api.openai.com;
embedding_api_key replace-with-embedding-api-key;
embedding_model text-embedding-3-small;
qdrant_url http://qdrant:6334;
qdrant_collection aif_semantic_cache;
qdrant_vector_size 1536;
cache_ttl_seconds 2592000;
exact_cache_enabled true;
exact_cache_fail_open true;
exact_cache_store_enabled true;
semantic_cache_enabled true;
semantic_similarity_threshold 0.92;
semantic_cache_fail_open true;
semantic_cache_store_enabled true;
request_timeout_seconds 120;
upstream_timeout_seconds 120;
embedding_timeout_seconds 30;
max_request_body_bytes 1048576;
max_prompt_chars 65536;
cache_bypass_header X-AIF-Cache-Bypass;
metrics_auth_required false;
readiness_requires_redis true;
readiness_requires_qdrant false;
readiness_requires_upstream false;
allow_unknown_models_pass_through false;
model_price gpt-4o-mini-2024-07-18 0.15 0.60;
model_price gpt-4.1-mini-2025-04-14 0.30 1.20;
embedding_price 0.020;
security_guard_enabled false;
security_guard_url http://vcal-security-guard:8091;
security_guard_timeout_seconds 3;
security_guard_block_response completion;
privacy_guard_enabled false;
privacy_guard_url http://vcal-privacy-guard:8090;
privacy_guard_mode detect_only;
privacy_guard_restore_enabled true;
privacy_guard_timeout_seconds 3;
usage_guard_enabled false;
usage_guard_url http://vcal-usage-guard:8095;
usage_guard_mode detect_only;
usage_guard_timeout_seconds 3;
usage_guard_block_response completion;
guard_fail_open false;
audit_enabled false;
audit_url http://vcal-audit:8092;
audit_producer_instance_id ai-firewall-01;
audit_queue_capacity 10000;
audit_batch_size 100;
audit_flush_interval_ms 1000;
audit_timeout_seconds 5;
audit_retry_max_attempts 5;
audit_retry_initial_backoff_ms 250;
audit_retry_max_backoff_ms 5000;
Secrets such as upstream, embedding, guard, Audit, Qdrant, and metrics-auth tokens should be supplied through the deployment's secrets-management mechanism rather than committed to source control.
Core directives
| Directive | Purpose |
|---|---|
listen_addr | Address and port where AI Cost Firewall listens. |
aif_enforcement_mode | AIF cache/optimization mode: enforce (default) or non-disruptive observe. |
redis_url | Redis connection URL for the exact cache. |
upstream_provider | Chat upstream provider type. Currently uses OpenAI-compatible behavior. |
upstream_base_url | Base URL for the chat-completion provider. |
upstream_api_key | API key for the chat-completion provider. |
streaming_enabled | Allows clients to request controlled OpenAI-compatible streaming with stream=true. |
max_stream_upstream_bytes | Maximum cumulative provider SSE bytes accepted for one controlled streaming request. |
embedding_provider | Embedding provider type. Currently uses OpenAI-compatible behavior. |
embedding_base_url | Base URL for the embedding provider. |
embedding_api_key | API key for the embedding provider. |
embedding_model | Embedding model used for semantic-cache vectors. |
qdrant_url | Qdrant URL for the semantic cache. |
qdrant_api_key | Optional Qdrant API key. |
qdrant_collection | Qdrant collection used by AI Cost Firewall. |
qdrant_vector_size | Vector size. Must match the configured embedding model. |
Cache controls
| Directive | Purpose |
|---|---|
exact_cache_enabled | Enables or disables exact-cache lookup. |
exact_cache_store_enabled | Controls whether new exact-cache entries are stored. |
exact_cache_fail_open | Controls whether Redis failures are skipped or treated as request failures. |
semantic_cache_enabled | Enables or disables semantic-cache lookup. |
semantic_cache_store_enabled | Controls whether new semantic-cache entries are stored. |
semantic_similarity_threshold | Minimum similarity score required for a semantic-cache hit. |
semantic_cache_fail_open | Controls whether embedding/Qdrant failures are skipped or treated as request failures. |
cache_ttl_seconds | Backward-compatible TTL default for cache layers. |
exact_cache_ttl_seconds | Optional explicit TTL for exact-cache entries. |
semantic_cache_retention_seconds | Optional retention period for semantic-cache entries. |
cache_bypass_header | Optional request header used to bypass cache lookup and cache store for one request. |
Metrics and readiness controls
| Directive | Purpose |
|---|---|
metrics_auth_required | Requires authentication for /metrics when enabled. |
metrics_auth_token | Token used to authorize access to /metrics when metrics auth is enabled. |
readiness_requires_redis | Makes Redis availability part of /readyz. |
readiness_requires_qdrant | Makes Qdrant availability part of /readyz. |
readiness_requires_upstream | Makes upstream availability part of /readyz. |
Readiness requirements are separate from request-path fail-open behavior in normal enforce operation. In observe mode, Redis/Qdrant/embedding dependencies used only for evaluation are non-blocking for /readyz; evaluation failures are reported separately. v0.8.2 /startupz is intentionally stricter for enabled Redis/Qdrant caches marked readiness-required and returns 503 if those clients did not initialize in the current process.
Request limits and timeouts
| Directive | Purpose |
|---|---|
request_timeout_seconds | Backward-compatible request timeout. |
upstream_timeout_seconds | Timeout for upstream chat-completion response headers and, for controlled streams, the maximum idle gap between provider SSE chunks. |
embedding_timeout_seconds | Timeout for embedding provider requests. |
max_request_body_bytes | Maximum accepted HTTP request body size. |
max_prompt_chars | Maximum accepted combined prompt size. |
Controlled generation also has an independent 15-minute absolute ceiling so a provider cannot retain an upstream concurrency permit indefinitely by continuously drip-feeding SSE data.
Model and pricing controls
| Directive | Purpose |
|---|---|
allow_unknown_models_pass_through | Allows or rejects models not defined with model_price. |
model_price | Defines chat-completion model pricing for savings and cost metrics. |
embedding_price | Optional embedding price used for net cost estimation. |
Optional VCAL Guard directives
Security Guard
| Directive | Purpose |
|---|---|
security_guard_enabled | Enables or disables Security Guard integration. |
security_guard_url | Base URL of VCAL Security Guard. |
security_guard_api_key | Service-to-service API key. |
security_guard_timeout_seconds | Timeout for Security Guard calls. |
security_guard_block_response | Client response mode for Security Guard blocks: completion or error_json. |
security_guard_block_message | Safe assistant message returned when block response mode is completion. |
Privacy Guard
| Directive | Purpose |
|---|---|
privacy_guard_enabled | Enables or disables Privacy Guard integration. |
privacy_guard_url | Base URL of VCAL Privacy Guard. |
privacy_guard_api_key | Service-to-service API key. |
privacy_guard_mode | detect_only, redact, or anonymize. |
privacy_guard_restore_enabled | Restores mapped placeholders in assistant responses. |
privacy_guard_timeout_seconds | Timeout for scan and restore calls. |
Usage Guard
| Directive | Purpose |
|---|---|
usage_guard_enabled | Enables or disables Usage Guard integration. |
usage_guard_url | Base URL of VCAL Usage Guard. |
usage_guard_api_key | Service-to-service API key. |
usage_guard_mode | Usage Guard scan/enforcement mode, for example detect_only or enforce. |
usage_guard_tenant_id | Optional tenant identifier supplied to Usage Guard. |
usage_guard_policy_id | Optional policy identifier supplied to Usage Guard. |
usage_guard_timeout_seconds | Timeout for Usage Guard evaluation. |
usage_guard_block_response | Client response mode for Usage Guard block or escalate: completion or error_json. |
usage_guard_block_message | Safe assistant message returned when block response mode is completion. |
Shared guard failure policy
| Directive | Purpose |
|---|---|
guard_fail_open | true skips a failed guard; false fails closed. |
In v0.8.0 the recommended full request-side order remains Security Guard -> Privacy Guard -> Usage Guard -> cache/upstream. Usage Guard therefore evaluates anonymized text when Privacy Guard anonymization is enabled. The AIF observe mode changes cache enforcement only; it does not put the guard modules into observe mode.
Environment variables
The AIF mode can be overridden with AIF_ENFORCEMENT_MODE=enforce|observe.
Most deployment examples use configuration files. Containerized deployments may also provide selected settings through environment variables.
Common environment variables include:
AIF_STREAMING_ENABLED
AIF_MAX_STREAM_UPSTREAM_BYTES
AIF_PRIVACY_GUARD_ENABLED
AIF_PRIVACY_GUARD_URL
AIF_PRIVACY_GUARD_API_KEY
AIF_PRIVACY_GUARD_MODE
AIF_PRIVACY_GUARD_RESTORE_ENABLED
AIF_SECURITY_GUARD_ENABLED
AIF_SECURITY_GUARD_URL
AIF_USAGE_GUARD_ENABLED
AIF_USAGE_GUARD_URL
AIF_GUARD_FAIL_OPEN
For production deployments, prefer secrets management for API keys instead of hard-coding credentials in committed configuration files.
Default paths
configs/ai-firewall.conf
/etc/ai-firewall/ai-firewall.conf
Example configurations
Configuration-only examples are available under:
configs/examples/
Runnable deployment examples are available under:
deploy/examples/
Use configs/examples/ for reusable snippets and deploy/examples/ for full Docker Compose evaluation patterns. OpenShift-specific Kustomize assets are provided separately under deploy/openshift/; they do not replace the generic Docker/Compose deployment model.
Guard Orchestration
AI Firewall v0.8.0 supports optional VCAL Security Guard, Privacy Guard, and Usage Guard orchestration with the same guard behavior as v0.7.0. AIF aif_enforcement_mode is independent of guard enforcement.
Typical full enterprise configuration:
security_guard_enabled true;
security_guard_url http://vcal-security-guard:8091;
security_guard_api_key dev-security-key;
security_guard_timeout_seconds 3;
security_guard_block_response completion;
security_guard_block_message "This interaction cannot be completed because it was blocked by your organization's AI security policy.";
privacy_guard_enabled true;
privacy_guard_url http://vcal-privacy-guard:8090;
privacy_guard_api_key dev-privacy-key;
privacy_guard_mode anonymize;
privacy_guard_restore_enabled true;
privacy_guard_timeout_seconds 3;
usage_guard_enabled true;
usage_guard_url http://vcal-usage-guard:8095;
usage_guard_api_key dev-usage-key;
usage_guard_mode enforce;
usage_guard_tenant_id your-tenant-id;
usage_guard_policy_id business-use-only;
usage_guard_timeout_seconds 3;
usage_guard_block_response completion;
usage_guard_block_message "This request appears to fall outside your organization's approved AI usage policy and cannot be processed.";
guard_fail_open false;
completion remains the default block-response mode for Security Guard and Usage Guard. It returns an OpenAI-compatible assistant completion with HTTP 200. Set the relevant *_block_response directive to error_json when the client should receive the structured HTTP 403 guard error instead.
For enterprise guard deployments, fail-closed behavior is recommended unless availability-over-enforcement is an explicit policy choice:
guard_fail_open false;
VCAL Audit evidence delivery
AI Firewall v0.8.0 can deliver vcal.evidence.event schema version 1.1 batches to VCAL Audit. Evaluation Mode adds attributes within the existing schema so hypothetical cache decisions remain distinguishable from applied actions.
audit_enabled false;
audit_url http://vcal-audit:8092;
# audit_api_key replace-with-shared-audit-token;
audit_producer_instance_id ai-firewall-01;
audit_queue_capacity 10000;
audit_batch_size 100;
audit_flush_interval_ms 1000;
audit_timeout_seconds 5;
audit_retry_max_attempts 5;
audit_retry_initial_backoff_ms 250;
audit_retry_max_backoff_ms 5000;
When enabled, delivery is asynchronous and uses a bounded in-memory queue. Batches are sent when the configured batch size is reached or the flush interval expires. Failed deliveries are retried with backoff. After retry exhaustion, the batch is dropped and the failure is logged rather than blocking the LLM request path.