Configuration Directives
Core
listen_addr 0.0.0.0:8080;
aif_enforcement_mode enforce;
redis_url redis://redis:6379;
aif_enforcement_mode supports:
enforce— default production behavior; eligible cache hits can be served;observe— evaluate exact/semantic cache opportunities using isolated shadow state while continuing to the live upstream provider.
Environment override:
AIF_ENFORCEMENT_MODE=observe
The directive controls AIF caching/cost optimization only. It does not change Security Guard, Privacy Guard, or Usage Guard enforcement behavior.
Upstream
upstream_provider openai_compatible;
upstream_base_url https://api.openai.com;
upstream_api_key sk-xxxx;
upstream_base_url may be the provider root URL or its /v1 base path.
Correct:
upstream_base_url http://ollama:11434/v1;
Wrong:
upstream_base_url http://ollama:11434/v1/chat/completions;
upstream_base_url http://vllm:8000/v1/models;
AIF builds /v1/chat/completions and /v1/models from upstream_base_url internally.
Placeholder no-auth values:
dummy
none
null
-
Controlled streaming
streaming_enabled true;
max_stream_upstream_bytes 8M;
upstream_timeout_seconds 120;
streaming_enabled allows clients to use "stream": true; it does not force streaming for ordinary requests.
max_stream_upstream_bytes is the maximum cumulative provider SSE payload accepted for one controlled streaming request. It is not an instantaneous parser-buffer limit.
For controlled streams, upstream_timeout_seconds also limits the maximum idle gap between provider SSE chunks after response headers are received. A separate 15-minute absolute generation ceiling prevents indefinite drip-feed streams.
The equivalent environment variables are AIF_STREAMING_ENABLED and AIF_MAX_STREAM_UPSTREAM_BYTES.
Embeddings
embedding_provider openai_compatible;
embedding_base_url https://api.openai.com;
embedding_api_key sk-xxxx;
embedding_model text-embedding-3-small;
embedding_price 0.020;
For local providers without authentication, use dummy, none, null, or -. Placeholder keys do not create upstream bearer auth headers.
Qdrant
qdrant_url http://qdrant:6334;
qdrant_api_key your-qdrant-key;
qdrant_collection aif_semantic_cache;
qdrant_vector_size 1536;
qdrant_vector_size must match the embedding model. Existing collections are validated at startup.
Cache lifecycle
cache_ttl_seconds 86400;
exact_cache_ttl_seconds 86400;
semantic_cache_retention_seconds 604800;
Request behavior
request_timeout_seconds 120;
upstream_timeout_seconds 120;
embedding_timeout_seconds 30;
max_request_body_bytes 1M;
Request limits
max_request_body_bytes 1M;
Supported formats:
1024
512K
1M
2M
Semantic cache
semantic_cache_enabled true;
semantic_cache_fail_open true;
semantic_similarity_threshold 0.92;
semantic_cache_fail_open controls normal enforce runtime behavior. In observe mode, Redis/Qdrant/embedding failures used only for evaluation are always non-blocking for the live request and are recorded as evaluation errors. Independently of fail-open behavior, v0.8.2 skips semantic lookup/store for requests containing tools, response_format, or any non-string message.content.
Model pricing
model_price gpt-4o-mini-2024-07-18 0.15 0.60;
allow_unknown_models_pass_through false;
For providers with variable model names, such as OpenRouter, use:
allow_unknown_models_pass_through true;
Metrics endpoint access
metrics_auth_required false;
# metrics_auth_token replace-with-prometheus-token;
When metrics_auth_required is true, callers must present the configured metrics token to access /metrics. The default remains unauthenticated scraping for local Prometheus deployments.
Readiness dependencies
readiness_requires_redis true;
readiness_requires_qdrant false;
readiness_requires_upstream false;
These directives control whether dependency availability affects /readyz during normal enforce operation. In observe mode, optional evaluation cache dependencies do not make AIF unready because evaluation infrastructure must not interrupt the live application path. /startupz additionally requires an enabled Redis/Qdrant cache marked readiness-required to have initialized successfully in the current process, which is useful for Kubernetes/OpenShift startupProbe handling.
Guard orchestration directives
security_guard_enabled
Enable VCAL Security Guard orchestration.
security_guard_enabled true;
Default: false.
security_guard_url
Base URL for VCAL Security Guard.
security_guard_url http://vcal-security-guard:8091;
Do not include /v1/scan; AI Firewall appends the endpoint internally.
security_guard_api_key
Service-to-service API key for VCAL Security Guard.
security_guard_api_key your-security-guard-key;
security_guard_timeout_seconds
Timeout for Security Guard calls.
security_guard_timeout_seconds 3;
security_guard_block_response
Controls the client-facing response when Security Guard returns block.
security_guard_block_response completion;
Supported values:
| Value | Behavior |
|---|---|
completion | Return HTTP 200 with a safe OpenAI-compatible assistant message |
error_json | Return HTTP 403 with the structured Security Guard error |
Default: completion.
security_guard_block_message
Safe assistant message used when security_guard_block_response completion; is configured.
security_guard_block_message "This interaction cannot be completed because it was blocked by your organization's AI security policy.";
privacy_guard_enabled
Enable VCAL Privacy Guard orchestration.
privacy_guard_enabled true;
Default: false.
privacy_guard_url
Base URL for VCAL Privacy Guard.
privacy_guard_url http://vcal-privacy-guard:8090;
Do not include /v1/scan or /v1/restore; AI Firewall appends endpoints internally.
privacy_guard_api_key
Service-to-service API key for VCAL Privacy Guard.
privacy_guard_api_key your-privacy-guard-key;
privacy_guard_mode
Privacy Guard scan mode.
privacy_guard_mode anonymize;
Common values are detect_only, redact, and anonymize.
privacy_guard_restore_enabled
Enable Privacy Guard placeholder restoration on assistant responses.
privacy_guard_restore_enabled true;
privacy_guard_timeout_seconds
Timeout for Privacy Guard scan and restore calls.
privacy_guard_timeout_seconds 3;
usage_guard_enabled
Enable VCAL Usage Guard orchestration.
usage_guard_enabled true;
Default: false.
usage_guard_url
Base URL for VCAL Usage Guard.
usage_guard_url http://vcal-usage-guard:8095;
usage_guard_api_key
Service-to-service API key for Usage Guard.
usage_guard_api_key your-usage-guard-key;
usage_guard_mode
Usage Guard mode.
usage_guard_mode enforce;
The example configuration defaults to detect_only. Use enforce when policy decisions should affect request handling.
usage_guard_tenant_id
Optional tenant identifier sent with Usage Guard evaluations.
usage_guard_tenant_id your-tenant-id;
usage_guard_policy_id
Optional policy identifier sent with Usage Guard evaluations.
usage_guard_policy_id business-use-only;
usage_guard_timeout_seconds
Timeout for Usage Guard evaluation.
usage_guard_timeout_seconds 3;
usage_guard_block_response
Controls the client-facing response when Usage Guard returns block or escalate.
usage_guard_block_response completion;
Supported values:
| Value | Behavior |
|---|---|
completion | Return HTTP 200 with a safe OpenAI-compatible assistant message and do not continue to cache/upstream |
error_json | Return HTTP 403 with the structured usage_request_blocked error |
Default: completion.
usage_guard_block_message
Safe assistant message used when usage_guard_block_response completion; is configured.
usage_guard_block_message "This request appears to fall outside your organization's approved AI usage policy and cannot be processed.";
guard_fail_open
Controls behavior when an enabled guard is unavailable, times out, or returns an invalid contract.
guard_fail_open false;
Recommended enterprise value: false.
| Value | Behavior |
|---|---|
true | Continue request processing when a guard fails |
false | Fail closed and reject the request when a guard fails |
VCAL Audit evidence delivery
AI Firewall can deliver vcal.evidence.event schema version 1.1 batches to VCAL Audit.
audit_enabled false;
audit_url http://vcal-audit:8092;
# audit_api_key replace-with-shared-audit-token;
audit_producer_instance_id ai-firewall-01;
audit_queue_capacity 10000;
audit_batch_size 100;
audit_flush_interval_ms 1000;
audit_timeout_seconds 5;
audit_retry_max_attempts 5;
audit_retry_initial_backoff_ms 250;
audit_retry_max_backoff_ms 5000;
| Directive | Purpose |
|---|---|
audit_enabled | Enables buffered delivery to VCAL Audit. |
audit_url | Base URL of VCAL Audit; batches are posted to /v1/events/batch. |
audit_api_key | Shared API key sent to VCAL Audit. |
audit_producer_instance_id | Stable producer identifier recorded with delivered batches. |
audit_queue_capacity | Maximum number of evidence events held in the in-memory queue. |
audit_batch_size | Maximum number of events sent in one batch. |
audit_flush_interval_ms | Maximum time before a partial batch is flushed. |
audit_timeout_seconds | HTTP request timeout for a batch delivery attempt. |
audit_retry_max_attempts | Maximum delivery attempts for a failed batch. |
audit_retry_initial_backoff_ms | Initial exponential retry backoff. |
audit_retry_max_backoff_ms | Maximum retry backoff between delivery attempts. |
When enabled, delivery is asynchronous and uses a bounded in-memory queue. Batches are sent when the configured batch size is reached or the flush interval expires. Failed deliveries are retried with backoff. After retry exhaustion, the batch is dropped and the failure is logged rather than blocking the LLM request path.