Troubleshooting
Confirm the running release
Before debugging a pilot deployment, confirm the running binary and compatibility model:
curl -s http://localhost:8080/version
For v0.8.2, the response should report the running 0.8.2 release, supported_api_style as openai_compatible, provider_specific_config_blocks as false, and the active AIF enforcement mode/cache scope.
Low cache hit rate
Check:
aif_cache_exact_hits
aif_cache_semantic_hits
aif_cache_misses
aif_semantic_threshold_results_total{result="fail"}
Common causes: threshold too high, prompts not similar, semantic cache disabled, retention too short.
Startup diagnostics fail before the server is ready
Check the firewall logs first. The startup diagnostics should identify whether the failure is related to configuration, Redis, Qdrant, vector-size compatibility, upstream connectivity, embedding provider behavior, DNS, connection errors, or TLS/certificate validation.
Use /healthz, /startupz, /readyz, and /version together:
curl -i http://localhost:8080/healthz
curl -i http://localhost:8080/startupz
curl -i http://localhost:8080/readyz
curl -s http://localhost:8080/version
/healthz only shows that the process is alive. /startupz verifies that enabled Redis/Qdrant caches marked readiness-required initialized in this process. /readyz shows whether the instance should currently receive traffic. /version confirms what release is running.
Redis connection failure
Docker Compose:
redis_url redis://redis:6379;
Local source run:
redis_url redis://127.0.0.1:6379;
Qdrant connection failure
Docker Compose:
qdrant_url http://qdrant:6334;
Local source run:
qdrant_url http://127.0.0.1:6334;
Qdrant vector size mismatch
Fix by recreating the collection, using a matching embedding model, or updating qdrant_vector_size. For a separate vLLM/Nomic embedding service, query its exact served model ID and issue one embedding request; set qdrant_vector_size to the actual returned embedding vector length rather than assuming a dimension from the model family name.
ai-firewall: command not found
After source build, use:
./target/release/ai-firewall
Or install it:
sudo install -m 0755 target/release/ai-firewall /usr/local/bin/ai-firewall
/metrics shows one in-flight request
This is normal because the metrics request itself is active.
Wrong OpenAI-compatible base URL
Use a provider root URL or /v1 base path.
Correct:
upstream_base_url http://ollama:11434/v1;
Wrong:
upstream_base_url http://ollama:11434/v1/chat/completions;
/v1/models fails or Open WebUI cannot discover a model
AIF v0.8.2 proxies /v1/models to upstream_base_url. Confirm that the configured chat upstream exposes an OpenAI-compatible model-list endpoint and that the base URL is a provider root or /v1 base, not the full /v1/models path.
curl -i http://localhost:8080/v1/models
For vLLM, test the chat service directly as well. If discovery succeeds but chat requests are rejected by AIF, check allow_unknown_models_pass_through and model_price; discovery does not automatically authorize a model.
Upstream provider errors
Check:
aif_errors_total{class="upstream_authentication_error"}
aif_errors_total{class="upstream_not_found"}
aif_errors_total{class="upstream_rate_limited"}
aif_errors_total{class="upstream_tls_error"}
aif_errors_total{class="upstream_dns_error"}
aif_errors_total{class="upstream_connect_error"}
Common causes:
wrong provider API key
- full endpoint path configured instead of base URL
- provider hostname cannot be resolved
- provider port is unreachable
- self-signed or hostname-mismatched TLS certificate
Embedding provider timeouts
Check:
aif_embedding_timeouts_total
aif_embedding_request_duration_seconds
Common causes:
- wrong provider API key
- full endpoint path configured instead of base URL
- provider hostname cannot be resolved
- provider port is unreachable
- self-signed or hostname-mismatched TLS certificate
Dashboards are empty
Common causes:
- no traffic has been sent yet
- Prometheus is not scraping the firewall
- Grafana datasource is not connected
- wrong Compose working directory
- dashboard provisioning paths are wrong
Check:
curl http://localhost:8080/metrics
and open:
http://localhost:9090/targets
Security Guard blocks or errors
Module overview: VCAL Security Guard.
Symptoms can include either a safe HTTP 200 assistant completion or a structured HTTP 403 guard error, depending on security_guard_block_response.
Structured error identifiers include:
security_request_blocked
security_response_blocked
security_guard_unavailable
security_guard_timeout
Common causes:
- Security Guard detected prompt injection, jailbreak, or system-prompt extraction text;
- Security Guard is running in
enforcemode; - Security Guard is unavailable and
guard_fail_open false; - API key mismatch between AI Firewall and Security Guard;
- wrong
security_guard_url; - Security Guard default mode is still
detect_onlywhen blocking was expected.
Recommended checks:
curl http://localhost:8091/healthz
curl http://localhost:8091/readyz
curl -s http://localhost:8080/metrics | grep 'aif_guard_requests_total'
Check configuration:
security_guard_enabled true;
security_guard_url http://vcal-security-guard:8091;
security_guard_api_key dev-security-key;
security_guard_timeout_seconds 3;
security_guard_block_response completion;
guard_fail_open false;
For production-like enforcement tests, Security Guard should normally use:
VCAL_SECURITY_GUARD_DEFAULT_MODE=enforce
To receive the structured HTTP 403 block instead of the default safe assistant completion, configure:
security_guard_block_response error_json;
Usage Guard blocks or policy errors
Module overview: VCAL Usage Guard.
Common symptoms include an organization-policy safe completion, usage_request_blocked, unexpected allow decisions, or Usage Guard availability errors.
Recommended checks:
curl http://localhost:8095/healthz
curl http://localhost:8095/readyz
curl http://localhost:8095/metrics
curl -s http://localhost:8080/metrics | grep -E 'aif_guard_requests_total|aif_usage_blocks_total'
Check AI Firewall configuration:
usage_guard_enabled true;
usage_guard_url http://vcal-usage-guard:8095;
usage_guard_api_key your-usage-guard-key;
usage_guard_mode enforce;
usage_guard_tenant_id your-tenant-id;
usage_guard_policy_id business-use-only;
usage_guard_timeout_seconds 3;
usage_guard_block_response completion;
guard_fail_open false;
If completion is configured, a policy block or escalate returns a safe HTTP 200 assistant message without cache or upstream execution. Set usage_guard_block_response error_json; when you need the structured HTTP 403 usage_request_blocked response for testing or client integration.
If semantic policy matching is enabled in Usage Guard, also verify the embeddings path used by that policy. Explicit deterministic rules do not require semantic matching.
Privacy Guard restore or anonymization errors
Module overview: VCAL Privacy Guard.
Symptoms:
privacy_guard_unavailable
privacy_guard_timeout
privacy_restore_failed
guard_contract_violation
or final responses contain placeholders such as [EMAIL_1] or [IP_1].
Common causes:
- Privacy Guard unavailable and
guard_fail_open false; - API key mismatch between AI Firewall and Privacy Guard;
- wrong
privacy_guard_url; - expired or missing mapping ID;
- restore disabled;
- response was blocked by Security Guard before restore;
- non-text content was expected to be scanned or restored.
Recommended checks:
curl http://localhost:8090/healthz
curl http://localhost:8090/readyz
curl -s http://localhost:8090/metrics | grep vcal_privacy
curl -s http://localhost:8080/metrics | grep -E 'aif_guard_requests_total|aif_privacy_restore_skipped_total'
VCAL Audit delivery failures
Module overview: VCAL Audit.
Symptoms include missing traces in Audit, delayed evidence delivery, repeated delivery retries, or dropped batches in AI Firewall logs.
Recommended checks:
curl http://localhost:8092/healthz
curl http://localhost:8092/readyz
curl http://localhost:8092/metrics
curl -s -H "X-API-Key: $AUDIT_API_KEY" \
"http://localhost:8092/v1/events?limit=1&after_sequence=0"
Verify:
audit_enabled true;audit_urlresolves from the AI Firewall container;- AI Firewall and VCAL Audit share the expected Docker network;
- the configured Audit API key matches the service;
- VCAL Audit accepts
POST /v1/events/batch; - AI Firewall logs show buffered Audit delivery initialization.
The sender queue is memory-backed. A batch that remains undeliverable after retry exhaustion can be dropped and is not replayed automatically after restart.
VCAL Compliance cannot read Audit evidence
Module overview: VCAL Compliance.
VCAL Compliance runs downstream of VCAL Audit. If Compliance is healthy but no evaluations or evidence-backed results appear, verify the Audit side first.
Recommended checks:
curl http://localhost:8093/healthz
curl http://localhost:8093/readyz
curl http://localhost:8093/metrics
curl http://localhost:8092/healthz
curl http://localhost:8092/readyz
Common causes:
- VCAL Audit is unavailable or not ready;
- Compliance is configured with the wrong Audit URL;
vcal-compliancecannot resolve or reachvcal-audit:8092on the Docker network;- API credentials between Compliance and Audit do not match;
- Audit has not received the expected AI Firewall evidence yet;
- the Compliance control configuration is missing or invalid.
From the shared Docker network, both services should be reachable at their service names:
vcal-audit:8092
vcal-compliance:8093
Troubleshoot the chain in order: AI Firewall -> VCAL Audit -> VCAL Compliance.
To determine whether synchronization is currently progressing, check:
curl -s http://localhost:8093/metrics | \
grep -E 'vcal_compliance_(imported_events_total|last_audit_sequence|sync_runs_total)'
A historical failed counter does not necessarily indicate a current outage. If vcal_compliance_last_audit_sequence and vcal_compliance_sync_runs_total{result="success"} continue to increase, the synchronization path is active.
Controlled streaming problems
AI Cost Firewall v0.8.2 retains controlled stream=true requests when streaming_enabled true; is configured.
If a streaming request is rejected before upstream processing, confirm:
streaming_enabled true;
If normal JSON requests work but stream=true fails at the provider boundary, verify that the configured upstream supports OpenAI-compatible SSE. Provider streaming support is required only for streaming requests.
For a provider that sends response headers and then stops producing body chunks, AIF applies upstream_timeout_seconds as the maximum idle gap between SSE chunks. The expected client result is a normal HTTP timeout error before downstream SSE commit.
A provider that continuously drip-feeds chunks cannot hold the upstream permit forever: controlled generation has a separate 15-minute absolute ceiling.
If the provider stream is malformed, truncated, oversized, idle, or exceeds the absolute generation ceiling, check:
aif_stream_errors_total
aif_stream_upstream_errors_total
aif_upstream_timeouts_total
and inspect Audit/evidence for upstream.stream.failed. Partial provider content must not be exposed to the client when committed_to_client=false.
Content arrays pass through but nested parts are not guard-inspected
AIF v0.8.2 preserves non-string message.content values such as OpenAI-style content arrays for upstream forwarding and exact-cache identity. Such requests are deliberately excluded from semantic cache. Requests containing tools or response_format are also semantically ineligible.
The current Security, Privacy, and Usage Guard integrations inspect string message content only. Text nested inside content arrays is not yet independently scanned, anonymized, or classified. If a client extracts OCR text, captions, or transcripts and sends them as ordinary string message content, that text can be processed normally.
Missing or duplicate terminal evidence events
Inspect request lifecycle events:
docker compose logs firewall 2>&1 | grep -E 'request\.(received|completed|failed)'
Every trace that emits request.received must have exactly one terminal event: either request.completed or request.failed.
Enable evidence logs with:
RUST_LOG=info,vcal_evidence=info
If lifecycle events are missing, confirm that the running image is the latest one and that vcal_evidence=info is not filtered by the logging configuration.
Observe mode reports no evaluation hits
Confirm the active runtime mode:
curl -s http://localhost:8080/version
Look for:
{
"aif_enforcement_mode": "observe",
"effective_cache_scope": "evaluation"
}
Then check evaluation metrics rather than production cache-hit counters:
curl -s http://localhost:8080/metrics | grep 'aif_evaluation_'
Would-have hits intentionally do not increment normal production cache-hit/savings metrics.
Redis or Qdrant fails during Evaluation Mode
Evaluation dependencies are non-blocking in observe mode. The live upstream request should continue and /readyz should remain healthy.
Check:
curl -i http://localhost:8080/readyz
curl -s http://localhost:8080/metrics | grep aif_evaluation_errors_total
If Redis/Qdrant initialized successfully and fails later at runtime, AIF can resume the corresponding shadow cache path after recovery.
A different rule applies when a cache backend was unavailable when the AIF process started and fail-open/Observe startup installed a no-op cache. That process does not replace the no-op cache automatically. If the enabled backend is configured as readiness-required, /startupz returns 503; restart the pod/process after the dependency is available so the real cache client is initialized.
For Qdrant, also confirm that the effective evaluation collection reported by /version is reachable and uses the expected vector dimension.
Observe-mode cost appears unchanged
That is expected for actual provider spend. Observe Mode still calls the live upstream provider even on a would-have cache hit.
Use aif_evaluation_gross_saved_micro_usd_total, aif_evaluation_net_saved_micro_usd_total, and related evaluation metrics to estimate the cost that enforcement could potentially avoid. Do not compare those counters directly with realized production savings without labeling them as prospective estimates.