VCAL Usage Guard
VCAL Usage Guard is an optional commercial module for applying organizational acceptable-use policy to AI requests before cache lookup or upstream model execution.
It evaluates the request against the selected tenant and policy and returns an explicit allow, warn, block, or escalate decision. Deterministic rules do not require a local LLM. Optional semantic matching can use a compatible embeddings path when the Usage Guard policy enables it.
Integration point
Client
-> AI Cost Firewall
-> VCAL Security Guard request scan (optional)
-> VCAL Privacy Guard anonymize/redact (optional)
-> VCAL Usage Guard policy evaluation
-> cache / upstream LLM
-> VCAL Security Guard response scan (optional)
-> VCAL Privacy Guard restore (optional)
-> Client
Usage Guard runs after Privacy Guard on the request path. When Privacy Guard anonymization is enabled, the usage policy is therefore evaluated against anonymized text.
Typical AI Firewall configuration:
usage_guard_enabled true;
usage_guard_url http://vcal-usage-guard:8095;
usage_guard_api_key your-usage-guard-key;
usage_guard_mode enforce;
usage_guard_tenant_id your-tenant-id;
usage_guard_policy_id business-use-only;
usage_guard_timeout_seconds 3;
usage_guard_block_response completion;
usage_guard_block_message "This request appears to fall outside your organization's approved AI usage policy and cannot be processed.";
guard_fail_open false;
Client response behavior
AI Cost Firewall separates the Usage Guard policy decision from the client-facing response.
usage_guard_block_response completion;
returns HTTP 200 with a safe OpenAI-compatible assistant message for block or escalate. No cache lookup, cache write, or upstream LLM call is performed.
To expose the structured guard error instead:
usage_guard_block_response error_json;
A blocked request then returns HTTP 403 with usage_request_blocked.
Example policy outcome
A business-use-only policy can classify a request such as:
Plan a two-week vacation in Thailand for my family.
with policy metadata such as:
decision: block
category: personal_travel
The exact categories and rules are defined by the deployed Usage Guard policy.
Observability
AI Cost Firewall exposes orchestration-level Usage Guard metrics, including:
aif_guard_requests_total
aif_guard_latency_seconds
aif_usage_blocks_total
aif_usage_blocks_total records Usage Guard block outcomes by available policy category and rule metadata.
Usage Guard can also expose its own module-specific metrics and dashboard.
Failure behavior
guard_fail_open controls what happens when Usage Guard is unavailable, times out, or returns an invalid response contract.
guard_fail_open false;
is the recommended enterprise default when policy enforcement must not be bypassed.
AI Cost Firewall v0.8.2 retains controlled stream=true requests. Usage Guard still evaluates the request before cache lookup or upstream processing; the assembled response then follows the same response-control pipeline before approved SSE delivery. aif_enforcement_mode observe affects AIF cache evaluation only and does not change Usage Guard enforcement behavior.
AI Cost Firewall can preserve OpenAI-style non-string message.content arrays/objects and forward them upstream, but the current Usage Guard integration evaluates string message content only. Nested text parts inside content arrays are not yet independently classified. Extracted OCR text, captions, or transcripts can be evaluated once supplied as ordinary string content.