Metrics
AI Cost Firewall exposes Prometheus-format metrics at:
/metrics
Example:
curl http://localhost:8080/metrics
The current v0.8.2 configuration model can optionally require authentication for /metrics with metrics_auth_required and metrics_auth_token. The default remains unauthenticated Prometheus scraping.
Request and cache metrics
aif_requests_total
aif_cache_exact_hits
aif_cache_semantic_hits
aif_cache_misses
aif_upstream_calls_total
These metrics show total request volume, cache outcomes, and how often requests are forwarded to the upstream LLM provider.
Per-model request and token metrics
aif_model_requests_total{model="..."}
aif_model_input_tokens_total{model="..."}
aif_model_output_tokens_total{model="..."}
These metrics show which models are being used and how many upstream input/output tokens they consume.
aif_model_requests_total counts upstream chat completion requests by model. It is useful as a denominator for average cost per upstream request.
Cost and savings metrics
AI Cost Firewall reports costs in micro-USD.
1 USD = 1,000,000 micro-USD
Backward-compatible aggregate cost metrics:
aif_chat_cost_saved_micro_usd
aif_embedding_cost_micro_usd
aif_cost_saved_micro_usd
Structured cost-intelligence metrics:
aif_model_cost_micro_usd_total{model="..."}
aif_request_cost_micro_usd_total{model="...", cost_type="chat|embedding"}
aif_gross_saved_micro_usd_total{model="...", cache_type="exact|semantic"}
aif_net_saved_micro_usd_total{model="...", cache_type="exact|semantic"}
aif_embedding_overhead_micro_usd_total{model="...", operation="lookup|store"}
aif_cache_hits_total{model="...", cache_type="exact|semantic"}
Cost accounting model
The main accounting model is:
gross savings = avoided upstream chat completion cost
embedding overhead = cost of semantic lookup/store embedding calls
net savings = gross savings - embedding overhead
For exact cache hits:
net savings ≈ gross savings
For semantic cache hits:
net savings = avoided chat cost - embedding overhead
Exact cache hits do not require embedding lookup. Semantic cache hits avoid upstream chat calls, but require embedding lookup to search for similar cached prompts.
Semantic cache storage can also create embedding overhead because cache misses may be embedded before they are stored for future semantic reuse.
If embedding_price is not configured, embedding overhead is treated as 0, and savings may be overestimated.
Evaluation Mode metrics
When aif_enforcement_mode observe; is active, would-have cache outcomes are kept separate from normal production cache-hit and savings counters.
aif_enforcement_mode_info{mode="observe"}
aif_evaluation_requests_total
aif_evaluation_cache_outcomes_total{cache_type="exact|semantic",result="hit|miss"}
aif_evaluation_upstream_calls_avoided_total{cache_type="exact|semantic"}
aif_evaluation_tokens_avoided_total{model="...",cache_type="exact|semantic"}
aif_evaluation_gross_saved_micro_usd_total{model="...",cache_type="exact|semantic"}
aif_evaluation_net_saved_micro_usd_total{model="...",cache_type="exact|semantic"}
aif_evaluation_shadow_store_total
aif_evaluation_errors_total{component="...",operation="..."}
These metrics represent potential behavior if normal enforcement were enabled. They are not actual cache hits or realized savings.
A shadow hit still calls the live upstream provider. Prospective token/cost avoidance is derived from the actual live upstream response for that request, representing the call enforcement would have avoided.
The following production metrics are therefore not incremented merely because Evaluation Mode finds a shadow hit:
aif_cache_exact_hits
aif_cache_semantic_hits
aif_cache_hits_total
aif_gross_saved_micro_usd_total
aif_net_saved_micro_usd_total
Example potential hit rate:
sum(increase(aif_evaluation_cache_outcomes_total{result="hit"}[$__range]))
/
clamp_min(sum(increase(aif_evaluation_requests_total[$__range])), 1)
Potential upstream calls avoided by cache type:
sum by (cache_type) (
increase(aif_evaluation_upstream_calls_avoided_total[$__range])
)
Runtime metrics
aif_inflight_requests
aif_readiness_state
aif_shutdown_in_progress
aif_shutdown_rejections_total
Note: aif_inflight_requests includes the /metrics request itself.
Provider diagnostics
Upstream provider diagnostics:
aif_upstream_timeouts_total
aif_upstream_request_duration_seconds
Embedding provider diagnostics:
aif_embedding_request_duration_seconds
aif_embedding_timeouts_total
These metrics help diagnose slow or unavailable upstream and embedding providers.
Controlled streaming metrics
aif_stream_requests_total
aif_stream_completed_total
aif_stream_errors_total
aif_stream_aborted_total
aif_stream_upstream_errors_total
aif_stream_upstream_chunks_total
aif_stream_upstream_bytes_total
aif_stream_upstream_time_to_first_byte_seconds
aif_stream_generation_duration_seconds
aif_stream_client_time_to_first_byte_seconds
aif_stream_upstream_response_bytes
aif_stream_client_buffer_bytes
aif_stream_duration_seconds
aif_stream_upstream_response_bytes records the cumulative provider SSE bytes consumed for a completed controlled streaming request. It is a response-size observation, not instantaneous parser-buffer occupancy. aif_stream_client_buffer_bytes describes the approved client replay payload materialized after response controls.
aif_upstream_timeouts_total also includes controlled-stream idle and absolute-generation timeout outcomes.
Error metrics
aif_errors_total{class="..."}
Common error classes include:
validation_error
upstream_error
upstream_timeout
upstream_authentication_error
upstream_not_found
upstream_rate_limited
upstream_tls_error
upstream_dns_error
upstream_connect_error
internal_error
These labels help distinguish request validation issues, provider connectivity problems, TLS/certificate errors, upstream rate limits, and internal failures.
Semantic diagnostics
aif_semantic_candidates_checked_total
aif_semantic_threshold_results_total{result="pass"}
aif_semantic_threshold_results_total{result="fail"}
aif_semantic_expired_entries_skipped_total
aif_semantic_lookup_duration_seconds
aif_semantic_store_total
aif_semantic_store_errors_total
These metrics explain semantic cache behavior:
- how many candidates are checked
- how often candidates pass or fail the similarity threshold
- how often expired entries are skipped
- how long semantic lookup takes
- whether semantic cache writes are succeeding
Useful PromQL examples
Estimated chat spend by model:
sum by (model) (
increase(aif_model_cost_micro_usd_total[$__range])
) / 1000000
Gross savings by cache type:
sum by (cache_type) (
increase(aif_gross_saved_micro_usd_total[$__range])
) / 1000000
Net savings by cache type:
sum by (cache_type) (
increase(aif_net_saved_micro_usd_total[$__range])
) / 1000000
Embedding overhead by operation:
sum by (operation) (
increase(aif_embedding_overhead_micro_usd_total[$__range])
) / 1000000
Average upstream chat cost per request:
sum by (model) (
rate(aif_model_cost_micro_usd_total[$__rate_interval])
)
/
clamp_min(
sum by (model) (
rate(aif_model_requests_total[$__rate_interval])
),
1
)
Average net saved cost per cache hit:
sum by (model, cache_type) (
rate(aif_net_saved_micro_usd_total[$__rate_interval])
)
/
clamp_min(
sum by (model, cache_type) (
rate(aif_cache_hits_total[$__rate_interval])
),
1
)
Semantic lookup overhead per semantic hit:
sum by (model) (
rate(aif_embedding_overhead_micro_usd_total{operation="lookup"}[$__rate_interval])
)
/
clamp_min(
sum by (model) (
rate(aif_cache_hits_total{cache_type="semantic"}[$__rate_interval])
),
1
)
Pricing configuration
Cost metrics depend on model pricing configured in AI Cost Firewall.
Example:
model_price gpt-4o-mini-2024-07-18 0.15 0.60;
embedding_price 0.020;
Model prices are configured in USD per 1M input/output tokens.
Embedding price is configured in USD per 1M embedding tokens.
Guard orchestration metrics
AI Firewall adds guard orchestration metrics when VCAL Security Guard, VCAL Privacy Guard, or VCAL Usage Guard is enabled.
aif_guard_requests_total
aif_guard_latency_seconds
aif_security_blocks_total
aif_privacy_restore_skipped_total
aif_usage_blocks_total
These metrics show:
- guard calls by guard, stage, and result;
- guard call latency;
- Security Guard blocks by request/response stage and rule ID;
- Privacy Guard restore skips when a response is blocked before restore;
- Usage Guard blocks by policy category and rule ID.
Useful check:
curl -s http://localhost:8080/metrics | grep -E 'aif_guard_requests_total|aif_security_blocks_total|aif_privacy_restore_skipped_total|aif_usage_blocks_total'
Evidence events are not Prometheus metrics
Trace-level evidence is emitted through structured application logs using vcal.evidence.event schema v1.1.
Enable evidence logging with:
RUST_LOG=info,vcal_evidence=info
Every trace that emits request.received ends with exactly one request.completed or request.failed event. VCAL Audit is available to consume and retain these events separately from Prometheus.