Grafana Dashboards
AI Cost Firewall includes preconfigured Grafana dashboards for production runtime visibility, cache behavior, provider diagnostics, guard orchestration, and cost/savings analysis. v0.8.0 also exposes dedicated Evaluation Mode Prometheus metrics; these should be visualized and labeled separately from realized production hits and savings.
The Docker Compose observability stack includes Prometheus and Grafana configuration.
Start the stack:
docker compose up -d
Open Grafana:
http://localhost:3000
Default dashboard files are located in the main repository under:
deploy/grafana/dashboards/
and are provided automatically by supported Compose examples.
Overview dashboard
The Grafana Overview dashboard is intended for detailed traffic, cache, and cost analysis.
It answers questions such as:
- How much traffic is going through the firewall?
- How many requests are served from exact or semantic cache?
- How much provider cost was incurred?
- How much cost was avoided by caching?
- How much embedding overhead was introduced by semantic caching?
- What is the net savings after embedding overhead?
- Which models generate the most spend or savings?
Main panels include:
- Total requests
- Provider cost
- Gross savings
- Embedding overhead
- Net savings
- Net savings percentage
- Cache hit rate
- Readiness
- Cost and savings breakdown
- Savings by cache type
- Spend and savings by model
- Exact vs semantic hit rate
- Cache bypass request rate
- Net saved per cache hit
- Request cost by cost type
- Embedding overhead by operation
Diagnostics dashboard
The Diagnostics dashboard is intended for troubleshooting and explaining why savings are high, low, or lower than expected.
It answers questions such as:
- Are semantic lookups happening?
- Are semantic candidates being checked?
- Are candidates passing or failing the similarity threshold?
- Are expired semantic entries being skipped?
- Are semantic cache writes succeeding?
- Is embedding overhead significant?
- Are providers returning authentication, connectivity, DNS, TLS, rate-limit, or timeout errors?
- Are guard orchestration latency or failures increasing?
- Are Usage Guard policy blocks concentrated in particular categories?
The Diagnostics dashboard focuses on runtime behavior rather than executive or operational summaries.
Controlled streaming diagnostics
The v0.7.0 Diagnostics dashboard adds controlled-stream visibility for:
- stream outcomes: requests, completions, errors, aborts, and upstream stream failures;
- upstream TTFB, generation duration, client commit TTFB, and total stream duration;
- upstream SSE intake rate and stream failures;
- cumulative upstream response size and approved client replay-buffer size.
These panels help distinguish provider-generation delay from AIF's post-generation control and replay stages.
Evaluation Mode dashboard guidance
When AIF runs with aif_enforcement_mode observe;, do not merge aif_evaluation_* series into normal production savings panels.
Recommended evaluation labels are Would-have, Potential, and Estimated. Useful pilot views include:
- requests evaluated;
- would-have exact and semantic hits;
- potential cache-hit rate;
- potentially avoidable upstream calls;
- potentially avoidable tokens;
- estimated gross/net cost avoidance;
- exact versus semantic contribution;
- evaluation errors and shadow backend recovery.
The shipped Overview and Diagnostics dashboards remain production-oriented unless your deployment adds dedicated evaluation panels.
Cost accounting model
AI Cost Firewall separates gross savings, embedding overhead, and net savings.
gross savings = avoided upstream chat completion cost
embedding overhead = cost of semantic lookup/store embedding calls
net savings = gross savings - embedding overhead
For exact cache hits:
net savings ≈ gross savings
For semantic cache hits:
net savings = avoided chat cost - embedding overhead
This distinction is important because semantic caching can avoid expensive chat completions, but it also requires embedding calls. Exact and semantic cache savings should therefore be evaluated separately.
Main cost metrics
The dashboards use the following cost-intelligence metrics:
aif_model_cost_micro_usd_total{model="..."}
aif_model_requests_total{model="..."}
aif_model_input_tokens_total{model="..."}
aif_model_output_tokens_total{model="..."}
aif_gross_saved_micro_usd_total{model="...", cache_type="exact|semantic"}
aif_net_saved_micro_usd_total{model="...", cache_type="exact|semantic"}
aif_embedding_overhead_micro_usd_total{model="...", operation="lookup|store"}
aif_request_cost_micro_usd_total{model="...", cost_type="chat|embedding"}
aif_cache_hits_total{model="...", cache_type="exact|semantic"}
All cost values are reported in micro-USD:
1 USD = 1,000,000 micro-USD
Useful PromQL examples
Provider spend by model:
sum by (model) (
increase(aif_model_cost_micro_usd_total[$__range])
) / 1000000
Gross savings by cache type:
sum by (cache_type) (
increase(aif_gross_saved_micro_usd_total[$__range])
) / 1000000
Net savings by cache type:
sum by (cache_type) (
increase(aif_net_saved_micro_usd_total[$__range])
) / 1000000
Embedding overhead by operation:
sum by (operation) (
increase(aif_embedding_overhead_micro_usd_total[$__range])
) / 1000000
Net savings percentage:
100 *
sum(increase(aif_net_saved_micro_usd_total[$__range]))
/
clamp_min(
sum(increase(aif_model_cost_micro_usd_total[$__range]))
+
sum(increase(aif_net_saved_micro_usd_total[$__range])),
1
)
Guard observability boundaries
AI Cost Firewall exposes orchestration-level guard metrics such as:
aif_guard_requests_total
aif_guard_latency_seconds
aif_security_blocks_total
aif_privacy_restore_skipped_total
aif_usage_blocks_total
Prometheus can also scrape VCAL Security Guard, VCAL Privacy Guard, and VCAL Usage Guard directly for module-specific metrics.
Request-level evidence belongs in VCAL Audit rather than Prometheus. Compliance can consume retained Audit evidence without turning Prometheus into an evidence store.
Notes
For meaningful cost panels, configure model pricing in the firewall configuration:
model_price <model> <input_usd_per_1m_tokens> <output_usd_per_1m_tokens>;
embedding_price <usd_per_1m_tokens>;
If embedding_price is not configured, embedding overhead is treated as 0, and semantic net savings may be overestimated.
Most deployment examples include an optional:
docker-compose.observability.yml
overlay for Prometheus and Grafana.