Skip to main content

Performance Characteristics

Performance depends on infrastructure, Redis latency, Qdrant latency, embedding provider latency, upstream model latency, and cache hit rate.

ScenarioExpected behavior
Exact cache hitfastest path; no upstream call; no embedding lookup
Semantic cache hitavoids upstream chat call; requires embedding and Qdrant lookup
Cache missdominated by upstream chat model latency
Observe-mode would-have hitstill incurs live upstream latency/cost; records the optimization opportunity instead of serving shadow cache

Semantic cache latency depends on embedding provider latency and Qdrant search latency.

In observe mode, cache evaluation does not reduce live request latency because AIF still calls the upstream provider. It is an assessment mode, not a latency optimization mode. Pilot capacity planning should therefore assume the same upstream request volume as the original application traffic.

Cost and latency trade-offs​

Exact cache hits usually provide the clearest performance and cost benefit because they avoid the upstream chat call without requiring an embedding lookup.

Semantic cache hits can avoid expensive upstream chat completions, but they introduce embedding overhead. This means semantic caching should be evaluated by both latency and net savings:

net savings = avoided upstream chat cost - embedding overhead

Semantic caching is usually most valuable when prompts are repeated or semantically similar, upstream models are relatively expensive, and cached responses are reused often enough to justify embedding cost.

For very cheap models or low-repeat workloads, semantic cache hit rates and gross savings may look positive while net savings are lower because embedding overhead reduces the benefit.

Evaluation metrics​

For Evaluation Mode, use the dedicated aif_evaluation_* metrics rather than production cache/savings counters. Would-have hits are intentionally excluded from normal production savings.

Useful evaluation metrics include:

aif_evaluation_requests_total
aif_evaluation_cache_outcomes_total
aif_evaluation_upstream_calls_avoided_total
aif_evaluation_tokens_avoided_total
aif_evaluation_gross_saved_micro_usd_total
aif_evaluation_net_saved_micro_usd_total
aif_evaluation_errors_total

Useful latency and runtime metrics​

aif_upstream_request_duration_seconds
aif_upstream_timeouts_total
aif_upstream_calls_total
aif_stream_generation_duration_seconds
aif_stream_client_time_to_first_byte_seconds
aif_stream_upstream_response_bytes
aif_cache_exact_hits
aif_cache_semantic_hits
aif_cache_misses
aif_embedding_request_duration_seconds
aif_embedding_timeouts_total
aif_semantic_lookup_duration_seconds

Useful cost and savings metrics​

aif_model_cost_micro_usd_total{model="..."}
aif_gross_saved_micro_usd_total{model="...", cache_type="exact|semantic"}
aif_net_saved_micro_usd_total{model="...", cache_type="exact|semantic"}
aif_embedding_overhead_micro_usd_total{model="...", operation="lookup|store"}
aif_request_cost_micro_usd_total{model="...", cost_type="chat|embedding"}
aif_model_requests_total{model="..."}
aif_cache_hits_total{model="...", cache_type="exact|semantic"}

All cost values are reported in micro-USD.

1 USD = 1,000,000 micro-USD