Skip to main content

Qdrant Backend

Qdrant stores embeddings and response payloads for semantic caching.

Ports​

AI Cost Firewall uses Qdrant gRPC:

6334

Qdrant REST usually runs on:

6333

REST is useful for health checks and manual inspection.

Local Docker example​

docker run -d --rm --name qdrant \
-p 6333:6333 \
-p 6334:6334 \
qdrant/qdrant

Health check:

curl http://127.0.0.1:6333/healthz

Firewall config:

qdrant_url http://127.0.0.1:6334;

Vector size​

qdrant_vector_size 1536;

This must match the embedding model. Existing production collections are validated for normal enforced semantic-cache use.

In observe mode, AIF uses an isolated evaluation collection derived from the configured production collection (for example, aif_semantic_cache_eval). Evaluation Qdrant failures are non-blocking for live application traffic and are reported as evaluation errors.

The configured qdrant_vector_size must match the dimension produced by embedding_model. This is especially important when using local embedding providers such as Ollama or LM Studio, where vector dimensions depend on the selected model.