OpenAI-Compatible Providers
AI Cost Firewall supports practical OpenAI-compatible chat and embedding endpoints while keeping the flat configuration model.
Compatibility scope
AI Cost Firewall focuses on OpenAI-compatible APIs only:
- chat requests use the OpenAI-compatible
POST /v1/chat/completionsshape - chat-side model discovery uses
GET /v1/modelsand is proxied to the configured chat/inference upstream - embeddings use the OpenAI-compatible
/v1/embeddingsshape when semantic cache is enabled - provider root URLs and
/v1base paths are supported - full endpoint paths such as
/v1/chat/completions,/v1/embeddings, or/v1/modelsshould not be configured as base URLs
AI Cost Firewall does not add native provider-specific API integrations or provider-specific configuration blocks. Use the existing flat directives for both cloud and local OpenAI-compatible providers.
Supported patterns
| Provider | Typical role |
|---|---|
| OpenAI | Chat and embeddings |
| Ollama | Local chat and embeddings |
| LM Studio | Local desktop inference |
| vLLM | Self-hosted GPU inference |
| LiteLLM | Aggregation/proxy layer |
| OpenRouter | Multi-provider upstream |
upstream_provider openai_compatible;
upstream_base_url <base-url>;
upstream_api_key <key-or-placeholder>;
embedding_provider openai_compatible;
embedding_base_url <base-url>;
embedding_api_key <key-or-placeholder>;
The base URL may be either the provider root URL or its /v1 base path:
https://api.openai.com
https://api.openai.com/v1
http://ollama:11434
http://ollama:11434/v1
http://lmstudio:1234/v1
http://vllm:8000/v1
http://litellm:4000/v1
Do not configure a full endpoint path:
# Wrong
upstream_base_url http://ollama:11434/v1/chat/completions;
upstream_base_url http://vllm:8000/v1/models;
# Correct
upstream_base_url http://ollama:11434/v1;
upstream_base_url http://vllm:8000/v1;
For local providers without authentication, use a placeholder key:
upstream_api_key dummy;
embedding_api_key dummy;
Accepted placeholder values are dummy, none, null, and -.
vLLM chat + embeddings
vLLM can be used as the chat/inference upstream and, independently, as the OpenAI-compatible embedding endpoint. A common self-hosted layout is:
OpenAI-compatible client / Open WebUI
-> AIF
-> vLLM chat service
-> vLLM embedding service -> Qdrant
Example base URLs:
upstream_base_url http://vllm-chat:8000/v1;
embedding_base_url http://vllm-embeddings:8000/v1;
GET /v1/models through AIF discovers models from upstream_base_url; it does not merge or expose the separate embedding service model list.
Before enabling semantic cache with a Nomic or other embedding model, confirm the exact served embedding model ID and issue a test embedding request. Set qdrant_vector_size to the actual returned vector length. Do not infer the dimension from the model family name alone.
OpenAI-style content arrays and cache behavior
AIF v0.8.2 accepts OpenAI-compatible non-string message.content values such as content arrays and preserves them for upstream forwarding and exact-cache identity. Top-level and message-level extension fields are also preserved in exact-cache identity, including tool-call and provider-specific fields.
Semantic caching is deliberately skipped when:
- a request contains
tools; - a request contains
response_format; or - any
message.contentvalue is not a JSON string.
This prevents image/file/audio payloads or structured content from being serialized into the text embedding path. The current Security, Privacy, and Usage Guard integrations inspect string message content only; nested text parts inside content arrays are not yet independently scanned or transformed.
Deployment examples
Runnable examples are available under:
deploy/examples/
Recommended examples:
openai-cloud/
local-ollama/
hybrid-openai-local-embeddings/
openrouter/
local-full-stack/
Runtime compatibility check
Use /version to confirm the running release and compatibility assumptions:
curl -s http://localhost:8080/version
The response includes supported_api_style and provider_specific_config_blocks so operators can verify that the deployment is using the intended compatibility model. It also reports the active AIF enforcement mode and effective cache scope.
Guard orchestration and provider compatibility
VCAL Security Guard, VCAL Privacy Guard, and VCAL Usage Guard operate at the AI Firewall layer and are independent of the selected OpenAI-compatible upstream provider.
Provider compatibility still matters for:
- OpenAI-compatible chat request and response shape;
- model naming;
- streaming behavior;
- tool/function response formats;
- embedding endpoint behavior when semantic cache is enabled.
AI Firewall supports both ordinary JSON chat completions and controlled OpenAI-compatible streaming. For stream=false or omitted-stream requests, the upstream only needs to return an OpenAI-compatible JSON chat completion. For stream=true, the configured upstream must additionally support OpenAI-compatible SSE.
AIF consumes provider SSE internally and assembles the complete canonical response before response controls and downstream replay. Provider streaming capability is therefore optional unless clients actually request streaming.
Evaluation Mode does not reduce upstream provider traffic: even a would-have cache hit still calls the live upstream provider. Capacity and rate-limit planning for an observe-mode pilot should therefore assume live request volume remains upstream-bound.
The current guard modules inspect string message content only. Non-string content arrays/objects can be preserved and proxied by AIF v0.8.2, but nested text parts inside those structures are not yet scanned, anonymized, or classified by the guard integrations.