Skip to main content

OpenAI-Compatible Providers

AI Cost Firewall supports practical OpenAI-compatible chat and embedding endpoints while keeping the flat configuration model.

Compatibility scope​

AI Cost Firewall focuses on OpenAI-compatible APIs only:

  • chat requests use the OpenAI-compatible POST /v1/chat/completions shape
  • chat-side model discovery uses GET /v1/models and is proxied to the configured chat/inference upstream
  • embeddings use the OpenAI-compatible /v1/embeddings shape when semantic cache is enabled
  • provider root URLs and /v1 base paths are supported
  • full endpoint paths such as /v1/chat/completions, /v1/embeddings, or /v1/models should not be configured as base URLs

AI Cost Firewall does not add native provider-specific API integrations or provider-specific configuration blocks. Use the existing flat directives for both cloud and local OpenAI-compatible providers.

Supported patterns​

ProviderTypical role
OpenAIChat and embeddings
OllamaLocal chat and embeddings
LM StudioLocal desktop inference
vLLMSelf-hosted GPU inference
LiteLLMAggregation/proxy layer
OpenRouterMulti-provider upstream
upstream_provider openai_compatible;
upstream_base_url <base-url>;
upstream_api_key <key-or-placeholder>;

embedding_provider openai_compatible;
embedding_base_url <base-url>;
embedding_api_key <key-or-placeholder>;

The base URL may be either the provider root URL or its /v1 base path:

https://api.openai.com
https://api.openai.com/v1
http://ollama:11434
http://ollama:11434/v1
http://lmstudio:1234/v1
http://vllm:8000/v1
http://litellm:4000/v1

Do not configure a full endpoint path:

# Wrong
upstream_base_url http://ollama:11434/v1/chat/completions;
upstream_base_url http://vllm:8000/v1/models;

# Correct
upstream_base_url http://ollama:11434/v1;
upstream_base_url http://vllm:8000/v1;

For local providers without authentication, use a placeholder key:

upstream_api_key dummy;
embedding_api_key dummy;

Accepted placeholder values are dummy, none, null, and -.

vLLM chat + embeddings​

vLLM can be used as the chat/inference upstream and, independently, as the OpenAI-compatible embedding endpoint. A common self-hosted layout is:

OpenAI-compatible client / Open WebUI
-> AIF
-> vLLM chat service
-> vLLM embedding service -> Qdrant

Example base URLs:

upstream_base_url http://vllm-chat:8000/v1;
embedding_base_url http://vllm-embeddings:8000/v1;

GET /v1/models through AIF discovers models from upstream_base_url; it does not merge or expose the separate embedding service model list.

Before enabling semantic cache with a Nomic or other embedding model, confirm the exact served embedding model ID and issue a test embedding request. Set qdrant_vector_size to the actual returned vector length. Do not infer the dimension from the model family name alone.

OpenAI-style content arrays and cache behavior​

AIF v0.8.2 accepts OpenAI-compatible non-string message.content values such as content arrays and preserves them for upstream forwarding and exact-cache identity. Top-level and message-level extension fields are also preserved in exact-cache identity, including tool-call and provider-specific fields.

Semantic caching is deliberately skipped when:

  • a request contains tools;
  • a request contains response_format; or
  • any message.content value is not a JSON string.

This prevents image/file/audio payloads or structured content from being serialized into the text embedding path. The current Security, Privacy, and Usage Guard integrations inspect string message content only; nested text parts inside content arrays are not yet independently scanned or transformed.

Deployment examples​

Runnable examples are available under:

deploy/examples/

Recommended examples:

openai-cloud/
local-ollama/
hybrid-openai-local-embeddings/
openrouter/
local-full-stack/

Runtime compatibility check​

Use /version to confirm the running release and compatibility assumptions:

curl -s http://localhost:8080/version

The response includes supported_api_style and provider_specific_config_blocks so operators can verify that the deployment is using the intended compatibility model. It also reports the active AIF enforcement mode and effective cache scope.

Guard orchestration and provider compatibility​

VCAL Security Guard, VCAL Privacy Guard, and VCAL Usage Guard operate at the AI Firewall layer and are independent of the selected OpenAI-compatible upstream provider.

Provider compatibility still matters for:

  • OpenAI-compatible chat request and response shape;
  • model naming;
  • streaming behavior;
  • tool/function response formats;
  • embedding endpoint behavior when semantic cache is enabled.

AI Firewall supports both ordinary JSON chat completions and controlled OpenAI-compatible streaming. For stream=false or omitted-stream requests, the upstream only needs to return an OpenAI-compatible JSON chat completion. For stream=true, the configured upstream must additionally support OpenAI-compatible SSE.

AIF consumes provider SSE internally and assembles the complete canonical response before response controls and downstream replay. Provider streaming capability is therefore optional unless clients actually request streaming.

Evaluation Mode does not reduce upstream provider traffic: even a would-have cache hit still calls the live upstream provider. Capacity and rate-limit planning for an observe-mode pilot should therefore assume live request volume remains upstream-bound.

The current guard modules inspect string message content only. Non-string content arrays/objects can be preserved and proxied by AIF v0.8.2, but nested text parts inside those structures are not yet scanned, anonymized, or classified by the guard integrations.