Health, Startup, and Readiness
AI Cost Firewall v0.8.2 exposes separate liveness, startup, and traffic-readiness endpoints. They intentionally answer different operational questions.
/healthz
curl -i http://localhost:8080/healthz
Returns 200 OK when the process is alive. It does not prove that cache backends were initialized or that the instance should receive application traffic.
/startupz
curl -i http://localhost:8080/startupz
/startupz is intended for Kubernetes/OpenShift startupProbe use. It is stricter than normal Observe-mode readiness for cache backends that are explicitly configured as readiness-required:
- when exact cache is enabled and
readiness_requires_redis true;, Redis must have initialized successfully in the current process; - when semantic cache is enabled and
readiness_requires_qdrant true;, Qdrant must have initialized successfully in the current process.
This closes an important startup race. If AIF starts while a fail-open/Observe cache dependency is unavailable, it may initialize a no-op cache implementation for that process lifetime. A failing startup probe lets the orchestrator restart the pod after the dependency becomes available so the real cache client is created.
A dependency outage that happens after successful initialization remains a runtime failure/recovery case and follows the normal fail-open, readiness, and Evaluation Mode behavior.
/readyz
curl -i http://localhost:8080/readyz
Returns 200 OK during normal operation and 503 during shutdown or when a dependency configured as readiness-required is unavailable.
| State | /healthz | /startupz | /readyz |
|---|---|---|---|
| Normal operation | 200 | 200 | 200 |
| Required cache backend failed to initialize | 200 | 503 | policy-dependent |
| Required runtime readiness dependency unavailable | 200 | usually 200 if it initialized earlier | 503 |
| Graceful shutdown | 200 | 503 | 503 |
| Process crashed | unavailable | unavailable | unavailable |
Readiness dependency policy is configured independently from cache fail-open behavior:
readiness_requires_redis true;
readiness_requires_qdrant false;
readiness_requires_upstream false;
For example, exact_cache_fail_open true; can allow an individual request to continue when Redis fails while readiness_requires_redis true; still removes the instance from readiness until Redis recovers in normal enforce mode.
Observe-mode exception
When aif_enforcement_mode observe; is active, Redis/Qdrant/embedding paths used only for evaluation are deliberately non-blocking for /readyz. Their runtime failure is recorded through evaluation error telemetry and does not interrupt the live upstream request.
/startupz is different: if an enabled Redis or Qdrant cache is marked readiness-required, that backend still has to initialize successfully in the current process. This is specifically intended to prevent a pod from remaining permanently on a startup no-op cache.
/version
curl -s http://localhost:8080/version
Returns release and compatibility metadata for the running binary. This is useful during pilot deployments, Docker image validation, and support checks.
Example fields include:
{
"version": "0.8.2",
"supported_api_style": "openai_compatible",
"provider_specific_config_blocks": false,
"aif_enforcement_mode": "observe",
"effective_cache_scope": "evaluation",
"effective_qdrant_collection": "aif_semantic_cache_eval"
}