API Reference
All endpoints exposed by the Prism proxy server. By default the server binds to :8080.
Base URL
http://localhost:8080
Common request headers
| Header | Required | Description |
|---|---|---|
| Authorization | Yes | Your LLM provider API key in Bearer sk-... format. Forwarded to the upstream provider on cache miss. |
| Content-Type | Yes | Must be application/json. |
| X-Tenant-ID | No | Namespace for cache isolation. Defaults to default if omitted. See Multi-Tenant. |
Response headers
| Header | Values | Description |
|---|---|---|
| X-Cache-Hit | true / false | Whether the response was served from cache. |
| X-Cache | L1 / L2a / L2b / backend | Which layer served the request. |
Endpoints
/proxy/openaiForwards to api.openai.com/v1/chat/completions. Accepts the full OpenAI request schema including streaming, function calling, and vision.
/proxy/anthropicForwards to api.anthropic.com/v1/messages. Supports all Claude models and tool use.
/proxy/groqForwards to api.groq.com/openai/v1/chat/completions. Uses the OpenAI-compatible schema.
/proxy/togetherForwards to api.together.xyz/v1/chat/completions. Uses the OpenAI-compatible schema.
/healthReturns {"status":"ready"} when the proxy is fully initialised (DB migrations run, embedding model loaded).
/metricsExposes cache_hits_total, cache_misses_total, request_latency_histogram, and upstream_errors_total labelled by layer and tenant.
/analytics/cost-savingsReturns per-tenant saved token counts and estimated USD savings since last reset.
/cache/queryLegacy endpoint for querying the cache directly without proxying to an upstream API. Useful for tooling and debugging.
Health check example
curl http://localhost:8080/health# → {"status":"ready"}
Cost savings response
GET /analytics/cost-savings{"tenants": [{"id": "team-alpha","cached_requests": 14823,"tokens_saved": 2940000,"estimated_usd_saved": 88.20},{"id": "default","cached_requests": 3201,"tokens_saved": 640200,"estimated_usd_saved": 19.21}],"total_usd_saved": 107.41}
Prometheus metrics
# Fetch all metricscurl http://localhost:8080/metrics# Filter for cache statscurl -s http://localhost:8080/metrics | grep cache_hits# cache_hits_total{layer="L1",tenant="team-alpha"} 9241# cache_hits_total{layer="L2a",tenant="team-alpha"} 3991# cache_hits_total{layer="L2b",tenant="team-alpha"} 1591