Usage>API Reference

API Reference

All endpoints exposed by the Prism proxy server. By default the server binds to :8080.

Base URL

text
http://localhost:8080

Common request headers

HeaderRequiredDescription
AuthorizationYesYour LLM provider API key in Bearer sk-... format. Forwarded to the upstream provider on cache miss.
Content-TypeYesMust be application/json.
X-Tenant-IDNoNamespace for cache isolation. Defaults to default if omitted. See Multi-Tenant.

Response headers

HeaderValuesDescription
X-Cache-Hittrue / falseWhether the response was served from cache.
X-CacheL1 / L2a / L2b / backendWhich layer served the request.

Endpoints

POST/proxy/openai

Forwards to api.openai.com/v1/chat/completions. Accepts the full OpenAI request schema including streaming, function calling, and vision.

POST/proxy/anthropic

Forwards to api.anthropic.com/v1/messages. Supports all Claude models and tool use.

POST/proxy/groq

Forwards to api.groq.com/openai/v1/chat/completions. Uses the OpenAI-compatible schema.

POST/proxy/together

Forwards to api.together.xyz/v1/chat/completions. Uses the OpenAI-compatible schema.

GET/health

Returns {"status":"ready"} when the proxy is fully initialised (DB migrations run, embedding model loaded).

GET/metrics

Exposes cache_hits_total, cache_misses_total, request_latency_histogram, and upstream_errors_total labelled by layer and tenant.

GET/analytics/cost-savings

Returns per-tenant saved token counts and estimated USD savings since last reset.

POST/cache/query

Legacy endpoint for querying the cache directly without proxying to an upstream API. Useful for tooling and debugging.

Health check example

bash
curl http://localhost:8080/health
# → {"status":"ready"}

Cost savings response

json
GET /analytics/cost-savings
{
"tenants": [
{
"id": "team-alpha",
"cached_requests": 14823,
"tokens_saved": 2940000,
"estimated_usd_saved": 88.20
},
{
"id": "default",
"cached_requests": 3201,
"tokens_saved": 640200,
"estimated_usd_saved": 19.21
}
],
"total_usd_saved": 107.41
}

Prometheus metrics

bash
# Fetch all metrics
curl http://localhost:8080/metrics
# Filter for cache stats
curl -s http://localhost:8080/metrics | grep cache_hits
# cache_hits_total{layer="L1",tenant="team-alpha"} 9241
# cache_hits_total{layer="L2a",tenant="team-alpha"} 3991
# cache_hits_total{layer="L2b",tenant="team-alpha"} 1591