Multi-Tenant Isolation
Deploy a single Prism instance to serve multiple teams, users, or applications. Cache namespaces are strictly isolated to prevent data leakage across boundaries.
How it works
Prism namespaces all cached responses (across L1, L2a, and L2b) using a tenant identifier. When a request comes in, Prism looks for the X-Tenant-ID header. If found, the request will only retrieve cache hits generated by that specific tenant, and any new backend responses will be saved strictly within that tenant's namespace.
If the X-Tenant-ID header is omitted, the request is assigned to the default tenant.
Using Tenant Headers
Team A requesting a completion
curl -X POST http://localhost:8080/proxy/openai \-H "Authorization: Bearer sk-your-key" \-H "X-Tenant-ID: team-alpha" \-H "Content-Type: application/json" \-d '{"model":"gpt-4o","messages":[{"role":"user","content":"Explain quantum mechanics"}]}'
Team B making the exact same request
Even though the prompt is identical, Team B will not hit Team A's cache. They will trigger a new upstream API call, which is then cached in the team-beta namespace.
curl -X POST http://localhost:8080/proxy/openai \-H "Authorization: Bearer sk-your-key" \-H "X-Tenant-ID: team-beta" \...
Per-Tenant Rate Limits
Rate limits are enforced at the tenant level. You can configure the global tenant rate limit via theRATE_LIMIT_RPM environment variable.
If RATE_LIMIT_RPM=1000, then team-alpha can make 1,000 requests per minute, and team-beta can make 1,000 requests per minute completely independently. Exceeding the limit returns a 429 Too Many Requests response.
Metrics & Cost Tracking
All Prometheus metrics exposed by Prism include a tenant label. This makes it trivial to bill individual teams for their API usage or attribute cost savings accurately.
cache_hits_total{layer="L1",tenant="team-alpha"} 1054cache_hits_total{layer="L1",tenant="team-beta"} 23cache_misses_total{tenant="team-alpha"} 12cache_misses_total{tenant="team-beta"} 405