Integration
Prism is a drop-in proxy. Change one line in your application — the base URL — and everything else works exactly as before. No new SDK, no new dependency, no refactor.
Provider endpoints
Prism exposes a separate path for each supported LLM provider. Route your traffic to the appropriate endpoint:
| Provider | Prism Endpoint |
|---|---|
| OpenAI | http://localhost:8080/proxy/openai |
| Anthropic | http://localhost:8080/proxy/anthropic |
| Groq | http://localhost:8080/proxy/groq |
| Together AI | http://localhost:8080/proxy/together |
Python — OpenAI SDK
Pass base_url when constructing the client. Every subsequent call — completions, streaming, function calling — works unchanged.
from openai import OpenAI# Before: direct OpenAI# client = OpenAI(api_key="sk-your-key")# After: route through Prism (one line change)client = OpenAI(api_key="sk-your-key",base_url="http://localhost:8080/proxy/openai")# Your existing code works unchangedresponse = client.chat.completions.create(model="gpt-4o",messages=[{"role": "user", "content": "What is machine learning?"}])print(response.choices[0].message.content)
Python — Anthropic SDK
from anthropic import Anthropic# Before# client = Anthropic(api_key="sk-ant-your-key")# Afterclient = Anthropic(api_key="sk-ant-your-key",base_url="http://localhost:8080/proxy/anthropic")message = client.messages.create(model="claude-3-5-sonnet-20241022",max_tokens=1024,messages=[{"role": "user", "content": "Explain transformers."}])print(message.content)
Node.js / TypeScript — OpenAI SDK
import OpenAI from "openai";const client = new OpenAI({apiKey: process.env.OPENAI_API_KEY,baseURL: "http://localhost:8080/proxy/openai", // ← only change});const response = await client.chat.completions.create({model: "gpt-4o",messages: [{ role: "user", content: "Summarize the Rust ownership model." }],});console.log(response.choices[0].message.content);
Go
package mainimport ("context"openai "github.com/sashabaranov/go-openai")func main() {cfg := openai.DefaultConfig("sk-your-key")cfg.BaseURL = "http://localhost:8080/proxy/openai" // ← only changeclient := openai.NewClientWithConfig(cfg)resp, err := client.CreateChatCompletion(context.Background(),openai.ChatCompletionRequest{Model: openai.GPT4o,Messages: []openai.ChatCompletionMessage{{Role: openai.ChatMessageRoleUser, Content: "What is Go?"},},},)// ...}
curl
Works with any HTTP client. The request body is forwarded verbatim to the upstream provider, so all model parameters, system prompts, and tools are supported.
curl -X POST http://localhost:8080/proxy/openai \-H "Authorization: Bearer sk-your-key" \-H "Content-Type: application/json" \-d '{"model": "gpt-4o","messages": [{"role": "system", "content": "You are a helpful assistant."},{"role": "user", "content": "Explain REST APIs"}]}'
Multi-tenant isolation
Add an X-Tenant-ID header to isolate cache namespaces. Each tenant sees only their own cached results. If omitted, the request is assigned to the default tenant.
# Team A's requests go into the "team-alpha" namespacecurl -X POST http://localhost:8080/proxy/openai \-H "Authorization: Bearer sk-your-key" \-H "X-Tenant-ID: team-alpha" \-H "Content-Type: application/json" \-d '{"model":"gpt-4o","messages":[...]}'# Team B cannot see Team A's cached answerscurl -X POST http://localhost:8080/proxy/openai \-H "Authorization: Bearer sk-your-key" \-H "X-Tenant-ID: team-beta" \...
Verifying cache hits
Every Prism response includes cache metadata. Check the X-Cache-Hit and X-Cache response headers, or the x_cache_metadata field in the response body.
Response headers
X-Cache-Hit: trueX-Cache: L1
Response body
{"x_cache_metadata": {"hit": true,"source": "L2b","latency_ms": 47,"similarity_score": 0.94}}