LLM proxy · a binary, not an app
One proxy.
The whole cache.
Intercept every LLM call. Serve hits from in-memory, Redis, and vector DB
before they ever reach the API. No code changes in your app.
# architecture
Built for scale. Tiered for speed.
Prism routes requests through a highly optimized multi-tier caching architecture. Deploy it on a tiny VPS or scale it across a Kubernetes cluster.
In-Memory
Lightning fast exact-match caching. Stores recent requests directly in RAM for sub-millisecond response times.
Latency: < 1ms
Distributed Redis
Shared exact-match cache across multiple Prism instances. Perfect for team environments and scalable deployments.
Latency: ~5ms
Semantic Vector DB
Fuzzy matching via pgvector and local embeddings. Catch variations in prompts without hitting the API.
Latency: ~50ms
# why different
Stop paying for identical tokens.
Most teams use raw OpenAI calls or basic Redis caches. Prism understands the semantic meaning of prompts to maximize cache hits.
Prism is a proxy. No need to rewrite your application code or change SDKs.
Keep your data on your servers. No third-party data collection or per-request fees.
Catch synonymous queries with local embedding models (Ollama / HuggingFace).
| Capability | Raw OpenAI | Basic Redis | Prism |
|---|---|---|---|
| Zero code changes | — | — | ✓ |
| Exact match caching | — | ✓ | ✓ |
| Semantic matching | — | — | ✓ |
| Local embeddings | — | — | ✓ |
| Cost per cached request | $0.01+ | $0 | $0 |
# extensibility
Make it yours.
Prism is built to be extended. Swap the default pgvector store for Pinecone or Qdrant. Use your own local embedding models via Ollama to ensure complete data privacy. Write custom eviction policies or rate limiters with simple plugins.