LLM proxy · a binary, not an app

One proxy.
The whole cache.

Intercept every LLM call. Serve hits from in-memory, Redis, and vector DB before they ever reach the API. No code changes in your app.

Install via
$docker-compose up -d

Linux · macOS · Docker · configuration guide

proxy.log
~/projects/prism $ docker-compose up -d
[+] Running 4/4
✔ Container redis Started
✔ Container postgres Started
✔ Container ollama Started
✔ Container prism Started

~/projects/prism $ curl http://localhost:8080/health
{"status":"ready"}

02:30:58INFOrouted → openai
02:30:59CACHEsaved to L1 + L2a
02:31:04SEMANTIChit L2b · score 0.94
02:31:04HITserved in 12ms · saved $0.002

# architecture

Built for scale. Tiered for speed.

Prism routes requests through a highly optimized multi-tier caching architecture. Deploy it on a tiny VPS or scale it across a Kubernetes cluster.

L1

In-Memory

Lightning fast exact-match caching. Stores recent requests directly in RAM for sub-millisecond response times.

Latency: < 1ms

L2a

Distributed Redis

Shared exact-match cache across multiple Prism instances. Perfect for team environments and scalable deployments.

Latency: ~5ms

L2b

Semantic Vector DB

Fuzzy matching via pgvector and local embeddings. Catch variations in prompts without hitting the API.

Latency: ~50ms


# why different

Stop paying for identical tokens.

Most teams use raw OpenAI calls or basic Redis caches. Prism understands the semantic meaning of prompts to maximize cache hits.

not an sdk wrapper

Prism is a proxy. No need to rewrite your application code or change SDKs.

not a cloud service

Keep your data on your servers. No third-party data collection or per-request fees.

more than exact match

Catch synonymous queries with local embedding models (Ollama / HuggingFace).

CapabilityRaw OpenAIBasic RedisPrism
Zero code changes——✓
Exact match caching—✓✓
Semantic matching——✓
Local embeddings——✓
Cost per cached request$0.01+$0$0

# extensibility

Make it yours.

Prism is built to be extended. Swap the default pgvector store for Pinecone or Qdrant. Use your own local embedding models via Ollama to ensure complete data privacy. Write custom eviction policies or rate limiters with simple plugins.

shell
# Add a custom embedding model
❯prism model add nomic-embed-text
Pulling model from ollama registry...
Updating vector dimensions to 768...
Reindexing L2b cache in background...
✓ Model ready