Quickstart
Get Prism up and running locally in under 3 minutes. You only need Docker. No code changes are required in your application.
Clone the repository
Clone the Prism repository and copy the environment template. The .env.example file documents every available option with sensible defaults.
git clone https://github.com/your-username/prism.gitcd prismcp .env.example .env
Configure credentials
Open .env and set the two required secrets. Everything else has a working default.
# REQUIRED — choose a strong password for PostgresPOSTGRES_PASSWORD=changeme_strong_password# REQUIRED — generate with: openssl rand -hex 32JWT_SECRET=changeme_generate_with_openssl_rand_hex_32# Your LLM provider key (forwarded on cache misses)OPENAI_API_KEY=sk-proj-...# Embedding model settings (defaults work out-of-the-box)OLLAMA_URL=http://ollama:11434EMBEDDING_MODEL=nomic-embed-textEMBEDDING_PROVIDER=ollama
If you set EMBEDDING_PROVIDER=openai instead of ollama, the Ollama container will not start and your OPENAI_API_KEY will be used for both LLM calls and embeddings.
Start the stack
A single command starts the Go proxy, Postgres (with pgvector), Redis, Ollama, and Grafana. Database migrations run automatically. On first boot, the init container pulls the embedding model, which takes about a minute.
docker-compose up -d
Verify the proxy is live:
curl http://localhost:8080/health# → {"status":"ready"}
Send your first request
Make a standard OpenAI chat completion request, but point it at localhost:8080/proxy/openai. The first request calls the upstream API normally. The second — even if phrased slightly differently — will be served from cache in milliseconds.
curl -X POST http://localhost:8080/proxy/openai \-H "Content-Type: application/json" \-H "Authorization: Bearer YOUR_OPENAI_KEY" \-d '{"model": "gpt-4o","messages": [{"role": "user", "content": "Explain quantum computing in one sentence."}]}'
Every response includes cache metadata in the headers and body:
X-Cache-Hit: trueX-Cache: L1
{"choices": [{ "message": { "content": "..." } }],"x_cache_metadata": {"hit": true,"source": "L1","latency_ms": 2}}
You're live!
Prism is now intercepting your LLM requests. Open the Grafana dashboard at http://localhost:4000 to see real-time hit rates and cost savings.