Getting Started>Quickstart

Quickstart

Get Prism up and running locally in under 3 minutes. You only need Docker. No code changes are required in your application.

Prerequisites — Docker and Docker Compose installed on your machine. That's it.
1

Clone the repository

Clone the Prism repository and copy the environment template. The .env.example file documents every available option with sensible defaults.

bash
git clone https://github.com/your-username/prism.git
cd prism
cp .env.example .env
2

Configure credentials

Open .env and set the two required secrets. Everything else has a working default.

.env
# REQUIRED — choose a strong password for Postgres
POSTGRES_PASSWORD=changeme_strong_password
# REQUIRED — generate with: openssl rand -hex 32
JWT_SECRET=changeme_generate_with_openssl_rand_hex_32
# Your LLM provider key (forwarded on cache misses)
OPENAI_API_KEY=sk-proj-...
# Embedding model settings (defaults work out-of-the-box)
OLLAMA_URL=http://ollama:11434
EMBEDDING_MODEL=nomic-embed-text
EMBEDDING_PROVIDER=ollama

If you set EMBEDDING_PROVIDER=openai instead of ollama, the Ollama container will not start and your OPENAI_API_KEY will be used for both LLM calls and embeddings.

3

Start the stack

A single command starts the Go proxy, Postgres (with pgvector), Redis, Ollama, and Grafana. Database migrations run automatically. On first boot, the init container pulls the embedding model, which takes about a minute.

bash
docker-compose up -d

Verify the proxy is live:

bash
curl http://localhost:8080/health
# → {"status":"ready"}
4

Send your first request

Make a standard OpenAI chat completion request, but point it at localhost:8080/proxy/openai. The first request calls the upstream API normally. The second — even if phrased slightly differently — will be served from cache in milliseconds.

bash
curl -X POST http://localhost:8080/proxy/openai \
-H "Content-Type: application/json" \
-H "Authorization: Bearer YOUR_OPENAI_KEY" \
-d '{
"model": "gpt-4o",
"messages": [
{"role": "user", "content": "Explain quantum computing in one sentence."}
]
}'

Every response includes cache metadata in the headers and body:

http
X-Cache-Hit: true
X-Cache: L1
json
{
"choices": [{ "message": { "content": "..." } }],
"x_cache_metadata": {
"hit": true,
"source": "L1",
"latency_ms": 2
}
}

You're live!

Prism is now intercepting your LLM requests. Open the Grafana dashboard at http://localhost:4000 to see real-time hit rates and cost savings.

Default credentials: admin / GRAFANA_ADMIN_PASSWORD from .env