Getting Started>Architecture
Architecture
Prism is a stateless Go binary that sits between your application and any LLM provider. Every request flows through a four-layer cache waterfall before ever hitting the upstream API.
Request Flow
Your Application
OpenAI SDK, curl, or any HTTP client
L0
Intent Normalizer
Strips typos and phrasing variations. What's 2+2? → what is 2+2
L1
In-Memory LRU
Served directly from process memory.
L2a
Redis Cache
Shared exact-match cache across instances.
L2b
Semantic Vector Search
Finds similar concepts via pgvector.
Upstream API
OpenAI / Anthropic / Groq (on complete miss)
Layer Details
L0
Intent Normalizer
Before any cache lookup, the incoming prompt is normalized: whitespace is collapsed, common typos are corrected, and filler words stripped. This dramatically increases hit rates for logically identical queries.
L1
In-Memory LRU
A custom LRU cache backed by Go process memory. Max size is controlled by L1_MAX_BYTES (default 128MB). Entries are evicted when the byte limit is reached.
L2a
Redis (Exact-Match)
Redis acts as a shared cache across all instances. Critical when running multiple replicas — a cache warm on one instance is immediately visible to all.
L2b
Postgres (Semantic)
Uses pgvector and local embedding models to find semantically similar past queries. Matches queries that are phrased differently but mean the same thing.
Embedding models
Prism supports two embedding providers. The default is Ollama, which runs fully locally with no API key or per-token cost.
| Provider | Default Model | Dimensions | Cost | Privacy |
|---|---|---|---|---|
| ollama | nomic-embed-text | 768 | Free | Fully local |
| openai | text-embedding-3-small | 1536 | $0.02/1M tokens | Sends to OpenAI |
Tech stack
EngineGo 1.22+ (single binary, ~15MB image)
L1 CacheCustom in-process LRU (Go)
L2a CacheRedis 7.2
L2b CachePostgreSQL 16 + pgvector
EmbeddingsOllama (nomic-embed-text) — local
ObservabilityPrometheus + Grafana
Useful commands
bash
# View proxy logs in real-timedocker-compose logs -f cache-proxy# Restart just the proxy (after config changes)docker-compose restart cache-proxy# Check raw cache hit statscurl http://localhost:8080/metrics | grep cache_hits# Wipe all data and start freshdocker-compose down -v && docker-compose up -d# Rebuild the proxy binary after code changesdocker-compose up -d --build cache-proxy