Getting Started>Architecture

Architecture

Prism is a stateless Go binary that sits between your application and any LLM provider. Every request flows through a four-layer cache waterfall before ever hitting the upstream API.

Request Flow

Your Application

OpenAI SDK, curl, or any HTTP client

L0

Intent Normalizer

Strips typos and phrasing variations. What's 2+2? → what is 2+2

L1

In-Memory LRU

Served directly from process memory.

< 1ms
L2a

Redis Cache

Shared exact-match cache across instances.

~5ms
L2b

Semantic Vector Search

Finds similar concepts via pgvector.

~50ms

Upstream API

OpenAI / Anthropic / Groq (on complete miss)

Layer Details

L0

Intent Normalizer

Before any cache lookup, the incoming prompt is normalized: whitespace is collapsed, common typos are corrected, and filler words stripped. This dramatically increases hit rates for logically identical queries.
L1

In-Memory LRU

A custom LRU cache backed by Go process memory. Max size is controlled by L1_MAX_BYTES (default 128MB). Entries are evicted when the byte limit is reached.
L2a

Redis (Exact-Match)

Redis acts as a shared cache across all instances. Critical when running multiple replicas — a cache warm on one instance is immediately visible to all.
L2b

Postgres (Semantic)

Uses pgvector and local embedding models to find semantically similar past queries. Matches queries that are phrased differently but mean the same thing.

Embedding models

Prism supports two embedding providers. The default is Ollama, which runs fully locally with no API key or per-token cost.

ProviderDefault ModelDimensionsCostPrivacy
ollamanomic-embed-text768FreeFully local
openaitext-embedding-3-small1536$0.02/1M tokensSends to OpenAI

Tech stack

EngineGo 1.22+ (single binary, ~15MB image)
L1 CacheCustom in-process LRU (Go)
L2a CacheRedis 7.2
L2b CachePostgreSQL 16 + pgvector
EmbeddingsOllama (nomic-embed-text) — local
ObservabilityPrometheus + Grafana

Useful commands

bash
# View proxy logs in real-time
docker-compose logs -f cache-proxy
# Restart just the proxy (after config changes)
docker-compose restart cache-proxy
# Check raw cache hit stats
curl http://localhost:8080/metrics | grep cache_hits
# Wipe all data and start fresh
docker-compose down -v && docker-compose up -d
# Rebuild the proxy binary after code changes
docker-compose up -d --build cache-proxy