Advanced>Observability

Observability

Prism is designed to be highly transparent. It ships with built-in Prometheus metrics, structured logging, and a pre-configured Grafana dashboard so you can immediately see your hit rates, latency, and cost savings.

Grafana Dashboard

If you deployed using the provided docker-compose.yml, Grafana is already running alongside Prism.

Hit Rates

Track exact-match and semantic hit rates over time, broken down by layer.

Latency

Monitor p50, p90, and p99 response times for both cache hits and upstream API calls.

Cost Savings

Calculate estimated USD saved based on cached token counts and model pricing.

Accessing Grafana

  • Open http://localhost:4000 in your browser.
  • Login with username admin
  • Password is your GRAFANA_ADMIN_PASSWORD from .env
  • Go to Dashboards > Prism Overview

Prometheus Metrics

Prism exposes a /metrics endpoint that can be scraped by any Prometheus instance.

/metrics
# HELP cache_hits_total Total number of cache hits
# TYPE cache_hits_total counter
cache_hits_total{layer="L1",tenant="default"} 1542
cache_hits_total{layer="L2a",tenant="default"} 314
cache_hits_total{layer="L2b",tenant="default"} 89
# HELP cache_misses_total Total number of cache misses
# TYPE cache_misses_total counter
cache_misses_total{tenant="default"} 402
# HELP upstream_errors_total Total upstream API errors
# TYPE upstream_errors_total counter
upstream_errors_total{provider="openai"} 2
# HELP request_latency_histogram_ms Request latency in milliseconds
# TYPE request_latency_histogram_ms histogram
...

Structured Logging

All logs are emitted as structured JSON to stdout, making them easy to ingest into Datadog, ELK, or CloudWatch. You can control the verbosity using the LOG_LEVEL environment variable.

json
{"level":"info","time":"2023-11-20T10:14:02Z","msg":"cache hit","tenant":"team-alpha","layer":"L2b","latency_ms":42,"similarity":0.93}
{"level":"warn","time":"2023-11-20T10:14:45Z","msg":"upstream rate limit reached","provider":"openai","retry_after_ms":1000}

Cost Savings API

You can programmatically fetch cost savings analytics via a dedicated endpoint, which is useful for building custom internal dashboards or billing reports.

bash
curl http://localhost:8080/analytics/cost-savings
json
{
"tenants": [
{
"id": "team-alpha",
"cached_requests": 14823,
"tokens_saved": 2940000,
"estimated_usd_saved": 88.20
}
],
"total_usd_saved": 88.20
}