OperationsMonitoring

Monitoring

Mnemosyne is observable but never demands an observability stack to run.

Built-in health endpoints

Several health checks for different consumers:

# Liveness — fast, no DB
curl http://localhost:3000/healthz
# 200 OK if the process is up
 
curl http://localhost:3000/health/live
# 200 OK if the process is up
 
# Readiness probe — checks dependencies
curl http://localhost:3000/health/ready
# 200 OK if PG (and Redis, if configured) are reachable

Per-workspace health snapshot

curl http://localhost:3000/v1/health \
  -H "Authorization: Bearer mns_live_xxx"

Returns the vitals (live facts, pinned count, embedded ratio, recall hit rate, last write timestamp) that the Brain Inspector UI consumes.

Prometheus metrics

curl http://localhost:3000/metrics

Standard Prometheus exposition format. Key metrics:

# Request rate / errors / duration (RED)
mnemo_requests_total{method, endpoint, status}
mnemo_errors_total{method, endpoint, type}
mnemo_request_duration_seconds{method, endpoint}

# Per-tier recall latency
mnemo_recall_duration_seconds{tier}
mnemo_recall_tier_hits_total{tier}

# Cognitive
mnemo_facts_total{workspace, type, status}
mnemo_worth_distribution{bucket}
mnemo_hot_tier_size{workspace}

# Cost
mnemo_llm_tokens_used_total{model, operation}
mnemo_tokens_saved_total{mechanism}

OpenTelemetry

Set OTEL_EXPORTER_OTLP_ENDPOINT and every request gets a full distributed trace:

mnemosyne:
  environment:
    OTEL_EXPORTER_OTLP_ENDPOINT: http://otel-collector:4317
    OTEL_SERVICE_NAME: mnemosyne

Mnemosyne emits domain-specific spans on top of the auto-instrumentation:

  • mnemo.recall (parent) with child spans per tier
  • mnemo.write with conflict-detection child span
  • mnemo.embed with the API call inside
  • mnemo.consolidation.gate with each condition checked

Dashboards

We ship Grafana dashboards in deploy/grafana/dashboards/:

  • Overview — RED metrics, error budget burn
  • Recall — tier hit distribution, latency p50/p95/p99
  • Cognitive — worth distribution, embedding coverage, drift
  • Cost — token spend, tokens saved, ROI multiplier

Import them with grafana-cli dashboard import or copy-paste the JSON.

Logs and traces

Alongside Prometheus and Grafana, Mnemosyne bundles a full observability stack: Loki + Promtail for log aggregation and Tempo for distributed traces. Point your OTLP exporter at Tempo and your Grafana explore views resolve metrics, logs, and traces in one place.

No external account required. Prometheus, Grafana, Loki, Promtail, and Tempo run wherever you run Mnemosyne. Self-host the whole observability stack if you want.