ConceptsZero-LLM features

Zero-LLM features

Most memory systems make agents more expensive to run. They add an extra LLM call for extraction, another for conflict resolution, another for reasoning. Mnemosyne ships four features that work the other way: they save tokens at zero LLM cost because they’re built on SQL, materialised views and rule engines.

Together they make the headline claim true: the only memory layer with a positive ROI multiplier.

1. Trust-anchor consensus

When two facts conflict, most systems ask an LLM to judge which one is right. Mnemosyne uses a deterministic trust ladder:

codebase_verified  → 100   (verified against actual code/config)
user_stated        →  90   (user explicitly said this)
user_confirmed     →  85   (user confirmed when asked)
api_response       →  80   (from an authoritative API)
document_extracted →  70   (from official docs)
agent_observed     →  60   (agent saw it happen)
agent_inferred     →  40   (agent concluded from reasoning)
llm_generated      →  30   (LLM extraction without confirmation)
federation_received → 20   (from a federated peer, untrusted by default)

Higher source wins. Same source → more recent wins. Same source and recency → higher Memory Worth wins. Truly ambiguous cases get queued for human review.

Savings: ~750 tokens per conflict. Conflicts are rare per session (~10% hit one), so this averages ~75 tokens/session — the figure carried into the aggregate below.

2. Composition engine

If fact A says “service X has a 3s timeout” and fact B says “service Y takes 5s to respond” and the graph says X calls Y, Mnemosyne can derive without an LLM:

“⚠️ Service X may time out when calling Y (3s < 5s).”

The engine runs declarative rules over recall results. Built-in rules include transitive dependency, timeout conflict, supersession chain, and shared dependency risk. Workspaces can register custom rules using the nine locked relation verbs.

Savings: ~100 tokens per inference, ~10 inferences/session.

3. Meta-cognition

The agent should know what it knows. Mnemosyne computes a per-entity confidence map as a materialised view:

coverage_level: expert | familiar | aware | minimal | blind
freshness:     current | recent  | aging  | stale

When the agent is about to recall against coverage=blind, freshness=stale, it can short-circuit and ask the user instead of burning tokens on a guaranteed-empty recall.

Savings: ~200 tokens per avoided empty recall, plus the “let me check… hmm I’m not sure” loop disappears entirely.

4. Predictive anticipation

Patterns in the event stream tell you what’s coming next:

  • Sequential: “after deploy.success, billing.spike follows within 24h 85% of the time (23 observations).”
  • Periodic: “Monday 9–11am, 80% of recalls are about weekly metrics.”
  • Conditional: “when entity production is mentioned, fact deploy-runbook is recalled 90% of the time.”
  • Co-occurrence: “Redis config + Docker compose + env vars — always together.”

The mining runs as pure SQL over mnemo_event and the recall telemetry. At session start the engine pre-injects the predicted context before the agent asks. Saves 1–3 recall round-trips per session — both their latency and their tokens.

Savings: ~300–600 tokens per session, plus 200–600ms of round-trip latency avoided.

The aggregate

Per session:
  Consensus       ~  75 tokens (~10% of conflicts)
  Composition     ~1000 tokens (10 inferences × 100)
  Meta-cognition  ~ 200 tokens
  Anticipation    ~ 450 tokens (avg)
  ────────────────────────────
  Total saved     ~1725 tokens / session

At typical token prices (~$0.003/1k for Sonnet input) that’s $0.005 saved per session — at zero extra cost. Mnemosyne pays for itself somewhere around five sessions per agent per month.

💡

These four features are the moat. No other memory system on the market (Letta, Zep, Mem0, Cognee, LangMem, Engram) ships any of them. They are what turn “yet another memory store” into “the layer that makes agents cheaper to run”.