ConceptsThree-tier retrieval

Three-tier retrieval

A naive recall hits the database every time and pays the full cost. A production-grade recall is a cascade: cheap and fast tiers serve the common case, expensive tiers handle the long tail, and every hit promotes itself upward so the next request is even faster.

Query
  │
  ▼
┌────────────────────────────────────────────────┐
│ Tier 1 — Hot (Redis ACT-R activation set)      │  < 5 ms
│  - Sorted set of "currently active" memories   │
│  - LRU by activation score                     │
│  - Sufficient when ≥ 3 hits above threshold    │
└────────────────────────────────────────────────┘
  │ fall through if insufficient
  ▼
┌────────────────────────────────────────────────┐
│ Tier 2 — Warm (Pointer index + LSH)            │  < 15 ms
│  - Entity drawer narrows the candidate set     │
│  - LSH on first 64 dims for O(1) shortlisting  │
│  - Full cosine verification on candidates only │
│  - Promotes hits → hot tier                    │
└────────────────────────────────────────────────┘
  │ fall through if insufficient
  ▼
┌────────────────────────────────────────────────┐
│ Tier 3 — Cold (Full 7-stage pipeline)          │  < 200 ms
│  - HyDE-style query prep                       │
│  - Hybrid vector + FTS retrieval               │
│  - Cross-encoder rerank (local or API)         │
│  - Trust ladder + trust decay                  │
│  - Graph expansion (1-hop)                     │
│  - Composition engine inferences appended      │
│  - Always promotes top 3 → hot tier            │
└────────────────────────────────────────────────┘
  │
  ▼
Post-processing
  - Deduplicate (cosine > 0.88 within result)
  - Memory Worth filter (exclude worth < 0.3)
  - Token-budget-adaptive truncation
  - Async: update ACT-R, Hebbian co-recall

Token-budget-adaptive recall

The pipeline reads the remaining context budget of the calling agent and adapts:

Budget leftBehaviour
< 500 tokensSkip recall entirely. The agent is full.
< 2 000Tier 1 only. Top 3, compact render.
< 8 000Tier 1+2. Top 5, include episodes.
≥ 8 000Full pipeline. Top 10 + 1-hop graph expansion.

This avoids the pathological case of recall returning fifteen facts that push the prompt over the token limit.

Graceful degradation

Component downResult
Redis unavailableTier 1 disabled. Tier 2+3 still work, ~10ms slower.
Embedding API downVector stages skipped. Mode A (FTS-only) kicks in.
pgvector unavailableMode A. Pure FTS recall over text_lemmatized.
All down except PGMode A. Mnemosyne still answers, just simpler.

The cascade is what makes Mnemosyne survive partial outages without going into a bad state.

Shipped in v3-alpha (Beta). All three tiers are live: Tier 3’s pipeline is in packages/core/src/recall/search.ts, Tier 2’s pointer index plus LSH shortlisting, and the Tier 1 hot tier (the @mnemosyne/adapter-redis ACT-R activation set, Stable). Interfaces may still change before 3.0 stable.