InternalsRecall pipeline

Recall pipeline

These seven stages are the Tier-3 (cold) pipeline — the deepest path in the three-tier retrieval cascade. Hot (L0/L1) and warm tiers answer most queries first; a query only falls through to this full pipeline when the faster tiers miss. Within Tier 3 a query flows through: query prep, embedding, hybrid retrieval (vector + FTS + pointer), rerank, trust decay, graph expansion, render.

Source of truth

Everything documented here is implemented in packages/core/src/recall/search.ts. 97 KB of pipeline. The headings below match its function names so the code is navigable.

The seven stages

1. Query preparation (query-prep.ts)

  • Normalisation (case, accent, whitespace).
  • HyDE-style expansion: synthesise hypothetical answers, embed those alongside the original query to widen recall.
  • Pointer extraction: extract entity mentions, decision topic keys, episode IDs from the query.

2. Embedding (embed.ts)

  • Workspace-keyed LRU cache (lru-cache) so repeated queries don’t re-embed.
  • Cached by (workspace, model, sha256(text)) — so the same query text on the same model is a free lookup.
  • Provider is injected — Mnemosyne never picks an embedder by default. Pass any function matching EmbedFn.

3. Hybrid retrieval

Three sub-queries run in parallel:

  • Vector (HNSW cosine over halfvec(1536))
  • FTS (PostgreSQL tsvector over text_lemmatized)
  • Pointer (entity drawer membership)

Results are merged with weights from trust-ladder.ts.

4. Rerank (rerank.ts)

Either a local cross-encoder or an external API depending on config. The rerank is gated — if the initial merge is already high-confidence (top scores cluster tightly, no contradictions), Mnemosyne skips the rerank to save latency and tokens.

5. Trust decay (trust-decay.ts)

Per-source half-lives:

  • codebase_verified decays slowly (high source reliability)
  • llm_generated decays fast (low source reliability)
  • user_stated is pinned to its creation source

This handles the “we believed X two months ago” problem — recent facts from low-trust sources don’t drown out older facts from high-trust sources.

6. Graph expansion

1-hop traversal over mnemo_relation for the top-K hits. If a fact about “the Stripe webhook” is returned, the graph expansion adds related facts about “Stripe”, “webhook secret”, “webhook retry policy” when they exist.

Triggered by triggering.ts — only when the initial pipeline yielded < 3 high-confidence hits or the query explicitly asks for graph context.

7. Render (render.ts)

  • Compact format for token-budget constrained agents.
  • Verbose format for debugging.
  • JSON format for SDKs (default).

Telemetry

Every call writes to mnemo_query_cache (which serves the predictive anticipation engine) and emits a span with the tier hit, latency per stage, and the result count. See Monitoring for what gets exported.

Mode A fallback

If no embedder is configured, stages 2 and 3a (vector) are skipped. Stages 3b (FTS) and 4–7 run unchanged. Results are still useful — just less semantic.