InternalsArchitecture

Mnemosyne is a cognitive memory layer for AI agents: it stores facts about a workspace, keeps them honest over time, and retrieves the relevant ones on demand. This document is the canonical system spec — the complete picture of what Mnemosyne is, what it guarantees, and how the moving parts fit together.

Audience: contributors, integrators evaluating Mnemosyne for production, and anyone wiring an agent runtime against it. Reader is assumed to know Postgres and a bit of LLM-app plumbing. Out of scope: tutorials — see the getting-started guide for that.


1. What Mnemosyne is (and is not)

Mnemosyne is the substrate between your agents and Postgres. Ten typed cognitive primitives — fact, decision, episode, entity, strategy, workflow, skill, task, event, and reasoning (plus hyperedge for n-ary relations) — sit behind a small, opinionated API:

recall()      → hybrid search: Postgres FTS + dense vectors + graph expansion + rerank
remember()    → write a fact with provenance, confidence, and valid-time
forget()      → bitemporal soft-invalidate — history is never destroyed
session()     → start a session, record turns, collapse it to typed memory

It is not a vector store with a nicer wrapper. The distinguishing claims:

  • Bitemporal by construction — every fact, relation, and decision carries valid_from / valid_to, so the system answers both “what is true now?” and “what did we believe last Tuesday?”.
  • Multi-tenant in the database, not in the WHERE clause — RLS + FORCE, per-transaction GUCs, role downgrade, integration-tested cross-tenant isolation. Fails closed when the workspace context is unset.
  • Provider-agnostic — bring your own LLM. Seven providers (OpenAI, Anthropic, Google, Cohere, Mistral, Voyage, Ollama) plug through the same LLMProvider interface, fronted by an LLMRouter (per-task routing) and a BudgetGuard (USD spend cap). The cognitive contract is independent of which embedder draws the vectors.
  • Three consumption modes — import as a library, run as an HTTP service, or expose as MCP tools for Claude Desktop / Cursor. Same model, same isolation guarantees, three deployment shapes.

2. The cognitive model: bitemporal facts

A fact in Mnemosyne is not a row you overwrite. It is a versioned assertion with two independent time axes:

  • Valid time (valid_from / valid_to) — when the fact is true in the world. A fact that became true on Monday and was superseded on Friday has valid_from = Monday, valid_to = Friday. An open-ended fact has valid_to = NULL.
  • Transaction time — when Mnemosyne recorded the fact (the row’s creation and its status transitions). This axis is append-only history: it never rewrites.

The mnemo_fact, mnemo_relation, and mnemo_decision tables all carry valid_from (NOT NULL DEFAULT now()) and a nullable valid_to, giving each the same bitemporal lifecycle.

Why two axes? Because the world changes and your knowledge of it changes, and those are different events. Bitemporality lets you correct a mistaken belief without erasing the fact that you once held it — essential for auditable, debuggable agent memory.

Each active fact also carries a memory_strength — a potentiation score with a neutral baseline of 1.0. Recall and decay nudge it over time so often-recalled facts surface more readily and stale ones fade. This is the Hebbian rule applied to retrieval: what fires together, recalls together.

Theory-of-mind discriminators

Beyond bitemporal time, every primitive carries three discriminators that let the agent reason about its own knowledge:

  • memory_type ∈ semantic / episodic / procedural / working — the classic cognitive-science cut. Procedural facts (“the API reset clock is the 1st of every month”) are weighted differently from working memory (“the user just said X”).
  • attribution ∈ user_stated / user_belief / objective_fact / inferred — who asserted this. Lets recall surface “the user thinks X” separately from “X is true”. Essential for honest agents.
  • protocol_version — the Memory Protocol revision the row was extracted under. Lets future extractors evolve without losing the ability to interpret older rows.

3. System topology

Mnemosyne is one Postgres database (plus a Redis hot tier) and a fleet of packages. The fleet ships as a pnpm monorepo so the cognitive contract (@mnemosyne/core) is shared across every consumption mode and the transport layers — REST, gRPC, WebSocket, and MCP — compose without duplication.

The packages

PackageRole
@mnemosyne/coreThe cognitive contract: the 10 primitives, recall pipeline, the zero-LLM engine (anticipation / composition / consensus), withMnemoTx, bitemporal queries, graph builders, PII detection. Framework-agnostic, no server-only.
@mnemosyne/typesShared type contract (PrimitiveKind, DTOs, errors) imported by every other package so the surface stays consistent.
@mnemosyne/serverA Hono-based HTTP service exposing /v1/* (REST), plus gRPC (ConnectRPC) and WebSocket on the same port. Owns cron scheduling, API-key auth, and the public surface. Run via docker compose up.
@mnemosyne/client-tsA typed TypeScript SDK over the v1 API. Zero runtime dependencies. The canonical way to integrate Mnemosyne from another Node app.
@mnemosyne/mcpA stdio-only MCP server exposing 26 memory tools (recall, remember, session, task, event, reasoning, self_edit, …). Binds no port. Plugs straight into Claude Desktop, Cursor, custom agents.
@mnemosyne/adapter-postgresThe Postgres adapter: schema, RLS policies, bitemporal queries, the Drizzle + postgres driver seam.
@mnemosyne/adapter-redisThe hot-tier adapter — ACT-R activation L0/L1 over Redis for sub-millisecond recall of the most-active memories.
@mnemosyne/adapter-vectorThe vector seam — HNSW search over halfvec(1536), pluggable so the embedding store can evolve.
@mnemosyne/adapter-llmThe BYO-LLM seam: a LLMProvider interface, 7 provider implementations (OpenAI, Anthropic, Google, Cohere, Mistral, Voyage, Ollama), the LLMRouter, and the BudgetGuard spend cap.
@mnemosyne/consolidationThe consolidation worker: ConsolidationGate + ReplayValidator + scheduler, wired at server boot.
@mnemosyne/federationMulti-instance peer sync: Ed25519 + mTLS, LWW / supermajority / consensus conflict resolution, 9 HTTP routes.
@mnemosyne/pluginsThe plugin runtime plus 9 builtin plugins. Stable.
@mnemosyne/cupThe Compatibility & Upgrade Path: v2→v3 migration tooling.
@mnemosyne/cliThe mnemo CLI (recall, remember, forget, status, sessions, tasks, events, health, export, import, login, …).
@mnemosyne/mock-serverAn in-memory server double for tests and SDK development — no Postgres required.
examples/graph-viewerA Next.js app demonstrating Memory Graph rendering against a live server. Reference implementation, not a shipped surface.

The three consumption modes

  1. Library mode — pnpm add @mnemosyne/core and call the primitives in your own process. Lowest latency, no separate service to operate, but the host app owns the DB connection and inherits the operational surface.
  2. Service mode (recommended) — run @mnemosyne/server next to your app (or anywhere — Fly, Render, K8s, the bundled Docker stack) and consume via @mnemosyne/client-ts. Decouples memory ops from the host process; one server scales independently and is the only thing that touches the memory DB. This is what production Orchester runs.
  3. MCP mode — host the stdio-only MCP server next to Claude Desktop / Cursor / custom agents. It binds no port; it is a thin wrapper over @mnemosyne/client-ts, so the API-key scoping and cognitive semantics are identical to direct REST.

All three modes go through the same @mnemosyne/core primitives. The isolation guarantees are identical regardless of mode: every operation runs under withMnemoTx, every API key pins a request to a single workspace.


4. The ten primitives

Mnemosyne models what an agent runtime needs as ten typed primitives — each with its own table, RLS policies, HNSW index, and MCP tool. The first four are the core memory shapes; the rest add procedural knowledge, work tracking, and the agent’s own reasoning trail.

PrimitiveTableWhat it is
Factmnemo_factDurable factual knowledge with hybrid recall (semantic + recency + decay + pin).
Decisionmnemo_decisionRecorded choice + rationale + supersedes link, queryable as a first-class object. Backed by facts via grounded_by.
Episodemnemo_episodeTimeline event with a duration and a bundle of facts that share its context.
Entitymnemo_entityCanonical person / organization / project / concept / place. Facts reference it via entity_id. Canonicalization collapses duplicates while preserving aliases.
Strategymnemo_strategyA procedural plan / approach — durable “how we do X” knowledge (procedural decay).
Workflowmnemo_workflowA multi-step procedure the agent can replay and refine over time (procedural decay).
Skillmnemo_skillA learned capability with a worth signal, evolved through use (procedural decay).
Taskmnemo_taskA unit of work with lifecycle state — first-class, queryable, never decays.
Eventmnemo_eventAn immutable record of something that happened. No decay, never rewritten.
Reasoningmnemo_reasoningA captured chain-of-thought / reasoning trace, retained conditionally on its worth.

An eleventh kind, hyperedge, represents n-ary associations (a single relation spanning more than two nodes) for the hypergraph layer.

Binary edges live in mnemo_relation with a typed relation column (derived_from, supersedes, part_of, member_of, scoped, related, plus contradiction-tracking variants) and provenance. The Memory Graph (§7) is literally a projection of these tables.

Adjacent state:

  • mnemo_query_cache — cross-process semantic cache for recall (§6).
  • mnemo_review_queue — low-confidence facts queued for human triage.
  • mnemo_summary — distilled per-workspace memory summary, for cached system-prompt prefixes on the agent side.
  • mnemo_fact_archive — bitemporal history sink for forgotten / superseded rows. Recall ignores it; audit + time-travel queries read from the UNION ALL.

5. Multi-tenant isolation (RLS + FORCE, Pattern A)

Every mnemo_* table is protected by row-level security with FORCE, gated on a per-transaction GUC, app.workspace_id. A row is only visible if its workspace_id matches the current setting; with the GUC unset, policies fail closed (return zero rows) rather than open.

FORCE is the crucial word. Without it, the table owner bypasses RLS — FORCE applies the policies even to the owning role. Without that, a misconfigured production role would silently skip the entire isolation layer.

The single entry point: withMnemoTx

withMnemoTx is the only sanctioned way to touch mnemo_* data. Inside one transaction it does two things, in order:

await tx.execute(sql`SET LOCAL ROLE app_user`);
await tx.execute(sql`SELECT set_config('app.workspace_id', workspaceId, true)`);
  1. SET LOCAL ROLE app_user — downgrades the transaction to app_user, a role defined in 0007_postgres_roles.sql as NOINHERIT LOGIN with no BYPASSRLS. This works even when the underlying connection is a superuser: SET ROLE to a non-BYPASSRLS role applies for the duration of the transaction. Because it’s LOCAL, it auto-reverts on commit/rollback, so the next reservation of the same pooled connection starts clean.
  2. set_config('app.workspace_id', …, true) — sets the workspace GUC LOCAL to the transaction. Releases on commit/rollback. Never bleeds across pooled requests.
import { withMnemoTx } from "@mnemosyne/core";
 
await withMnemoTx("ws_acme", async (tx) => {
  // Every mnemo_* query in here is scoped to ws_acme. A query that
  // forgot to filter by workspace still can't see another tenant's rows.
});

Layered defense

LayerWhat it does
TableRLS + FORCE + Pattern A policy gating on app.workspace_id, failing closed when unset.
TransactionwithMnemoTx downgrades to app_user and sets the workspace GUC LOCAL. Even a superuser connection is safe.
DeploymentConnect production directly as app_user so the elevated role is never on the wire.
CredentialAPI keys are bcrypt-hashed at rest; the key pins the request to a single workspace.
ContentPII regex detector + redactor; severity-weighted categories surfaced before persistence.

Optional per-actor isolation

For workspaces that need per-end-user privacy within a tenant, the MnemoTxOptions form of withMnemoTx sets app.actor_id and flips app.enforce_actor_isolation. When enabled, the SELECT policy on mnemo_fact restricts visible rows to actor_id IS NULL OR actor_id = $actorId. Defaults off to preserve existing read scope. An empty-string actor id is treated as unset so a sentinel can’t accidentally match a real id.

This matters when one server hosts many end-user agents on a single pooled connection: Alice’s private facts cannot leak into Bob’s recall even though both run against the same workspace.

Integration-tested, not just asserted

The isolation is verified end-to-end against a live pgvector container, not just asserted in policy SQL. The integration suite confirms that a second workspace (ws_b) sees zero of the first workspace’s (ws_a) facts, and that cross-tenant writes are rejected — RLS holds for both reads and writes.

See security for the full threat model.


6. The hybrid recall pipeline

Recall is not a single vector search. Mnemosyne blends lexical, dense, and graph signals, then reranks, because each catches what the others miss: Postgres full-text search (ts_rank_cd over tsvector) nails exact tokens and rare terms, dense vectors catch paraphrase, and graph expansion pulls in facts that are connected to a hit even if they don’t match the query text. The blend is what makes a 3-fact recall feel like the agent actually remembers, instead of finding the nearest cosine neighbor.

The pipeline, top to bottom:

  1. Query prep. Normalizes the query and (optionally) contextualizes it against recent turns (“what’s the same thing?” → resolve “it”). When HyDE is on, generates a hypothetical answer to embed instead of the bare question — empirically improves dense recall for under-specified queries.
  2. Cache layers. Three of them, fail-open:
    • L1 — in-process workspace LRU (60s TTL, ~5K entries, keyed on (workspaceId, queryHash, scope, scopeRef, topK, agentId)).
    • L2 — embedding LRU lives inside embedMnemo, workspace-keyed.
    • L3 — mnemo_query_cache table. Cross-pod semantic cache. A query whose embedding has cosine ≥ 0.95 against a row created in the last 5 minutes short-circuits the full pipeline. Skipped when asOf is set (time-travel must hit live SQL).
  3. Lexical search over fact statements via Postgres full-text ts_rank_cd (tsvector).
  4. Dense vector search over halfvec(1536) embeddings with cosine similarity. Embeddings are produced by your configured LLM provider (BYO) — Mnemosyne never makes that decision for you.
  5. Merge + normalize. Brings the two score scales onto a comparable [0, 1] range and blends them. The default blend is documented; the helper exposes per-call overrides.
  6. Graph expansion. Walks mnemo_relation edges from the strong hits to pull in connected facts that didn’t surface from the lexical/dense pass. Bounded to one hop in v1 — wider expansion is a known sharp edge.
  7. Cross-encoder rerank. Cohere when COHERE_API_KEY is set; otherwise a local lexical reranker over Postgres full-text scores. Identity rerank is considered a regression — there is always some rerank signal in v1.
  8. Stage caps + dedup. Drops near-duplicates (cosine > 0.88 against an already-kept fact), enforces per-stage caps, and applies a final hard cap (default topK=3 for agent system prompts; configurable per call).

Every stage is wrapped in try / catch. A flaky reranker, KB provider, or embedding error never breaks recall — it degrades to whatever earlier stages produced. Recall is optimisation; never a hard dependency for a turn.

KB chunks (optional fusion)

Unified recall can blend in knowledge-base chunks from a host-injected KbChunkProvider. Mnemosyne owns the blending policy (facts are weighted as dense conversational signal, KB chunks normalized against the max score in the response) and degrades gracefully to memory-only if the KB provider throws. The host app keeps full ownership of where KB content lives.


7. The Memory Graph

The same facts that recall searches also form a graph. Entities (people, orgs, projects, concepts, places) are connected by typed relations with a confidence and provenance, and episodes and decisions hang off them.

The graph layer is split into a client-safe surface (types + canvas drawing primitives + layout config, zero DB) and a server surface (buildGraphData / buildGraphQuery, which reads Postgres). This keeps the Postgres driver out of browser bundles.

buildGraphQuery runs inside withMnemoTx, so the graph is workspace-scoped by the same RLS guarantee as everything else — it only emits canonical entities, non-synthetic episodes, active decisions, and live (valid_to IS NULL, non-dismissed) relations.

The data contract, the client/server split rationale, and rendering recipes (2D + 3D, force-directed) live in rendering the graph. The reference renderer ships in examples/graph-viewer.


8. Bring your own LLM

Mnemosyne never hard-codes a provider. Embeddings, contextualization, HyDE, consolidation summaries — every LLM dispatch goes through the @mnemosyne/adapter-llm LLMProvider interface, fronted by an LLMRouter that picks a provider per task and a BudgetGuard that enforces a USD spend cap:

interface LLMProvider {
  embed(input: string | string[]): Promise<number[][]>;
  chat(messages: ChatMessage[], opts?: ChatOptions): Promise<ChatResponse>;
  capabilities(): ProviderCapabilities;  // dim, max tokens, supports vision, etc.
}

Ships with seven provider implementations:

ProviderEmbeddingsChat
OpenAItext-embedding-3-small/largegpt-4o-mini / gpt-4o family
Anthropic(delegates to a configured embedder)Claude 3.5 / Claude 4 family
GoogleGemini embeddingsGemini family
CohereCohere embeddingsCommand family (+ rerank)
MistralMistral embeddingsMistral family
VoyageVoyage embeddings—
Ollamalocal OpenAI-shaped endpointslocal OpenAI-shaped endpoints

Local stacks land through the Ollama / OpenAI-compatible path: Ollama, vLLM, LM Studio, LiteLLM all expose OpenAI-shaped APIs. Point the base URL at the local endpoint and Mnemosyne treats it like any other provider.

The LLMRouter lives in the workspace’s API-key context — different workspaces can run on different providers without code changes, and a workspace can swap providers without losing existing facts (the embedding dimension stays fixed at 1536 via halfvec(1536), so a workspace migrating providers re-embeds on the next write).


9. HTTP API surface (@mnemosyne/server)

The server exposes a Hono-based /v1/* REST surface — roughly 130 endpoints across 40 route modules. Every endpoint takes Authorization: Bearer mns_live_… and resolves the workspace from the key server-side — the workspace id is never accepted from the client (there is no X-Workspace-Id header). gRPC (ConnectRPC) and WebSocket are served on the same port.

A representative slice of the surface:

GroupEndpoints
FactsGET / POST /v1/facts, GET /v1/facts/:id, PATCH /v1/facts/:id, POST /v1/facts/:id/pin, POST /v1/facts/:id/forget, GET /v1/facts/:id/citations
RecallPOST /v1/recall (single op, hybrid pipeline)
EntitiesGET / POST /v1/entities, GET / PATCH /v1/entities/:id, GET /v1/entities/:id/facts
EpisodesGET /v1/episodes
DecisionsGET /v1/decisions, POST /v1/decisions, GET /v1/decisions/:id
RelationsGET / POST /v1/relations, PATCH /v1/relations/:id, DELETE /v1/relations/:id
Sessions/v1/session/* — start a session, record turns, collapse to typed memory
Self/v1/self/* — self-edit, reflect, profile (the agent’s autonomy surface)
Tasks/v1/tasks/* — task lifecycle as a first-class primitive
Events/v1/events/* — immutable event records
Reasoning/v1/reasoning/* — record and resolve reasoning traces
GraphGET /v1/graph (server-rendered graph payload for the Memory Graph)
ReviewGET /v1/review, POST /v1/review/:id/resolve
TimelineGET /v1/timeline (bitemporal event stream over facts)
AuditGET /v1/audit (UndoClient contract: items / total / available)
Federation/v1/federation/* — peer sync (9 routes)
ExportGET /v1/export (workspace dump for GDPR / migration)
HealthGET /v1/health (per-workspace stats), GET /healthz (unauthenticated liveness)

The full schemas (request, response, errors) are emitted as OpenAPI 3.1 and served at /v1/openapi.json; the server’s built-in Scalar reference UI lives at /v1/docs. The TypeScript SDK (@mnemosyne/client-ts) is the canonical client and is regenerated against the same schemas, so a client mismatch is a build-time error, not a runtime one.

Errors

Every endpoint returns { error: { code, message, details? } } on failure with the appropriate HTTP status. code is a stable string enum (UNAUTHORIZED, WORKSPACE_NOT_FOUND, FACT_NOT_FOUND, VALIDATION_FAILED, BLOCKED_BY_POISONING_GUARD, etc.). Clients gate on code, not on message text.


10. MCP — model context protocol

@mnemosyne/mcp exposes the Mnemosyne primitives as MCP tools so any MCP client (Claude Desktop, Cursor, Gemini, custom agents) can use them without a code integration. 26 tools, mapped onto the cognitive surface:

GroupTools
Core recallmemory_recall, memory_remember, memory_pin, memory_forget, memory_timeline, memory_search, memory_status
Graphmemory_relate, memory_entity, memory_hyperedge
Decisionsmemory_decide
Sessions / episodesmemory_session, memory_episode_start, memory_episode_end
Workmemory_task, memory_event, memory_workflow, memory_skill, memory_strategy
Reasoningmemory_reasoning, memory_self_edit, memory_introspect
Zero-LLMmemory_anticipate, memory_compose
Maintenancememory_consolidate, memory_remind_when

Transport is stdio only — the MCP server binds no port. It goes through the same @mnemosyne/client-ts SDK, so API-key scoping, workspace pinning, and isolation guarantees are identical to direct REST.


11. Operations — maintenance crons

@mnemosyne/server runs a small fleet of background jobs to keep the fact store healthy. All run inside withMnemoTx(workspaceId, …) and are pure functions of the workspace’s current state — restartable, idempotent, and budget-respecting (every LLM dispatch records against the workspace’s spend cap when the host wires one).

JobDefault schedulePurpose
embed-batch*/1 * * * *Fill embedding = NULL rows in batched API calls (createFactAsync defers cost by writing the row first).
summary0 5 * * *Distill per-workspace memory summary into mnemo_summary for cached system-prompt prefixes.
health0 6 * * *Per-workspace memory-drift snapshot (recall hit rate, embedding coverage, backlog).
dedup0 3 * * 0Semantic dedup; archive losers into mnemo_fact_archive.
prune30 3 * * 0Drop inactive facts past the grace window.
review-sweep0 4 * * *Queue low-confidence facts for human triage in mnemo_review_queue.
auto-pin30 4 * * *Apply pure auto-pin rules (decideAutoPin) to high-recall facts.
consolidation0 2 * * 0REM-style: cluster related facts, ask the cheap LLM for a one-sentence supersede summary.
decay0 4 * * *Exponential decay of memory_strength (true half-life, H=30d, pinned facts exempt).

Consolidation runs before dedup on Sunday so the janitor sees a stable graph: the new summary has diverged enough from its members not to be collapsed back. Each cron is wrapped in withMnemoTx per workspace, never one tx for the whole sweep — long-running maintenance never holds a pool connection.

Per-workspace cadence overrides live in mnemo_cron_schedule: a workspace can disable, throttle, or accelerate any of the above without code changes.


12. Deployment

Mnemosyne is one binary + one Postgres. Two reference topologies, both single-tenant from Mnemosyne’s perspective (the multi-tenant boundary is the workspace, not the deployment).

Single-node — bundled Docker stack

Docker host
├── mnemosyne-server     Hono app (3000)
└── mnemosyne-postgres   Postgres 17 + pgvector

The repository ships a docker/compose.yaml that brings both up with sane defaults. Suits dev, demo, and small production. Bundle Postgres in docker-compose for self-hosting; swap in managed Postgres for production by pointing MNEMO_DATABASE_URL at it.

Service + managed Postgres

Fly / Render / Railway / K8s
└── mnemosyne-server     Stateless; horizontally scalable

Managed Postgres (Supabase, Neon, RDS, Aiven, CrunchyData)
└── pgvector ≥ 0.7 enabled

The server is stateless once the DB connection is established. Horizontal scaling is uncomplicated — crons are guarded by a per-job advisory lock so only one replica runs each sweep, but every replica serves API traffic.

Migrations run as a dedicated migrate service (npx mnemo-migrate) that must exit 0 before the server starts — schema changes are never applied automatically on first boot, and they are not reversible from a CLI command. The journal is 75 numbered files (0001–0075).


13. Load-bearing decisions

The decisions that, if reversed, would force re-architecture:

DecisionWhy
Postgres only — no separate vector DBOne ACID store. Bitemporal facts and their vectors live in the same row, so transactional consistency between “I remembered X” and “X is searchable” is free.
Bitemporal columns on every primitiveLets the system answer time-travel queries without an event-sourcing rewrite. Forget / supersede are O(1) UPDATEs, not log replays.
RLS + FORCE + app_user role downgradeThe only model where a misconfigured deployment fails closed instead of leaking. Integration-tested across pooled connections.
halfvec(1536) embeddingsRoughly half the storage and IO of vector(1536) for a negligible recall delta. The dimension is fixed so a workspace can swap providers without a schema change.
BYO LLM via LLMProviderMnemosyne’s value is the cognitive contract, not the embedding model. Letting the host bring their own keys keeps Mnemosyne under any AUP / privacy constraint they already have.
HTTP service as the canonical modeLibrary mode is supported (and fast), but the service decouples ops. One thing scales, one thing owns the DB, one thing rolls upgrades.
MCP as a first-class transportThe fastest way for an end-user to give an LLM persistent memory is to make Mnemosyne look like a tool. The MCP server is thin (it just adapts the SDK) so it stays in lockstep.
Cohere reranker preferred, local FTS-based fallbackA good cross-encoder is the single biggest recall-quality lever. Defaulting to one (when available) means new deployments get the upgrade for free; the local fallback (Postgres full-text scores) means the pipeline still produces ranked output without paid signal.
Per-actor isolation is opt-in, off by defaultThe common case is workspace-shared memory. Forcing per-actor on every read would break those workloads; making it opt-in (with a fail-closed semantics when on) gives the per-end-user use case a strong guarantee.

14. Further reading