Mnemosyne is a cognitive memory layer for AI agents: it stores facts about a workspace, keeps them honest over time, and retrieves the relevant ones on demand. This document is the canonical system spec — the complete picture of what Mnemosyne is, what it guarantees, and how the moving parts fit together.
Audience: contributors, integrators evaluating Mnemosyne for production, and anyone wiring an agent runtime against it. Reader is assumed to know Postgres and a bit of LLM-app plumbing. Out of scope: tutorials — see the getting-started guide for that.
1. What Mnemosyne is (and is not)
Mnemosyne is the substrate between your agents and Postgres. Ten typed cognitive
primitives — fact, decision, episode, entity, strategy,
workflow, skill, task, event, and reasoning (plus
hyperedge for n-ary relations) — sit behind a small, opinionated API:
recall() → hybrid search: Postgres FTS + dense vectors + graph expansion + rerank
remember() → write a fact with provenance, confidence, and valid-time
forget() → bitemporal soft-invalidate — history is never destroyed
session() → start a session, record turns, collapse it to typed memoryIt is not a vector store with a nicer wrapper. The distinguishing claims:
- Bitemporal by construction — every fact, relation, and decision carries
valid_from/valid_to, so the system answers both “what is true now?” and “what did we believe last Tuesday?”. - Multi-tenant in the database, not in the
WHEREclause — RLS +FORCE, per-transaction GUCs, role downgrade, integration-tested cross-tenant isolation. Fails closed when the workspace context is unset. - Provider-agnostic — bring your own LLM. Seven providers (OpenAI,
Anthropic, Google, Cohere, Mistral, Voyage, Ollama) plug through the same
LLMProviderinterface, fronted by anLLMRouter(per-task routing) and aBudgetGuard(USD spend cap). The cognitive contract is independent of which embedder draws the vectors. - Three consumption modes — import as a library, run as an HTTP service, or expose as MCP tools for Claude Desktop / Cursor. Same model, same isolation guarantees, three deployment shapes.
2. The cognitive model: bitemporal facts
A fact in Mnemosyne is not a row you overwrite. It is a versioned assertion with two independent time axes:
- Valid time (
valid_from/valid_to) — when the fact is true in the world. A fact that became true on Monday and was superseded on Friday hasvalid_from = Monday,valid_to = Friday. An open-ended fact hasvalid_to = NULL. - Transaction time — when Mnemosyne recorded the fact (the row’s creation and its status transitions). This axis is append-only history: it never rewrites.
The mnemo_fact, mnemo_relation, and mnemo_decision tables all carry
valid_from (NOT NULL DEFAULT now()) and a nullable valid_to, giving each
the same bitemporal lifecycle.
Why two axes? Because the world changes and your knowledge of it changes, and those are different events. Bitemporality lets you correct a mistaken belief without erasing the fact that you once held it — essential for auditable, debuggable agent memory.
Each active fact also carries a memory_strength — a potentiation score with a
neutral baseline of 1.0. Recall and decay nudge it over time so often-recalled
facts surface more readily and stale ones fade. This is the Hebbian rule applied
to retrieval: what fires together, recalls together.
Theory-of-mind discriminators
Beyond bitemporal time, every primitive carries three discriminators that let the agent reason about its own knowledge:
memory_type∈semantic/episodic/procedural/working— the classic cognitive-science cut. Procedural facts (“the API reset clock is the 1st of every month”) are weighted differently from working memory (“the user just said X”).attribution∈user_stated/user_belief/objective_fact/inferred— who asserted this. Lets recall surface “the user thinks X” separately from “X is true”. Essential for honest agents.protocol_version— the Memory Protocol revision the row was extracted under. Lets future extractors evolve without losing the ability to interpret older rows.
3. System topology
Mnemosyne is one Postgres database (plus a Redis hot tier) and a fleet of
packages. The fleet ships as a pnpm monorepo so the cognitive contract
(@mnemosyne/core) is shared across every consumption mode and the transport
layers — REST, gRPC, WebSocket, and MCP — compose without duplication.
The packages
| Package | Role |
|---|---|
@mnemosyne/core | The cognitive contract: the 10 primitives, recall pipeline, the zero-LLM engine (anticipation / composition / consensus), withMnemoTx, bitemporal queries, graph builders, PII detection. Framework-agnostic, no server-only. |
@mnemosyne/types | Shared type contract (PrimitiveKind, DTOs, errors) imported by every other package so the surface stays consistent. |
@mnemosyne/server | A Hono-based HTTP service exposing /v1/* (REST), plus gRPC (ConnectRPC) and WebSocket on the same port. Owns cron scheduling, API-key auth, and the public surface. Run via docker compose up. |
@mnemosyne/client-ts | A typed TypeScript SDK over the v1 API. Zero runtime dependencies. The canonical way to integrate Mnemosyne from another Node app. |
@mnemosyne/mcp | A stdio-only MCP server exposing 26 memory tools (recall, remember, session, task, event, reasoning, self_edit, …). Binds no port. Plugs straight into Claude Desktop, Cursor, custom agents. |
@mnemosyne/adapter-postgres | The Postgres adapter: schema, RLS policies, bitemporal queries, the Drizzle + postgres driver seam. |
@mnemosyne/adapter-redis | The hot-tier adapter — ACT-R activation L0/L1 over Redis for sub-millisecond recall of the most-active memories. |
@mnemosyne/adapter-vector | The vector seam — HNSW search over halfvec(1536), pluggable so the embedding store can evolve. |
@mnemosyne/adapter-llm | The BYO-LLM seam: a LLMProvider interface, 7 provider implementations (OpenAI, Anthropic, Google, Cohere, Mistral, Voyage, Ollama), the LLMRouter, and the BudgetGuard spend cap. |
@mnemosyne/consolidation | The consolidation worker: ConsolidationGate + ReplayValidator + scheduler, wired at server boot. |
@mnemosyne/federation | Multi-instance peer sync: Ed25519 + mTLS, LWW / supermajority / consensus conflict resolution, 9 HTTP routes. |
@mnemosyne/plugins | The plugin runtime plus 9 builtin plugins. Stable. |
@mnemosyne/cup | The Compatibility & Upgrade Path: v2→v3 migration tooling. |
@mnemosyne/cli | The mnemo CLI (recall, remember, forget, status, sessions, tasks, events, health, export, import, login, …). |
@mnemosyne/mock-server | An in-memory server double for tests and SDK development — no Postgres required. |
examples/graph-viewer | A Next.js app demonstrating Memory Graph rendering against a live server. Reference implementation, not a shipped surface. |
The three consumption modes
- Library mode —
pnpm add @mnemosyne/coreand call the primitives in your own process. Lowest latency, no separate service to operate, but the host app owns the DB connection and inherits the operational surface. - Service mode (recommended) — run
@mnemosyne/servernext to your app (or anywhere — Fly, Render, K8s, the bundled Docker stack) and consume via@mnemosyne/client-ts. Decouples memory ops from the host process; one server scales independently and is the only thing that touches the memory DB. This is what production Orchester runs. - MCP mode — host the stdio-only MCP server next to Claude Desktop /
Cursor / custom agents. It binds no port; it is a thin wrapper over
@mnemosyne/client-ts, so the API-key scoping and cognitive semantics are identical to direct REST.
All three modes go through the same @mnemosyne/core primitives. The
isolation guarantees are identical regardless of mode: every operation runs
under withMnemoTx, every API key pins a request to a single workspace.
4. The ten primitives
Mnemosyne models what an agent runtime needs as ten typed primitives — each with its own table, RLS policies, HNSW index, and MCP tool. The first four are the core memory shapes; the rest add procedural knowledge, work tracking, and the agent’s own reasoning trail.
| Primitive | Table | What it is |
|---|---|---|
| Fact | mnemo_fact | Durable factual knowledge with hybrid recall (semantic + recency + decay + pin). |
| Decision | mnemo_decision | Recorded choice + rationale + supersedes link, queryable as a first-class object. Backed by facts via grounded_by. |
| Episode | mnemo_episode | Timeline event with a duration and a bundle of facts that share its context. |
| Entity | mnemo_entity | Canonical person / organization / project / concept / place. Facts reference it via entity_id. Canonicalization collapses duplicates while preserving aliases. |
| Strategy | mnemo_strategy | A procedural plan / approach — durable “how we do X” knowledge (procedural decay). |
| Workflow | mnemo_workflow | A multi-step procedure the agent can replay and refine over time (procedural decay). |
| Skill | mnemo_skill | A learned capability with a worth signal, evolved through use (procedural decay). |
| Task | mnemo_task | A unit of work with lifecycle state — first-class, queryable, never decays. |
| Event | mnemo_event | An immutable record of something that happened. No decay, never rewritten. |
| Reasoning | mnemo_reasoning | A captured chain-of-thought / reasoning trace, retained conditionally on its worth. |
An eleventh kind, hyperedge, represents n-ary associations (a single
relation spanning more than two nodes) for the hypergraph layer.
Binary edges live in mnemo_relation with a typed relation column
(derived_from, supersedes, part_of, member_of, scoped, related,
plus contradiction-tracking variants) and provenance. The Memory Graph (§7) is
literally a projection of these tables.
Adjacent state:
mnemo_query_cache— cross-process semantic cache for recall (§6).mnemo_review_queue— low-confidence facts queued for human triage.mnemo_summary— distilled per-workspace memory summary, for cached system-prompt prefixes on the agent side.mnemo_fact_archive— bitemporal history sink for forgotten / superseded rows. Recall ignores it; audit + time-travel queries read from the UNION ALL.
5. Multi-tenant isolation (RLS + FORCE, Pattern A)
Every mnemo_* table is protected by row-level security with FORCE, gated
on a per-transaction GUC, app.workspace_id. A row is only visible if its
workspace_id matches the current setting; with the GUC unset, policies fail
closed (return zero rows) rather than open.
FORCE is the crucial word. Without it, the table owner bypasses RLS — FORCE
applies the policies even to the owning role. Without that, a misconfigured
production role would silently skip the entire isolation layer.
The single entry point: withMnemoTx
withMnemoTx is the only sanctioned way to touch mnemo_* data. Inside one
transaction it does two things, in order:
await tx.execute(sql`SET LOCAL ROLE app_user`);
await tx.execute(sql`SELECT set_config('app.workspace_id', workspaceId, true)`);SET LOCAL ROLE app_user— downgrades the transaction toapp_user, a role defined in0007_postgres_roles.sqlasNOINHERIT LOGINwith noBYPASSRLS. This works even when the underlying connection is a superuser:SET ROLEto a non-BYPASSRLSrole applies for the duration of the transaction. Because it’sLOCAL, it auto-reverts on commit/rollback, so the next reservation of the same pooled connection starts clean.set_config('app.workspace_id', …, true)— sets the workspace GUCLOCALto the transaction. Releases on commit/rollback. Never bleeds across pooled requests.
import { withMnemoTx } from "@mnemosyne/core";
await withMnemoTx("ws_acme", async (tx) => {
// Every mnemo_* query in here is scoped to ws_acme. A query that
// forgot to filter by workspace still can't see another tenant's rows.
});Layered defense
| Layer | What it does |
|---|---|
| Table | RLS + FORCE + Pattern A policy gating on app.workspace_id, failing closed when unset. |
| Transaction | withMnemoTx downgrades to app_user and sets the workspace GUC LOCAL. Even a superuser connection is safe. |
| Deployment | Connect production directly as app_user so the elevated role is never on the wire. |
| Credential | API keys are bcrypt-hashed at rest; the key pins the request to a single workspace. |
| Content | PII regex detector + redactor; severity-weighted categories surfaced before persistence. |
Optional per-actor isolation
For workspaces that need per-end-user privacy within a tenant, the
MnemoTxOptions form of withMnemoTx sets app.actor_id and flips
app.enforce_actor_isolation. When enabled, the SELECT policy on mnemo_fact
restricts visible rows to actor_id IS NULL OR actor_id = $actorId. Defaults
off to preserve existing read scope. An empty-string actor id is treated as
unset so a sentinel can’t accidentally match a real id.
This matters when one server hosts many end-user agents on a single pooled connection: Alice’s private facts cannot leak into Bob’s recall even though both run against the same workspace.
Integration-tested, not just asserted
The isolation is verified end-to-end against a live pgvector container, not just
asserted in policy SQL. The integration suite confirms that a second workspace
(ws_b) sees zero of the first workspace’s (ws_a) facts, and that
cross-tenant writes are rejected — RLS holds for both reads and writes.
See security for the full threat model.
6. The hybrid recall pipeline
Recall is not a single vector search. Mnemosyne blends lexical, dense, and
graph signals, then reranks, because each catches what the others miss:
Postgres full-text search (ts_rank_cd over tsvector) nails exact tokens and
rare terms, dense vectors catch paraphrase, and graph expansion pulls in facts
that are connected to a hit even if they don’t match the query text. The blend
is what makes a 3-fact recall feel like the agent actually remembers, instead of
finding the nearest cosine neighbor.
The pipeline, top to bottom:
- Query prep. Normalizes the query and (optionally) contextualizes it against recent turns (“what’s the same thing?” → resolve “it”). When HyDE is on, generates a hypothetical answer to embed instead of the bare question — empirically improves dense recall for under-specified queries.
- Cache layers. Three of them, fail-open:
- L1 — in-process workspace LRU (60s TTL, ~5K entries, keyed on
(workspaceId, queryHash, scope, scopeRef, topK, agentId)). - L2 — embedding LRU lives inside
embedMnemo, workspace-keyed. - L3 —
mnemo_query_cachetable. Cross-pod semantic cache. A query whose embedding has cosine ≥ 0.95 against a row created in the last 5 minutes short-circuits the full pipeline. Skipped whenasOfis set (time-travel must hit live SQL).
- L1 — in-process workspace LRU (60s TTL, ~5K entries, keyed on
- Lexical search over fact statements via Postgres full-text
ts_rank_cd(tsvector). - Dense vector search over
halfvec(1536)embeddings with cosine similarity. Embeddings are produced by your configured LLM provider (BYO) — Mnemosyne never makes that decision for you. - Merge + normalize. Brings the two score scales onto a comparable
[0, 1]range and blends them. The default blend is documented; the helper exposes per-call overrides. - Graph expansion. Walks
mnemo_relationedges from the strong hits to pull in connected facts that didn’t surface from the lexical/dense pass. Bounded to one hop in v1 — wider expansion is a known sharp edge. - Cross-encoder rerank. Cohere when
COHERE_API_KEYis set; otherwise a local lexical reranker over Postgres full-text scores. Identity rerank is considered a regression — there is always some rerank signal in v1. - Stage caps + dedup. Drops near-duplicates (cosine > 0.88 against an
already-kept fact), enforces per-stage caps, and applies a final hard cap
(default
topK=3for agent system prompts; configurable per call).
Every stage is wrapped in try / catch. A flaky reranker, KB provider, or
embedding error never breaks recall — it degrades to whatever earlier stages
produced. Recall is optimisation; never a hard dependency for a turn.
KB chunks (optional fusion)
Unified recall can blend in knowledge-base chunks from a host-injected
KbChunkProvider. Mnemosyne owns the blending policy (facts are weighted as
dense conversational signal, KB chunks normalized against the max score in the
response) and degrades gracefully to memory-only if the KB provider throws. The
host app keeps full ownership of where KB content lives.
7. The Memory Graph
The same facts that recall searches also form a graph. Entities (people,
orgs, projects, concepts, places) are connected by typed relations with a
confidence and provenance, and episodes and decisions hang off them.
The graph layer is split into a client-safe surface (types + canvas drawing
primitives + layout config, zero DB) and a server surface (buildGraphData
/ buildGraphQuery, which reads Postgres). This keeps the Postgres driver out
of browser bundles.
buildGraphQuery runs inside withMnemoTx, so the graph is workspace-scoped by
the same RLS guarantee as everything else — it only emits canonical entities,
non-synthetic episodes, active decisions, and live (valid_to IS NULL,
non-dismissed) relations.
The data contract, the client/server split rationale, and rendering recipes
(2D + 3D, force-directed) live in
rendering the graph. The reference renderer
ships in examples/graph-viewer.
8. Bring your own LLM
Mnemosyne never hard-codes a provider. Embeddings, contextualization, HyDE,
consolidation summaries — every LLM dispatch goes through the
@mnemosyne/adapter-llm LLMProvider interface, fronted by an LLMRouter that
picks a provider per task and a BudgetGuard that enforces a USD spend cap:
interface LLMProvider {
embed(input: string | string[]): Promise<number[][]>;
chat(messages: ChatMessage[], opts?: ChatOptions): Promise<ChatResponse>;
capabilities(): ProviderCapabilities; // dim, max tokens, supports vision, etc.
}Ships with seven provider implementations:
| Provider | Embeddings | Chat |
|---|---|---|
OpenAI | text-embedding-3-small/large | gpt-4o-mini / gpt-4o family |
Anthropic | (delegates to a configured embedder) | Claude 3.5 / Claude 4 family |
Google | Gemini embeddings | Gemini family |
Cohere | Cohere embeddings | Command family (+ rerank) |
Mistral | Mistral embeddings | Mistral family |
Voyage | Voyage embeddings | — |
Ollama | local OpenAI-shaped endpoints | local OpenAI-shaped endpoints |
Local stacks land through the Ollama / OpenAI-compatible path: Ollama, vLLM, LM Studio, LiteLLM all expose OpenAI-shaped APIs. Point the base URL at the local endpoint and Mnemosyne treats it like any other provider.
The LLMRouter lives in the workspace’s API-key context — different workspaces
can run on different providers without code changes, and a workspace can swap
providers without losing existing facts (the embedding dimension stays fixed at
1536 via halfvec(1536), so a workspace migrating providers re-embeds on the
next write).
9. HTTP API surface (@mnemosyne/server)
The server exposes a Hono-based /v1/* REST surface — roughly 130 endpoints
across 40 route modules. Every endpoint takes Authorization: Bearer mns_live_… and resolves the workspace from the key server-side — the workspace
id is never accepted from the client (there is no X-Workspace-Id header). gRPC
(ConnectRPC) and WebSocket are served on the same port.
A representative slice of the surface:
| Group | Endpoints |
|---|---|
| Facts | GET / POST /v1/facts, GET /v1/facts/:id, PATCH /v1/facts/:id, POST /v1/facts/:id/pin, POST /v1/facts/:id/forget, GET /v1/facts/:id/citations |
| Recall | POST /v1/recall (single op, hybrid pipeline) |
| Entities | GET / POST /v1/entities, GET / PATCH /v1/entities/:id, GET /v1/entities/:id/facts |
| Episodes | GET /v1/episodes |
| Decisions | GET /v1/decisions, POST /v1/decisions, GET /v1/decisions/:id |
| Relations | GET / POST /v1/relations, PATCH /v1/relations/:id, DELETE /v1/relations/:id |
| Sessions | /v1/session/* — start a session, record turns, collapse to typed memory |
| Self | /v1/self/* — self-edit, reflect, profile (the agent’s autonomy surface) |
| Tasks | /v1/tasks/* — task lifecycle as a first-class primitive |
| Events | /v1/events/* — immutable event records |
| Reasoning | /v1/reasoning/* — record and resolve reasoning traces |
| Graph | GET /v1/graph (server-rendered graph payload for the Memory Graph) |
| Review | GET /v1/review, POST /v1/review/:id/resolve |
| Timeline | GET /v1/timeline (bitemporal event stream over facts) |
| Audit | GET /v1/audit (UndoClient contract: items / total / available) |
| Federation | /v1/federation/* — peer sync (9 routes) |
| Export | GET /v1/export (workspace dump for GDPR / migration) |
| Health | GET /v1/health (per-workspace stats), GET /healthz (unauthenticated liveness) |
The full schemas (request, response, errors) are emitted as OpenAPI 3.1 and
served at /v1/openapi.json; the server’s built-in Scalar reference UI
lives at /v1/docs. The TypeScript SDK (@mnemosyne/client-ts) is the
canonical client and is regenerated against the same schemas, so a client
mismatch is a build-time error, not a runtime one.
Errors
Every endpoint returns { error: { code, message, details? } } on failure with
the appropriate HTTP status. code is a stable string enum
(UNAUTHORIZED, WORKSPACE_NOT_FOUND, FACT_NOT_FOUND,
VALIDATION_FAILED, BLOCKED_BY_POISONING_GUARD, etc.). Clients gate on
code, not on message text.
10. MCP — model context protocol
@mnemosyne/mcp exposes the Mnemosyne primitives as MCP tools so any MCP client
(Claude Desktop, Cursor, Gemini, custom agents) can use them without a code
integration. 26 tools, mapped onto the cognitive surface:
| Group | Tools |
|---|---|
| Core recall | memory_recall, memory_remember, memory_pin, memory_forget, memory_timeline, memory_search, memory_status |
| Graph | memory_relate, memory_entity, memory_hyperedge |
| Decisions | memory_decide |
| Sessions / episodes | memory_session, memory_episode_start, memory_episode_end |
| Work | memory_task, memory_event, memory_workflow, memory_skill, memory_strategy |
| Reasoning | memory_reasoning, memory_self_edit, memory_introspect |
| Zero-LLM | memory_anticipate, memory_compose |
| Maintenance | memory_consolidate, memory_remind_when |
Transport is stdio only — the MCP server binds no port. It goes through the
same @mnemosyne/client-ts SDK, so API-key scoping, workspace pinning, and
isolation guarantees are identical to direct REST.
11. Operations — maintenance crons
@mnemosyne/server runs a small fleet of background jobs to keep the fact
store healthy. All run inside withMnemoTx(workspaceId, …) and are pure
functions of the workspace’s current state — restartable, idempotent, and
budget-respecting (every LLM dispatch records against the workspace’s spend cap
when the host wires one).
| Job | Default schedule | Purpose |
|---|---|---|
embed-batch | */1 * * * * | Fill embedding = NULL rows in batched API calls (createFactAsync defers cost by writing the row first). |
summary | 0 5 * * * | Distill per-workspace memory summary into mnemo_summary for cached system-prompt prefixes. |
health | 0 6 * * * | Per-workspace memory-drift snapshot (recall hit rate, embedding coverage, backlog). |
dedup | 0 3 * * 0 | Semantic dedup; archive losers into mnemo_fact_archive. |
prune | 30 3 * * 0 | Drop inactive facts past the grace window. |
review-sweep | 0 4 * * * | Queue low-confidence facts for human triage in mnemo_review_queue. |
auto-pin | 30 4 * * * | Apply pure auto-pin rules (decideAutoPin) to high-recall facts. |
consolidation | 0 2 * * 0 | REM-style: cluster related facts, ask the cheap LLM for a one-sentence supersede summary. |
decay | 0 4 * * * | Exponential decay of memory_strength (true half-life, H=30d, pinned facts exempt). |
Consolidation runs before dedup on Sunday so the janitor sees a stable
graph: the new summary has diverged enough from its members not to be collapsed
back. Each cron is wrapped in withMnemoTx per workspace, never one tx for the
whole sweep — long-running maintenance never holds a pool connection.
Per-workspace cadence overrides live in mnemo_cron_schedule: a workspace can
disable, throttle, or accelerate any of the above without code changes.
12. Deployment
Mnemosyne is one binary + one Postgres. Two reference topologies, both single-tenant from Mnemosyne’s perspective (the multi-tenant boundary is the workspace, not the deployment).
Single-node — bundled Docker stack
Docker host
├── mnemosyne-server Hono app (3000)
└── mnemosyne-postgres Postgres 17 + pgvectorThe repository ships a docker/compose.yaml that brings both up with sane
defaults. Suits dev, demo, and small production. Bundle Postgres in
docker-compose for self-hosting; swap in managed Postgres for production by
pointing MNEMO_DATABASE_URL at it.
Service + managed Postgres
Fly / Render / Railway / K8s
└── mnemosyne-server Stateless; horizontally scalable
Managed Postgres (Supabase, Neon, RDS, Aiven, CrunchyData)
└── pgvector ≥ 0.7 enabledThe server is stateless once the DB connection is established. Horizontal scaling is uncomplicated — crons are guarded by a per-job advisory lock so only one replica runs each sweep, but every replica serves API traffic.
Migrations run as a dedicated migrate service (npx mnemo-migrate) that must
exit 0 before the server starts — schema changes are never applied
automatically on first boot, and they are not reversible from a CLI command. The
journal is 75 numbered files (0001–0075).
13. Load-bearing decisions
The decisions that, if reversed, would force re-architecture:
| Decision | Why |
|---|---|
| Postgres only — no separate vector DB | One ACID store. Bitemporal facts and their vectors live in the same row, so transactional consistency between “I remembered X” and “X is searchable” is free. |
| Bitemporal columns on every primitive | Lets the system answer time-travel queries without an event-sourcing rewrite. Forget / supersede are O(1) UPDATEs, not log replays. |
RLS + FORCE + app_user role downgrade | The only model where a misconfigured deployment fails closed instead of leaking. Integration-tested across pooled connections. |
halfvec(1536) embeddings | Roughly half the storage and IO of vector(1536) for a negligible recall delta. The dimension is fixed so a workspace can swap providers without a schema change. |
BYO LLM via LLMProvider | Mnemosyne’s value is the cognitive contract, not the embedding model. Letting the host bring their own keys keeps Mnemosyne under any AUP / privacy constraint they already have. |
| HTTP service as the canonical mode | Library mode is supported (and fast), but the service decouples ops. One thing scales, one thing owns the DB, one thing rolls upgrades. |
| MCP as a first-class transport | The fastest way for an end-user to give an LLM persistent memory is to make Mnemosyne look like a tool. The MCP server is thin (it just adapts the SDK) so it stays in lockstep. |
| Cohere reranker preferred, local FTS-based fallback | A good cross-encoder is the single biggest recall-quality lever. Defaulting to one (when available) means new deployments get the upgrade for free; the local fallback (Postgres full-text scores) means the pipeline still produces ranked output without paid signal. |
| Per-actor isolation is opt-in, off by default | The common case is workspace-shared memory. Forcing per-actor on every read would break those workloads; making it opt-in (with a fail-closed semantics when on) gives the per-end-user use case a strong guarantee. |
14. Further reading
- the getting-started guide — from empty DB to first recall.
security— the threat model and the test evidence.rendering the graph— the Memory Graph data contract, client/server split rationale, and rendering recipes.examples/graph-viewer— runnable Next.js reference renderer.- changelog — release notes.
- roadmap — what’s next.