Mnemosyne is a multi-tenant cognitive memory layer for AI agents — bitemporal, provider-agnostic (bring your own LLM), backed by Postgres + pgvector, and isolated per workspace with row-level security. This guide takes you from an empty database to your first recall, through whichever of the three consumption paths fits your stack.
Requirements: Node 20+, pnpm 9+, and a Postgres 14+ database with the
pgvector extension available.
1. Provision Postgres + pgvector
The fastest path is the bundled Docker stack, which brings up Postgres with pgvector preinstalled:
git clone https://github.com/lucasmailland/mnemosyne.git
cd mnemosyne
pnpm docker:up # docker compose -f docker/docker-compose.yml up -dIf you bring your own Postgres, enable the extension first — the migrator does not create it for you:
psql "$DATABASE_URL" -c "CREATE EXTENSION IF NOT EXISTS vector;"2. Run migrations (mnemo-migrate)
@mnemosyne/core ships a migration runner as a binary. It discovers the bundled
SQL migrations, creates a mnemo_migration_history tracking table on first run,
and applies each migration in its own transaction (so a failure halts cleanly
without leaving a half-applied schema).
export DATABASE_URL="postgres://user:pass@localhost:5432/mnemosyne"
npx mnemo-migrate # apply all pending migrations
npx mnemo-migrate --dry-run # show what would run, apply nothing
npx mnemo-migrate --target 0042 # stop after migration 0042Migration
0015_mnemo_rls_foundation.sqldefines the RLS helper functions (apply_pattern_a(),current_workspace_id()) that every later migration depends on. Don’t reorder or skip migrations.
This is a forward-only runner. The bundled .down.sql files exist for
completeness, but rolling back is intentionally a manual psql operation to
avoid accidentally dropping production memory.
3. Choose a consumption path
There are three ways to consume Mnemosyne. They’re not mutually exclusive — the
HTTP server and MCP server are both thin shells over @mnemosyne/core.
Path A — Library (in-process)
Import @mnemosyne/core directly for sub-millisecond, in-process recall. Every
tenant-scoped operation runs inside withMnemoTx, which sets the workspace GUC
and downgrades the transaction role so RLS actually enforces (see
the security model).
import { setDb, getDb, withMnemoTx, createCore } from "@mnemosyne/core";
import postgres from "postgres";
import { drizzle } from "drizzle-orm/postgres-js";
import { schema } from "@mnemosyne/core";
// Wire the DB once at boot.
setDb(drizzle(postgres(process.env.DATABASE_URL!), { schema }));
const core = createCore({ /* inject your LLM provider for embeddings */ });
// Remember something, then recall it — all scoped to one workspace.
await withMnemoTx("ws_acme", async (tx) => {
await core.createFact(tx, "ws_acme", { statement: "User prefers dark mode." });
});
const { hits } = await withMnemoTx("ws_acme", (tx) =>
core.recall(tx, "ws_acme", { query: "what UI theme does the user like?" }),
);Path B — HTTP API (@mnemosyne/server)
Run the Hono server and talk to it over HTTP — ideal when your agent runtime
isn’t TypeScript, or when you want a network boundary between your app and the
memory store. It defaults to port 3000.
DATABASE_URL="postgres://…" \
MNEMO_LLM_PROVIDER=openai \
MNEMO_LLM_API_KEY="sk-…" \
npx mnemosyne-serverUse the typed SDK from a TypeScript client:
import { MnemosyneClient } from "@mnemosyne/client-ts";
const memory = new MnemosyneClient({
url: "http://localhost:3000",
apiKey: process.env.MNEMO_KEY!, // a per-workspace API key
});
await memory.createFact({ statement: "Ship date moved to Q3." });
const { hits } = await memory.recall({ query: "when do we ship?" });The API key both authenticates the caller and pins the request to a workspace, so the SDK never sends a workspace id over the wire.
Path C — MCP server (@mnemosyne/mcp)
Run the standalone Model Context Protocol server so MCP-capable clients (Claude
Desktop, IDEs, custom agents) get mnemosyne_* tools over stdio. It’s a thin MCP
shell over the HTTP API:
// e.g. an MCP client config
{
"mcpServers": {
"mnemosyne": {
"command": "npx",
"args": ["mnemosyne-mcp"],
"env": {
"MNEMO_URL": "http://localhost:3000",
"MNEMO_KEY": "your-workspace-api-key"
}
}
}
}4. Bring your own LLM
Mnemosyne never bundles model credentials — you pay for exactly what you use.
Embeddings (for vector recall) and any extraction steps run through whichever
provider you configure via @mnemosyne/llm-providers, which ships adapters for
OpenAI, Anthropic, Google, Cohere, Mistral, Voyage, and Ollama, plus an
LLMRouter and a BudgetGuard.
import { createProvider, LLMRouter, BudgetGuard } from "@mnemosyne/llm-providers";
const provider = createProvider({ name: "openai", apiKey: process.env.OPENAI_API_KEY! });
// or run fully local with Ollama, or route across providers + cap spend:
const router = new LLMRouter(/* … */);
const guard = new BudgetGuard(/* … */);For the HTTP server, the same choice is environment-driven:
MNEMO_LLM_PROVIDER=openai # openai | anthropic | google | cohere | mistral | voyage | ollama
MNEMO_LLM_API_KEY=sk-…
MNEMO_EMBED_MODEL=text-embedding-3-small # defaults to this when unsetIf no provider is configured, POST /v1/recall still works for lexical search
but returns 501 when a request needs server-side embedding.
Next steps
- the architecture spec — the bitemporal model, hybrid recall pipeline, and the Memory Graph.
- the security model — how per-workspace RLS isolation actually holds.
- rendering the graph — visualize a workspace’s memory.