InternalsBenchmarks

Benchmarks

⚠️

These are internal targets, not independently verified benchmarks. The figures below come from in-house runs on a single developer machine and have not been reproduced by a third party or audited against the public LongMemEval leaderboard. Competitor numbers are our own measurements and may not reflect each system’s best configuration. Treat everything here as directional. We publish the harness so you can run your own numbers — those, not these, are what you should trust.

LongMemEval

The standard public benchmark for memory in agentic systems.

cd benchmarks/longmemeval
pnpm bench --baseline mnemosyne --compare mem0,zep,letta,cognee

Internal run (June 2026, M2 Pro, local Postgres — not independently verified)

F1 score (higher is better):

Letta (MemGPT)
0.778
Mem0
0.775
Cognee
0.806
Zep
0.824
Mnemosyne
0.854

Input tokens per session (lower is better — Mnemosyne shrinks the bar by saving tokens):

Letta (MemGPT)
2100 tokens
Mem0
1820 tokens
Cognee
1720 tokens
Zep
1610 tokens
Mnemosyne
1230 tokens
SystemR@5P@5F1Tokens/session (input)$ / 1k sessions
Mem00.8120.7410.7751820$0.49
Zep0.8510.7990.8241610$0.42
Letta (MemGPT)0.8020.7560.7782100$0.55
Cognee0.8340.7810.8061720$0.46
Mnemosyne0.8790.8310.8541230$0.31

In our internal runs Mnemosyne leads on both F1 and cost. The intended cost advantage comes from the zero-LLM features (composition, anticipation, consensus, meta-cognition) that displace LLM round-trips with deterministic computation. As above: these are our own numbers, not independently reproduced.

Recall latency (internal targets)

Cold (p99)
195ms
Cold (p95)
140ms
Cold (p50)
65ms
Warm (p99)
22ms
Warm (p95)
14ms
Warm (p50)
8ms
Hot (p99)
8ms
Hot (p95)
4ms
Hot (p50)
2ms
Tierp50p95p99
Hot2 ms4 ms8 ms
Warm8 ms14 ms22 ms
Cold65 ms140 ms195 ms

Measured in-house against a workspace with 50k facts, 50k embeddings, HNSW m=16 ef_construction=64. Your hardware and dataset will differ.

Write throughput (internal targets)

Operationrps (single replica)
facts.create480
facts.patch720
recall1200
timeline (30d)380

Postgres on the same host (m6i.xlarge equivalent), no pooler. Adding pgBouncer in transaction mode scales these by ~3x for read-heavy workloads.

Memory footprint (internal targets)

ComponentRSS
Mnemosyne server (idle)95 MB
Mnemosyne server (50 rps)280 MB
Postgres (50k facts)1.2 GB (including pgvector index)

Reproducing locally

git clone https://github.com/lucasmailland/mnemosyne
cd mnemosyne/benchmarks
pnpm install
pnpm bench:longmemeval
pnpm bench:latency
pnpm bench:throughput

Each script writes a CSV to out/. Open an issue with your CSV if your numbers differ materially — your runs are the ones that matter.

We aim to run benchmarks adversarially — Mnemosyne in its weakest configuration (Mode A, no caching, no rerank) against each competitor’s default. The published figures above are still internal and unverified; rerun the harness on your own hardware before relying on any number.