Benchmarks
These are internal targets, not independently verified benchmarks. The figures below come from in-house runs on a single developer machine and have not been reproduced by a third party or audited against the public LongMemEval leaderboard. Competitor numbers are our own measurements and may not reflect each system’s best configuration. Treat everything here as directional. We publish the harness so you can run your own numbers — those, not these, are what you should trust.
LongMemEval
The standard public benchmark for memory in agentic systems.
cd benchmarks/longmemeval
pnpm bench --baseline mnemosyne --compare mem0,zep,letta,cogneeInternal run (June 2026, M2 Pro, local Postgres — not independently verified)
F1 score (higher is better):
Input tokens per session (lower is better — Mnemosyne shrinks the bar by saving tokens):
| System | R@5 | P@5 | F1 | Tokens/session (input) | $ / 1k sessions |
|---|---|---|---|---|---|
| Mem0 | 0.812 | 0.741 | 0.775 | 1820 | $0.49 |
| Zep | 0.851 | 0.799 | 0.824 | 1610 | $0.42 |
| Letta (MemGPT) | 0.802 | 0.756 | 0.778 | 2100 | $0.55 |
| Cognee | 0.834 | 0.781 | 0.806 | 1720 | $0.46 |
| Mnemosyne | 0.879 | 0.831 | 0.854 | 1230 | $0.31 |
In our internal runs Mnemosyne leads on both F1 and cost. The intended cost advantage comes from the zero-LLM features (composition, anticipation, consensus, meta-cognition) that displace LLM round-trips with deterministic computation. As above: these are our own numbers, not independently reproduced.
Recall latency (internal targets)
| Tier | p50 | p95 | p99 |
|---|---|---|---|
| Hot | 2 ms | 4 ms | 8 ms |
| Warm | 8 ms | 14 ms | 22 ms |
| Cold | 65 ms | 140 ms | 195 ms |
Measured in-house against a workspace with 50k facts, 50k embeddings, HNSW
m=16 ef_construction=64. Your hardware and dataset will differ.
Write throughput (internal targets)
| Operation | rps (single replica) |
|---|---|
| facts.create | 480 |
| facts.patch | 720 |
| recall | 1200 |
| timeline (30d) | 380 |
Postgres on the same host (m6i.xlarge equivalent), no pooler. Adding pgBouncer in transaction mode scales these by ~3x for read-heavy workloads.
Memory footprint (internal targets)
| Component | RSS |
|---|---|
| Mnemosyne server (idle) | 95 MB |
| Mnemosyne server (50 rps) | 280 MB |
| Postgres (50k facts) | 1.2 GB (including pgvector index) |
Reproducing locally
git clone https://github.com/lucasmailland/mnemosyne
cd mnemosyne/benchmarks
pnpm install
pnpm bench:longmemeval
pnpm bench:latency
pnpm bench:throughputEach script writes a CSV to out/. Open an issue with your CSV if your
numbers differ materially — your runs are the ones that matter.
We aim to run benchmarks adversarially — Mnemosyne in its weakest configuration (Mode A, no caching, no rerank) against each competitor’s default. The published figures above are still internal and unverified; rerun the harness on your own hardware before relying on any number.