Without an embedding API
Maybe you can’t reach OpenAI from your network. Maybe you don’t want to pay for embeddings on every write. Maybe you’re air-gapped. Maybe you just want to try Mnemosyne with the absolute minimum stack.
Mode A works. Mnemosyne ships with no embedding default and recall gracefully runs FTS-only when no embedder is configured.
How recall works in Mode A
Query
↓
Tier 3 cold path
↓
Vector stages → SKIPPED (no embedder)
FTS over text_lemmatized → ALWAYS RUNS
Trust ladder + recency + worth → ALWAYS RUNS
Compose top-K → returnThe pipeline still ranks results by recency × worth × pin-status × FTS score. You lose semantic similarity but you keep textual similarity and all of the cognitive layer.
What you can still do in Mode A
- Create / read / update / pin / forget facts.
- Bitemporal time-travel queries (
asOf). - Entity graph and 1-hop expansion.
- All decisions, episodes, citations, audit.
- All MCP tools.
- All zero-LLM features (consensus, composition, meta-cognition, anticipation). They are SQL-based; embeddings aren’t required.
What you give up
- Cross-phrasing similarity. A query for “morning beverage” won’t match a fact about “coffee” unless the words overlap.
- HNSW speed. FTS over
tsvectoris slower than HNSW at large scale.
When you should pick Mode A
- Privacy-sensitive deployments: nothing leaves the machine.
- Cost-sensitive deployments: zero per-write cost.
- CI / testing: deterministic, fast, no external calls.
- Air-gapped intranet: no outbound network needed.
Switching to Mode B later
When you want vector recall, configure an embedder by setting the provider env vars and restarting:
MNEMO_LLM_PROVIDER=openai
MNEMO_LLM_API_KEY=sk-...Future writes embed inline as they’re created, and reads start using the HNSW index as soon as rows have vectors.
No data is lost. Mode A → Mode B is non-destructive. The vectors are added alongside existing rows, and the search blender starts using them the moment they’re available.