Mnemosyne never ships a model or a key — you wire in your own LLM provider and
your users pay for exactly what they use. @mnemosyne/llm-providers is the
bring-your-own-model layer: seven provider adapters plus an LLMRouter (best
provider per task) and a BudgetGuard (hard spend cap).
The seven providers
@mnemosyne/core never imports a vendor SDK directly — it accepts an
LLMProvider interface, and this package implements it for seven backends.
Cognitive memory has two LLM-shaped needs — embeddings (text → vector) and
judgments (fact extraction, conflict resolution, summarization) — and the
best provider for each is rarely the same vendor.
| Provider | Embed | Complete | Streaming | Local |
|---|---|---|---|---|
OpenAIProvider | ✓ | ✓ | ✓ | |
AnthropicProvider | ✓ | ✓ | ||
GoogleProvider | ✓ | ✓ | ✓ | |
CohereProvider | ✓ | ✓ | ✓ | |
MistralProvider | ✓ | ✓ | ✓ | |
VoyageProvider | ✓ | |||
OllamaProvider | ✓ | ✓ | ✓ |
Anthropic does not embed; Voyage is embedding-only; Ollama runs fully local /
air-gapped. Routing embed() to a provider that does not support it raises
CapabilityNotSupportedError.
LLMRouter — best provider per task
LLMRouter takes a default provider plus a perTask map, then dispatches each
call to the right backend. Pass the router straight to @mnemosyne/core — it
sees a single LLMProvider.
import {
LLMRouter,
OpenAIProvider,
VoyageProvider,
AnthropicProvider,
} from "@mnemosyne/llm-providers";
const router = new LLMRouter({
default: new OpenAIProvider({ apiKey: process.env.OPENAI_API_KEY! }),
perTask: {
embed: new VoyageProvider({ apiKey: process.env.VOYAGE_API_KEY! }),
judge: new AnthropicProvider({ apiKey: process.env.ANTHROPIC_API_KEY! }),
summary: new AnthropicProvider({
apiKey: process.env.ANTHROPIC_API_KEY!,
defaultCompleteModel: "claude-3-5-haiku-latest", // cheaper for summaries
}),
},
});
await router.embed({ texts: ["..."] }); // → Voyage
await router.complete({ prompt: "...", task: "judge" }); // → Claude
await router.complete({ prompt: "...", task: "summary" }); // → Claude Haiku
await router.complete({ prompt: "..." }); // → OpenAI (default)BudgetGuard — hard spend cap
BudgetGuard wraps any provider and enforces a daily USD ceiling and a
per-request USD ceiling. Spend is tracked per response via the costUSD
field; when a cap is hit the call throws BudgetExceededError — surface it, do
not retry.
import { BudgetGuard, OpenAIProvider } from "@mnemosyne/llm-providers";
const guarded = new BudgetGuard({
provider: new OpenAIProvider({ apiKey: process.env.OPENAI_API_KEY! }),
dailyUSDCap: 25.0,
perRequestCap: 0.5,
onSpend: (delta, total) => metrics.gauge("llm_spend_usd", total),
});
try {
await guarded.embed({ texts: ["..."] });
} catch (err) {
if (err instanceof BudgetExceededError) {
// surface to user, do NOT retry
}
}
// Reset from a midnight cron / scheduled function
guarded.reset();| Option | Purpose |
|---|---|
dailyUSDCap | Hard ceiling on total spend per day |
perRequestCap | Hard ceiling on the cost of any single request |
onSpend(delta, total) | Callback after each charge — wire to your metrics |
unknownCostPolicy | What to do when a model’s cost is unknown |
The daily counter is reset by calling guarded.reset() — wire it to a midnight
cron or scheduled function.
The costUSD field is powered by a static pricing table (src/pricing.ts),
with rates current as of 2026-Q2. Models not in the table return
costUSD: undefined. BudgetGuard’s default unknownCostPolicy is "block",
so a misconfigured model name fails loudly instead of silently overspending.
Combining the two
BudgetGuard works inside a router: wrap individual per-task providers to keep
your judgments budget independent from your embeddings budget.
const router = new LLMRouter({
default: new BudgetGuard({
provider: new OpenAIProvider({ apiKey: process.env.OPENAI_API_KEY! }),
dailyUSDCap: 25,
perRequestCap: 0.5,
unknownCostPolicy: "block",
}),
perTask: {
judge: new BudgetGuard({
provider: new AnthropicProvider({ apiKey: process.env.ANTHROPIC_API_KEY! }),
dailyUSDCap: 10,
}),
},
});Configuration
Every provider accepts the same ProviderConfig — at minimum an apiKey (hosted
providers) or a baseURL (Azure / proxy / local emulators; Ollama ignores
apiKey):
{
apiKey?: string;
baseURL?: string;
defaultEmbedModel?: string;
defaultCompleteModel?: string;
timeoutMs?: number; // default 60_000 (120_000 for Ollama)
}At the server level, the active provider is selected with MNEMO_LLM_PROVIDER
and MNEMO_LLM_API_KEY (the embedding model defaults to
text-embedding-3-small, overridable with MNEMO_EMBED_MODEL). Routes that
require an LLM return 501 when no provider is configured.