OperationsCost control

Mnemosyne never ships a model or a key — you wire in your own LLM provider and your users pay for exactly what they use. @mnemosyne/llm-providers is the bring-your-own-model layer: seven provider adapters plus an LLMRouter (best provider per task) and a BudgetGuard (hard spend cap).


The seven providers

@mnemosyne/core never imports a vendor SDK directly — it accepts an LLMProvider interface, and this package implements it for seven backends. Cognitive memory has two LLM-shaped needs — embeddings (text → vector) and judgments (fact extraction, conflict resolution, summarization) — and the best provider for each is rarely the same vendor.

ProviderEmbedCompleteStreamingLocal
OpenAIProvider✓✓✓
AnthropicProvider✓✓
GoogleProvider✓✓✓
CohereProvider✓✓✓
MistralProvider✓✓✓
VoyageProvider✓
OllamaProvider✓✓✓

Anthropic does not embed; Voyage is embedding-only; Ollama runs fully local / air-gapped. Routing embed() to a provider that does not support it raises CapabilityNotSupportedError.


LLMRouter — best provider per task

LLMRouter takes a default provider plus a perTask map, then dispatches each call to the right backend. Pass the router straight to @mnemosyne/core — it sees a single LLMProvider.

import {
  LLMRouter,
  OpenAIProvider,
  VoyageProvider,
  AnthropicProvider,
} from "@mnemosyne/llm-providers";
 
const router = new LLMRouter({
  default: new OpenAIProvider({ apiKey: process.env.OPENAI_API_KEY! }),
  perTask: {
    embed: new VoyageProvider({ apiKey: process.env.VOYAGE_API_KEY! }),
    judge: new AnthropicProvider({ apiKey: process.env.ANTHROPIC_API_KEY! }),
    summary: new AnthropicProvider({
      apiKey: process.env.ANTHROPIC_API_KEY!,
      defaultCompleteModel: "claude-3-5-haiku-latest", // cheaper for summaries
    }),
  },
});
 
await router.embed({ texts: ["..."] });                    // → Voyage
await router.complete({ prompt: "...", task: "judge" });   // → Claude
await router.complete({ prompt: "...", task: "summary" }); // → Claude Haiku
await router.complete({ prompt: "..." });                  // → OpenAI (default)

BudgetGuard — hard spend cap

BudgetGuard wraps any provider and enforces a daily USD ceiling and a per-request USD ceiling. Spend is tracked per response via the costUSD field; when a cap is hit the call throws BudgetExceededError — surface it, do not retry.

import { BudgetGuard, OpenAIProvider } from "@mnemosyne/llm-providers";
 
const guarded = new BudgetGuard({
  provider: new OpenAIProvider({ apiKey: process.env.OPENAI_API_KEY! }),
  dailyUSDCap: 25.0,
  perRequestCap: 0.5,
  onSpend: (delta, total) => metrics.gauge("llm_spend_usd", total),
});
 
try {
  await guarded.embed({ texts: ["..."] });
} catch (err) {
  if (err instanceof BudgetExceededError) {
    // surface to user, do NOT retry
  }
}
 
// Reset from a midnight cron / scheduled function
guarded.reset();
OptionPurpose
dailyUSDCapHard ceiling on total spend per day
perRequestCapHard ceiling on the cost of any single request
onSpend(delta, total)Callback after each charge — wire to your metrics
unknownCostPolicyWhat to do when a model’s cost is unknown

The daily counter is reset by calling guarded.reset() — wire it to a midnight cron or scheduled function.

The costUSD field is powered by a static pricing table (src/pricing.ts), with rates current as of 2026-Q2. Models not in the table return costUSD: undefined. BudgetGuard’s default unknownCostPolicy is "block", so a misconfigured model name fails loudly instead of silently overspending.


Combining the two

BudgetGuard works inside a router: wrap individual per-task providers to keep your judgments budget independent from your embeddings budget.

const router = new LLMRouter({
  default: new BudgetGuard({
    provider: new OpenAIProvider({ apiKey: process.env.OPENAI_API_KEY! }),
    dailyUSDCap: 25,
    perRequestCap: 0.5,
    unknownCostPolicy: "block",
  }),
  perTask: {
    judge: new BudgetGuard({
      provider: new AnthropicProvider({ apiKey: process.env.ANTHROPIC_API_KEY! }),
      dailyUSDCap: 10,
    }),
  },
});

Configuration

Every provider accepts the same ProviderConfig — at minimum an apiKey (hosted providers) or a baseURL (Azure / proxy / local emulators; Ollama ignores apiKey):

{
  apiKey?: string;
  baseURL?: string;
  defaultEmbedModel?: string;
  defaultCompleteModel?: string;
  timeoutMs?: number; // default 60_000 (120_000 for Ollama)
}

At the server level, the active provider is selected with MNEMO_LLM_PROVIDER and MNEMO_LLM_API_KEY (the embedding model defaults to text-embedding-3-small, overridable with MNEMO_EMBED_MODEL). Routes that require an LLM return 501 when no provider is configured.