MCP toolsmemory_recall

memory_recall

Search your memory for relevant knowledge. Returns facts, decisions and episodes ranked by similarity, recency and Memory Worth. Call this before answering to surface durable preferences, prior commitments, and learned context.

Input

{
  "query": "natural-language question or topic",
  "topK": 5,
  "types": ["fact", "decision", "episode"],
  "includeGraph": false,
  "asOf": "2026-06-07T12:00:00Z"
}
FieldTypeRequiredDefaultNotes
querystringyes—Free-form natural language
topKnumberno51 – 20
typesstring[]noallFilter by primitive type
includeGraphbooleannofalseExpand the hits with 1-hop graph neighbours
asOfstringnonow()Bitemporal time travel (ISO 8601)

Output

{
  "hits": [
    {
      "id": "fact_xxx",
      "type": "fact",
      "statement": "user prefers espresso to filter coffee",
      "score": 0.87,
      "worth": 0.72,
      "validFrom": "2026-06-03T18:30:00Z",
      "attribution": { "source": "user_stated" },
      "graph": null
    }
  ],
  "debug": {
    "embedded": true,
    "tier_hit": "warm",
    "latency_ms": 12,
    "inferences": []
  }
}

When the agent should call it

  • Before answering a question that could reference past context.
  • When the user uses a definite article (“the deploy runbook”, “our pricing”) — the agent shouldn’t guess; it should look it up.
  • Before deciding to ask the user a clarifying question — if the answer is already in memory, asking is annoying.

When the agent should NOT call it

  • Pure arithmetic, fresh news, or code generation that doesn’t reference prior context.
  • Queries about real-time external state (use a web search tool instead).
  • During session collapse — the system handles that internally.

Example

// Tool call
{
  "name": "memory_recall",
  "arguments": {
    "query": "what does Lucas usually drink in the morning?",
    "topK": 3
  }
}
 
// Tool response
{
  "hits": [
    {
      "id": "fact_seed_0",
      "type": "fact",
      "statement": "Lucas prefers espresso to filter coffee in the morning",
      "score": 0.435,
      "worth": 0.5
    }
  ],
  "debug": { "tier_hit": "cold", "latency_ms": 21 }
}

Latency budget

  • Tier 1 (hot): < 5 ms
  • Tier 2 (warm): < 15 ms
  • Tier 3 (cold): < 200 ms

See Three-tier retrieval for the cascade.