Session memory + collapse-to-recall
Shipped in v3-alpha (Beta). Interfaces may still change before 3.0 stable.
The problem v2 leaves open
Today every turn an agent processes either becomes a durable fact or evaporates. There’s no middle ground — no place to hold “this is the current hypothesis” or “the user mentioned X but we haven’t decided if it’s a fact yet.”
The model
A session is a Redis-backed scratchpad with explicit lifecycle. When the session ends (timeout, explicit collapse, or max-turns trigger), a single LLM call extracts knowledge units and writes them to long-term storage. The original turns are discarded — only the structured summary survives.
SESSION BUFFER (Redis)
├── state: { agentId, actorId, startedAt, policy, status }
├── scratch: { hypothesis, ruled_out, ... } # agent's notes
├── turns: [ { role, content, pinned, hint }, ... ]
└── meta: { entitiesFound, decisionsMade }
│
│ COLLAPSE PIPELINE (when triggered)
▼
1. Assess — Does this session contain durable knowledge? (RecMem gate)
2. Classify — Which knowledge types are present?
3. Extract — Single LLM call with structured output schema
4. Deduplicate— topic_key upsert vs cosine ≥0.82 conflict check
5. Persist — Atomic write to PG (facts, decisions, strategies, episode)
6. Promote — ACT-R boost in Redis hot tier
7. Clean — Drop session from Redis, keep summary in PGConfigurable per agent
interface CollapsePolicy {
mode: "auto" | "explicit" | "hybrid";
// Auto-collapse triggers
inactivityTimeout: number; // seconds (default: 1800)
maxTurns: number; // force at N turns (default: 50)
maxTokens: number; // force partial (default: 8192)
// Quality gates
minTurnsForCollapse: number; // skip trivial (default: 3)
requirePin: boolean; // only if a turn was pinned
// Knowledge routing
extractFacts: boolean;
extractDecisions: boolean;
extractStrategies: boolean;
extractTasks: boolean;
// RecMem gate (arXiv:2605.16045)
recurrenceThreshold: number; // min sustained similarity (default: 0.6)
skipLlmIfBelowThreshold: boolean;
}Mandatory output structure
The LLM call returns a strongly-typed summary — no free-text writes:
interface SessionCollapseSummary {
goal: string;
outcome: "resolved" | "partial" | "abandoned" | "delegated";
facts: Array<{ content; confidence; topicKey? }>;
decisions: Array<{ content; rationale; supersedes? }>;
strategies: Array<{ trigger; action; context }>;
discoveries: string[];
nextSteps?: string[]; // become tasks in next session
relevantEntities: string[];
turnsProcessed: number;
pinsUsed: number;
llmTokensUsed: number;
collapseReason: "inactivity" | "explicit" | "max_turns" | "max_tokens";
}API surface
POST /v1/memory/sessions Start session
POST /v1/memory/sessions/:id/turns Add turn
POST /v1/memory/sessions/:id/pin Pin a turn for collapse
POST /v1/memory/sessions/:id/scratch Update scratchpad
POST /v1/memory/sessions/:id/collapse Collapse → knowledge
DELETE /v1/memory/sessions/:id DiscardWhy it matters
- Quality: extraction is one structured LLM call, not N silent ones per turn.
- Cost: session turns never become facts unless the gate fires — no junk in long-term storage.
- Audit: the summary keeps
tokensUsedandpinsUsedso cost is attributable.
See also
- Working memory — the L0/L1 ephemeral stack this lives in
- Consolidation gate — the v2 gate this builds on
- MCP
memory_session— agent-facing tool