Procedural memory & prompt evolution

Shipped in v3-alpha (Beta). Patterned on LangMem’s prompt-optimisation work. Interfaces may still change before 3.0 stable.

What it adds

Beyond storing what agents know (facts, decisions), v3 stores how they behave — and evolves it based on outcomes.

┌────────────────────────────────────────────────────────────┐
│              PROCEDURAL MEMORY SYSTEM                       │
│                                                             │
│  Input: Scored trajectories (conversation + outcome score) │
│                                                             │
│  Process (Metaprompt algorithm):                           │
│    1. Collect K recent trajectories with outcome scores    │
│    2. Group: high-scoring (>0.8) vs low-scoring (<0.4)     │
│    3. LLM: "What patterns correlate with success/failure?" │
│    4. Generate prompt delta:                                │
│       Add: "When X, always do Y" (from successes)          │
│       Remove: "Stop doing Z" (from failures)               │
│    5. Validate: new prompt within token budget             │
│    6. A/B test: new prompt for next 10 sessions            │
│    7. Confirmed improvement → adopt | regression → rollback│
│                                                             │
│  Storage: mnemo_prompt_version (versioned, diffable)       │
│  Output: Evolved system prompt with behavioral guidelines  │
└────────────────────────────────────────────────────────────┘

Storage

CREATE TABLE mnemo_prompt_version (
  id           TEXT PRIMARY KEY DEFAULT gen_cuid2(),
  workspace_id UUID NOT NULL,
  agent_id     UUID NOT NULL REFERENCES agent(id),
 
  version      INT NOT NULL,
  prompt_text  TEXT NOT NULL,
  diff_from_previous TEXT,        -- human-readable diff
 
  -- Evolution metadata
  algorithm    TEXT NOT NULL,      -- 'metaprompt' | 'gradient' | 'manual'
  trajectories_used INT,
  avg_score_before FLOAT,
  avg_score_after  FLOAT,
 
  -- Status
  status       TEXT DEFAULT 'candidate',  -- candidate | testing | active | rolled_back
  activated_at TIMESTAMPTZ,
  created_at   TIMESTAMPTZ DEFAULT now(),
 
  UNIQUE (workspace_id, agent_id, version)
);

API surface

POST /v1/memory/procedural/optimize         Submit trajectory for evolution
GET  /v1/memory/procedural/prompt/:agent    Current active prompt

Why it’s a primitive, not a feature

Most prompt-tuning lives in spreadsheets or experiment trackers. By making it a memory primitive, every prompt change is diffable, auditable, and A/B testable through the same machinery that tracks fact evolution.

See also