Chain-of-thought distillation

Shipped in v3-alpha (Beta). Interfaces may still change before 3.0 stable.

The cost it eliminates

Every time an agent reasons through a problem from scratch, it pays for tokens that produced the same answer last week. Distillation caches the reasoning chain itself, scored by outcome, and returns it when similar problems arrive.

Architecture

Agent query arrives
        ↓
┌─────────────────────────────────────────────────────────────┐
│  1. RESOLVE                                                  │
│  a) Topic-key exact match (O(1), ~1ms)                      │
│  b) Miss → embed query → HNSW search on mnemo_reasoning     │
│  c) If cosine ≥ 0.88 AND success_rate > 0.7 AND worth > 0.5:│
│     Return: { cached_cot, output, suggested_model: haiku }  │
│  d) Else: Return: { action: "execute", similar_attempts: …} │
│  Budget: <20ms topic-key, <50ms vector                       │
└─────────────────────────────────────────────────────────────┘
        ↓ (if executed)
┌─────────────────────────────────────────────────────────────┐
│  2. RECORD (post-execution)                                  │
│  Success: embed input → check ≥0.95 similar → upsert/insert │
│  Init memory_worth counters: { success: 1, fail: 0 }        │
│  Failure: record context for future avoidance                │
└─────────────────────────────────────────────────────────────┘
        ↓ (on subsequent uses)
┌─────────────────────────────────────────────────────────────┐
│  3. FEEDBACK                                                 │
│  Agent reports outcome → update memory_worth counters       │
│  Worth < 0.3 → quarantine                                   │
│  Worth > 0.8 after 50 uses → golden                         │
└─────────────────────────────────────────────────────────────┘

The Memory Worth signal

Caches are gated by a two-counter Wilson confidence interval:

interface ReasoningWorth {
  successCount: number;
  failureCount: number;
  worth: number;       // successCount / (successCount + failureCount)
  confidence: number;  // Wilson 95% lower bound
}
 
// Only serve if:  worth > 0.5  AND  confidence > 0.3 (≥5 uses)
// Converges to true utility after ~10 uses (paper: Spearman ρ=0.89)

API surface

POST /v1/memory/reasoning/resolve     Check cache before execution
POST /v1/memory/reasoning/record      Store successful chain
POST /v1/memory/reasoning/feedback    Report success/failure of cached result

ROI

The doc estimates ~500–2,000 reasoning tokens saved per cache hit. At 30% hit rate on repeat problems, a 100-session/day agent saves ~6M tokens/month.

See also