Rate limits
Defaults (per API key)
| Operation class | Limit | Burst |
|---|---|---|
| Recall | 600 / minute | 50 |
| Write | 120 / minute | 20 |
| Export | 10 / day | — |
| Admin | 30 / minute | 5 |
| Session | 10 concurrent | — |
| Federation sync | 60 / minute | 10 |
These are per-key — multiple keys on the same workspace each get their own bucket.
Response headers
Every response carries:
X-RateLimit-Limit: 600
X-RateLimit-Remaining: 587
X-RateLimit-Reset: 1717500060 # Unix secondsOn a 429 response we add:
Retry-After: 5 # seconds to waitSelf-hosted: defaults are tunable
When you run Mnemosyne yourself, the limits live in mnemo_api_keys
(per-key) and can be overridden by env (MNEMO_RATE_LIMIT_DEFAULTS).
UPDATE mnemo_api_keys
SET recall_rpm = 6000, write_rpm = 1200
WHERE id = 'key_xxx';Hosted: tiers
If you use the (future) hosted version, the tiers will look something like Free 10 rps / Pro 100 rps / Enterprise unlimited. Nothing on the OSS side enforces these tiers — they are only relevant if you use the hosted service.
Fail-open behaviour
If the rate-limit store (Redis) is unavailable, Mnemosyne fails open — requests are allowed and a warning is logged. We never let infrastructure breakage become a denial of service.