SPM SPM Docs

Benchmarks

These first-party results are reproducible from published evaluation manifests. We report bands rather than best runs and exclude degraded-validity runs from anchors. Scorecard date: 2026-08-14.

Headline

Metric Value Scope
Context reduction 99.507% Terminal-Bench v2: 160,509 → 791 tokens per query
Payload gold recall 99.7% LME v2-Small, n = 630 evidence recall
Reader-token savings 8.4× vs full-context replay, 20.2k reader tokens/query (LoCoMo arm)
LoCoMo accuracy 67.9–69.2% full 10 conversations, 1,489 questions
LongMemEval v2-Small 48.1% 451 questions, fixed reader + judge
Agentic SWE-Atlas 9/9 10-checkpoint probe; N ≤ 9, indicative only

How to read these

Method

Benchmarks use frozen-anchor paired evaluations. Each checkpoint has one anchor configuration, and both arms (with-memory and baseline) receive the same canonical evidence stream. Runs missing forced compaction or transport symmetry are retained but marked degraded. Reader and judge models are pinned per anchor; readers cannot change mid-comparison.