Agent Memory's benchmarks measure performance across long-horizon sessions where context accumulates; WideSearch shows 61% token savings and 51% higher task pass rate; PersonaMem long-term memory accuracy gains 28 percentage points.
All benchmark results are measured over continuous long-horizon sessions, not isolated turns — for example, SWE-bench runs 50 consecutive tasks per session to simulate the context-accumulation pressure of real-world long-horizon agents.[1] On the WideSearch short-term memory benchmark (integrated with OpenClaw), TencentDB-Agent-Memory cuts token usage by 61.38% (221.31M → 85.64M tokens) and raises task pass rate by 51.52% (33% → 50%) relative to baseline.[1] On the PersonaMem long-term memory benchmark, the plugin raises accuracy from 48% to 76%, a 59% relative improvement.[1] All benchmark results compare Agent Memory against a baseline of the same agent running without memory assistance; reported gains reflect the incremental improvement Agent Memory adds.
Sources