TierMem: Balancing Compressed Memory and Raw Evidence for Long-Horizon Agent Memory
Abstract
Long-horizon agents often rely on compressed memory representations to make long interaction histories tractable at inference time. But compression is inherently lossy: it may discard details that later become necessary for a faithful answer, leaving compressed memory topically relevant yet evidentially insufficient. Systems that cannot return to authoritative raw evidence cannot reliably recover such information at inference time. Always grounding on raw logs avoids this failure mode, but using raw history as the default evidence source is costly and often unnecessary. We therefore argue that long-horizon agent memory should support \emph{inference-time evidence allocation}: for each query, the system should use the cheapest evidence that is still sufficient for faithful and traceable answering. We instantiate this principle in \textbf{TierMem}, a provenance-linked two-tier memory framework with three components: summary-first retrieval, selective escalation to immutable raw logs, and evidence-backed write-back. Across two long-horizon memory benchmarks, TierMem improves the accuracy--efficiency frontier over summary-only memory while substantially reducing the cost of always-raw grounding. On LoCoMo, TierMem reaches \textbf{0.854} accuracy versus \textbf{0.873} for the raw-grounded reference, while reducing average input tokens by \textbf{49.4\%} and latency by \textbf{66.2\%}. On LongMemEval, TierMem reaches \textbf{0.752} versus \textbf{0.808}, while reducing input tokens by \textbf{34.2\%} and latency by \textbf{33.6\%}.