测试记忆随无关信息增加而失效的极限,揭示不同模型表现差异。
When Stored Evidence Stops Being Usable: Scale-Conditioned Evaluation of Agent Memory
- 固定任务证据,逐步加入无关会话,观察记忆可靠性变化
- 发现可靠性的下降幅度达16-20个百分点,且受模型规模影响
- 适用于评估大模型在复杂交互中的可扩展性,适合系统设计者
现有记忆代理评估仅报告固定快照下的准确率或召回率,无法反映当与任务无关的会话持续累积时,证据是否仍可用。本文提出一种尺度相关的评估协议,在保持任务证据不变的前提下,逐步增加无关会话,记录代理-记忆轨迹并报告四项诊断指标:预算合规可靠性、尾部调用负担、失败模式分解,以及可靠性低于目标值的可用尺度边界。在LongMemEval和LoCoMo数据集上,针对扁平、平面和分层记忆接口进行测试发现,可靠性下降并非单一现象。在LongMemEval上,HippoRAG虽保持在两调用预算内,但可靠性下降16–20个百分点;LiCoMemory的失败行为强烈依赖于代理,Qwen3-8B超出预算,而Qwen3-32B与Qwen3-235B在测试范围内仍保持可靠。结果支持对可扩展记忆声明应基于代理、接口、尺度范围和交互预算进行条件化。
原文摘要 · Abstract (English)
Memory-agent evaluations report fixed-snapshot accuracy or retrieval quality, but these scores do not show whether evidence remains usable as irrelevant sessions (sessions not annotated as task-relevant evidence for the query) accumulate. We present a scale-conditioned evaluation protocol for agent memory under evidence-preserving growth: for each query, task evidence is held fixed while irrelevant sessions are added. The protocol logs agent--memory trajectories and reports four diagnostics: budget-compliant reliability, tail memory-call burden, failure-regime decomposition, and the usable-scale boundary where reliability falls below the target. Applied to LongMemEval and LoCoMo across flat, planar, and hierarchical memory interfaces, the protocol shows reliability loss is not a single phenomenon. On LongMemEval, HippoRAG stays within the two-call budget but loses 16--20 percentage points in budget-compliant reliability as irrelevant sessions are added; LiCoMemory's observed failures depend strongly on the agent, with Qwen3-8B exceeding the budget while Qwen3-32B and Qwen3-235B remain reliable in the tested range. The result supports a framework for making scalable-memory claims conditional on agent, interface, scale range, and interaction budget.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。