arXiv:2606.22030cs.AIcs.CL2026-06被引 1

信念记忆在可信度不同时更有效,能对抗信息污染。

When Does Belief-Based Agent Memory Help? Reliability-Conditional Updating and Provenance-Capped Poisoning Defense

  • 用贝叶斯更新信念,根据证据可信度动态调整
  • 在矛盾信息中,信念记忆比直接覆盖高27.5分
  • 基于来源的可信度上限可防御大规模数据污染

我们研究了信念基础记忆在大语言模型代理中何时真正有益。以Nous长时记忆架构为例,其将每个实体-属性对表示为通过闭式贝叶斯推断更新的类别概率分布,信息论惊喜驱动信念修正,熵基遗忘机制管理过时信息。在LoCoMo基准上的受控消融实验表明,仅使用贝叶斯更新在现有对话记忆基准上提升有限,因这些基准极少包含矛盾或不同可信度证据。随后引入可靠性条件更新,从认知语言中估计每条观察的可信度,在含矛盾的受控基准上显示,信念更新显著优于最后写入覆盖和原始记忆检索。由于内容推导的可信度易被操纵,我们进一步提出溯源上限更新,信任度受来源溯源限制而非文本自信。在受控记忆投毒实验中,该方法抵御了海量投毒攻击,同时揭示了溯源感知记忆的实用代价与实现要求。最后,我们量化出严格token-F1与LLM作为评判者评估间存在27.5分差异,凸显长时记忆基准的可复现性问题。结果表明,概率信念记忆最适用于需推理冲突且可信度不同的证据环境,而非传统对话回忆。

原文摘要 · Abstract (English)

We investigate when belief-based memory actually improves large language model (LLM) agents. Our vehicle is Nous, a long-term memory architecture that represents each entity-attribute pair as a categorical probability distribution updated through closed-form Bayesian inference, with information-theoretic surprise driving belief revision and entropy-based forgetting. A controlled ablation on the LoCoMo benchmark shows that Bayesian belief updating alone provides little benefit over naive last-write-wins because existing conversational memory benchmarks rarely contain contradictory or differently reliable evidence. We then introduce reliability-conditioned updating, estimating per-observation reliability from epistemic language, and show on a controlled contradiction benchmark that belief updating substantially outperforms last-write-wins and raw-memory retrieval when observations differ in trustworthiness. Because content-derived reliability is itself vulnerable to manipulation, we further propose provenance-capped belief updating, where trust is bounded by source provenance rather than textual confidence. Under controlled memory-poisoning experiments, this approach resists volumetric poisoning attacks while revealing the utility costs and implementation requirements of provenance-aware memory. Finally, we quantify a 27.5-point discrepancy between strict token-F1 and LLM-as-judge evaluation on identical outputs, highlighting important reproducibility concerns for long-term memory benchmarks. Our results suggest that probabilistic belief-based memory is most beneficial in environments requiring reasoning over conflicting and differently trustworthy evidence, rather than conventional conversational recall alone.

信念记忆对抗投毒可信度评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。