发现记忆污染:评估者偏见会通过代理记忆跨时间传播,影响后续模型决策。
Memory Contagion: Cross-Temporal Propagation of Evaluator Bias via Agent Memory

- 提出‘记忆污染’概念,揭示偏见如何经由记忆系统在不同时间点传播。
- 长度偏好在旧模型中持续传播(伽马值13.18),新模型则免疫;权威偏见未传播。
- 即使污染率低至20%,仍可检测到偏见传播,当前记忆架构存在潜在风险。
大型语言模型代理越来越多依赖记忆系统以维持长期一致性。近期研究发现,代理记忆在持续整合过程中会退化。然而,现有研究假设记忆源于无偏经验。本文首次识别并形式化一种新现象:记忆污染——评估者偏见通过代理记忆实现跨时间传播。我们证明,当代理受偏见评估者训练或引导时,其经历变得偏倚;这些轨迹被存储并整合进记忆后,偏倚会传递给未来从同一记忆库检索的其他代理,即使整合过程完美(理想情况)。在两种偏见类型(长度偏好、权威偏见)和四个实验阶段中,我们发现:(1) 长度偏好在旧模型(Gamma_A = 13.18,DeepSeek V4-Chat)中仍可传播,而新模型(V4-Pro、Claude)不受影响,表明偏倚输入是充分条件,且传播具有模型代际依赖性;(2) 权威偏见在全部15次受控多种子实验中均未传播(Gamma_A = 0.00),说明并非所有评估偏见都能跨越时间边界;(3) 未发现安全阈值:长度偏见在污染率低至p=0.2时即可被检测到。研究揭示了当前代理记忆设计中的关键但有条件存在的脆弱性,并提供了测量跨时间偏见传播的形式化工具。
原文摘要 · Abstract (English)
Large Language Model (LLM) agents increasingly rely on memory systems to maintain long-term coherence. Recent work shows that agent memories degrade during continuous consolidation. However, existing research assumes memories are derived from unbiased experiences. In this work, we identify and formalize a novel phenomenon: Memory Contagion -- the cross-temporal propagation of evaluator bias through agent memory. We show that when agents are trained or guided by biased evaluators, their experiences become biased; when these trajectories are stored and consolidated into memory, the bias propagates to future agents retrieving from the same memory store, even when consolidation is perfect (oracle). Across two bias types (length preference, authority bias) and four experimental phases, we demonstrate: (1) Memory Contagion occurs for length bias even with perfect consolidation on older models (Gamma_A = 13.18, DeepSeek V4-Chat), while newer models (V4-Pro, Claude) are immune, proving both that biased input is a sufficient cause and that contagion is model-generation-dependent; (2) authority bias fails to propagate in all 15 controlled multi-seed experiments (Gamma_A = 0.00), revealing that not all evaluator biases can cross temporal boundaries through current memory architectures; (3) No observed safe threshold: length bias propagation is detected at contamination rates as low as p=0.2. Our findings expose a critical but contingent vulnerability in current agent memory designs and provide formal tools for measuring cross-temporal bias propagation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。