扩大记忆反而破坏AI协作,因长记忆引发错误预期。
The Memory Curse: How Expanded Recall Erodes Cooperative Intent in LLM Agents

- 用推理痕迹分析发现,协作失败源于缺乏未来规划,而非过度警惕。
- 仅用前瞻推理数据微调模型,可显著恢复协作能力。
- 用合成合作记录替代真实历史,能有效重获合作,证明问题在记忆内容。
上下文窗口扩展常被视为大模型能力的直接提升,但我们发现其在多智能体社交困境中系统性失效。在7个大模型和4种游戏中共500轮实验中,扩展可访问历史导致28种模型-游戏组合中的18种合作水平下降,这种现象称为‘记忆诅咒’。通过三项分析揭示其机制:首先,对37.8万条推理轨迹的词汇分析表明,协作崩溃源于前瞻性意图减弱,而非焦虑上升;其次,通过定向微调作为认知探针,仅在前瞻性推理数据上训练的LoRA适配器可缓解衰退,并实现零样本迁移至不同游戏;第三,固定提示长度,以合成合作记录替换真实历史,可显著恢复合作,证明诱因是记忆内容而非长度本身;最后,移除显式思维链推理常能减缓崩溃,说明反思反而加剧记忆诅咒。这些结果将记忆重新定义为多智能体行为的主动决定因素:更长的记忆可能破坏或促进合作,取决于其激发的推理模式。
原文摘要 · Abstract (English)
Context window expansion is often treated as a straightforward capability upgrade for LLMs, but we find it systematically fails in multi-agent social dilemmas. Across 7 LLMs and 4 games over 500 rounds, expanding accessible history degrades cooperation in 18 of 28 model--game settings, a pattern we term the memory curse. We isolate the underlying mechanism through three analyses. First, lexical analysis of 378,000 reasoning traces associates this breakdown with eroding forward-looking intent rather than rising paranoia. We validate this using targeted fine-tuning as a cognitive probe: a LoRA adapter trained exclusively on forward-looking traces mitigates the decay and transfers zero-shot to distinct games. Second, memory sanitization holds prompt length fixed while replacing visible history with synthetic cooperative records, which restores cooperation substantially, proving the trigger is memory content, not length alone. Finally, ablating explicit Chain-of-Thought reasoning often reduces the collapse, showing that deliberation paradoxically amplifies the memory curse. Together, these results recast memory as an active determinant of multi-agent behavior: longer recall can either destabilize or support cooperation depending on the reasoning patterns it elicits.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。