大模型处理更新信息时,旧数据比新数据更难遗忘,存在系统性优先效应。
LLMs Remember First, Forget Last: Dual-Process Interference in Large Language Models
- 用认知心理学干扰实验对比新旧信息冲突下的记忆表现
- 所有模型均显示主动干扰远强于回溯干扰(效应量d=1.73)
- 适合关注上下文演化场景中模型可靠性的研究者与开发者
大型语言模型可处理数百万个标记,但其在上下文中处理矛盾信息的方式仍不清晰。从追踪生命体征变化的病历记录到包含后续条款的法律文件,现实应用常需模型从语义相似且相互竞争的连续更新中检索特定值。本文借鉴认知心理学中的经典干扰范式,对比了39个参数规模从10亿到2.5万亿的大型语言模型在回溯干扰(回忆初始值)与主动干扰(在先前编码干扰下回忆最新值)下的表现。结果显示,所有模型均呈现相同模式:主动干扰导致的性能下降远大于回溯干扰(效应量d = 1.73),这与人类典型情况相反,即新信息更易干扰旧记忆。四方面证据表明两类干扰机制不同:模型规模仅预测回溯干扰抵抗能力,不预测主动干扰;两者得分相关性极弱;错误分析显示回溯干扰失败表现为被动检索失败,而主动干扰失败表现为活跃的优先效应侵入;两种条件下幻觉率均低于1%。自适应启动设计进一步揭示,回溯干扰随干扰强度平滑衰减,而主动干扰在高干扰下呈双峰崩溃,符合渐进稀释与胜者通吃注意力机制。这些发现揭示了变压器注意力中的系统性优先偏差,对信息动态演化场景中模型的可靠部署具有重要启示。
原文摘要 · Abstract (English)
Large language models can process millions of tokens, yet how they handle conflicting information within context remains poorly understood. From patient health logs tracking evolving vital signs to legal documents with superseding clauses, real-world applications routinely require models to retrieve specific values from streams of semantically similar, competing updates. We adapt classical interference paradigms from cognitive psychology to compare retroactive interference (RI; recalling initial values after updates) and proactive interference (PI; recalling recent values despite competing prior encodings) across 39 LLMs spanning 1B to 2.5T parameters. Every model exhibits the same pattern: PI causes substantially greater degradation than RI (d = 1.73), the opposite of the typical human finding where new information more readily disrupts old. Four lines of evidence indicate that RI and PI engage distinct mechanisms: model size predicts RI resistance but not PI; the two scores correlate only weakly; error analysis reveals qualitatively different failure profiles, with RI failures reflecting passive retrieval failure and PI failures reflecting active primacy intrusion; and hallucination rates remain below 1% in both conditions. An adaptive bootstrap design further shows that RI decays smoothly while PI collapses bimodally at high interference, consistent with gradual dilution versus winner-take-all attention dynamics. These findings characterize a systematic primacy bias in transformer attention, with implications for reliable deployment in settings where in-context information evolves over time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。