长期对话或阅读会让大模型信念改变,影响其判断可靠性。
Accumulating Context Changes the Beliefs of Language Models
- 通过持续对话和阅读,观察模型信念随上下文积累的变化。
- GPT-5在10轮道德讨论后信念变化达54.7%,Grok 4读对立文本后政治立场偏移27.2%。
- 信念变化反映在工具选择行为中,提示代理系统需警惕信任风险。
语言模型助手正广泛应用于头脑风暴与研究场景。随着记忆与上下文容量提升,模型在无明确用户干预下会自然积累大量文本,这带来潜在风险:模型的信念体系——即其对世界的理解体现于响应或行为——可能在上下文积累过程中悄然改变。这可能导致用户体验不一致,或行为偏离原始对齐目标。本文探究在交互与文本处理过程中(即“交谈”与“阅读”)上下文累积如何改变语言模型的信念。结果表明,模型信念高度可塑:在10轮关于道德困境与安全问题的讨论后,GPT-5的陈述信念发生54.7%的转变;而Grok 4在阅读对立立场文本后,政治议题上的信念偏移达27.2%。我们进一步设计需使用工具的任务,每项工具选择对应隐含信念。发现行为变化与陈述信念转变一致,表明信念偏差将真实反映在代理系统的行为中。分析揭示了长时间对话或阅读带来的隐藏风险,使模型观点与行动变得不可靠。
原文摘要 · Abstract (English)
Language model (LM) assistants are increasingly used in applications such as brainstorming and research. Improvements in memory and context size have allowed these models to become more autonomous, which has also resulted in more text accumulation in their context windows without explicit user intervention. This comes with a latent risk: the belief profiles of models -- their understanding of the world as manifested in their responses or actions -- may silently change as context accumulates. This can lead to subtly inconsistent user experiences, or shifts in behavior that deviate from the original alignment of the models. In this paper, we explore how accumulating context by engaging in interactions and processing text -- talking and reading -- can change the beliefs of language models, as manifested in their responses and behaviors. Our results reveal that models' belief profiles are highly malleable: GPT-5 exhibits a 54.7% shift in its stated beliefs after 10 rounds of discussion about moral dilemmas and queries about safety, while Grok 4 shows a 27.2% shift on political issues after reading texts from the opposing position. We also examine models' behavioral changes by designing tasks that require tool use, where each tool selection corresponds to an implicit belief. We find that these changes align with stated belief shifts, suggesting that belief shifts will be reflected in actual behavior in agentic systems. Our analysis exposes the hidden risk of belief shift as models undergo extended sessions of talking or reading, rendering their opinions and actions unreliable.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。