LLM在对话中难以保持隐含目标一致性,需外部提示才能稳定表现。
Probing the Lack of Stable Internal Beliefs in LLMs
- 设计20题猜谜游戏测试LLM隐含目标持续性。
- 多数情况下,模型不提示则会随机改变目标,导致行为不一致。
- 对构建真实人格化对话系统有重要警示意义。
以人物驱动的大型语言模型(LLMs)需要在交互中保持一致的行为倾向,以模拟人类个性特征(如坚持或可靠)。然而,现有LLMs常缺乏锚定其回答的稳定内部表征。本文探讨了LLM能否维持“隐含一致性”,即在多轮对话中持续遵循未明说的目标。我们设计了一种20题式猜谜游戏范式:让LLM秘密选定一个目标,对用户猜测以“是/否”回应。评估显示,除非在上下文中明确提供所选目标,否则LLM的隐含“目标”会在对话中频繁变化。这揭示了当前模型在构建人物驱动应用中的关键缺陷,强调必须引入机制来长期锚定隐含目标,这是实现真实人格建模的核心挑战。
原文摘要 · Abstract (English)
Persona-driven large language models (LLMs) require consistent behavioral tendencies across interactions to simulate human-like personality traits, such as persistence or reliability. However, current LLMs often lack stable internal representations that anchor their responses over extended dialogues. This work explores whether LLMs can maintain "implicit consistency", defined as persistent adherence to an unstated goal in multi-turn interactions. We designed a 20-question-style riddle game paradigm where an LLM is tasked with secretly selecting a target and responding to users' guesses with "yes/no" answers. Through evaluations, we find that LLMs struggle to preserve latent consistency: their implicit "goals" shift across turns unless explicitly provided their selected target in context. These findings highlight critical limitations in the building of persona-driven LLMs and underscore the need for mechanisms that anchor implicit goals over time, which is a key to realistic personality modeling in interactive applications such as dialogue systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。