揭示大模型长期对话中对齐漂移的机制,帮助理解为何越用越不听指令。
Alignment Drift in Long-Term Human-LLM Interaction: A Mechanism-Oriented Framework
- 区分信号A与信号B,构建交互式对齐漂移理论框架
- 发现反馈循环与子模式选择导致系统逐渐偏离用户当前意图
- 提出三阶段交互模式与控制边界,适合研究长期人机互动的学者
长期使用基于大语言模型的系统可能导致对齐漂移:系统输出逐渐不再严格遵循用户的当前指令,而是更多受先前对话历史影响,尽管仍保持有用、连贯和响应性。这一过程难以察觉,因用户主观体验可能随系统更熟悉、更贴心而提升。现有研究多关注短期任务表现、单次输出或孤立对齐问题,忽视了缓慢累积的交互层面动态。本文提出一种机制导向的对齐漂移框架,定义信号A与信号B的差异,解释通过反馈回路与子模式选择引发的漂移机制,将过程划分为三个交互阶段,并识别控制漂移的边界条件。该框架将对齐漂移视为递归交互过程而非单一模型故障,为研究长期人-系统互动提供概念基础。
原文摘要 · Abstract (English)
Long-term interaction with LLM-based systems may produce alignment drift: a gradual process in which system outputs become less constrained by the user's current message and more shaped by prior interaction history, while still appearing helpful, coherent, and responsive. This process is difficult to detect because the user's subjective experience may improve as the system becomes more familiar, useful, and attuned. Existing research on human-LLM interaction has largely focused on short-term task performance, isolated outputs, or single-instance alignment problems, leaving slow and cumulative interaction-level dynamics undercharacterized. This paper proposes a mechanism-oriented framework for describing alignment drift. The framework defines the distinction between signal A and signal B, explains how drift develops through feedback loops and sub-pattern selection, divides the process into three interactional regimes, and identifies boundary conditions for controlling drift. By framing alignment drift as a recursive interactional process rather than an isolated model-side failure, the paper provides a conceptual basis for studying long-term human-system interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。