对话中语言模型的语义表示会剧烈变化,可能让事实变非事实
Linear representations in language models can change dramatically over a conversation
- 通过分析对话中线性语义方向的演变,发现信息表征可随上下文大幅改变
- 事实性信息在对话末尾可能被反转为非事实,且该现象跨模型、跨层均存在
- 适合关注模型动态理解、解释性与可控生成的研究者阅读
语言模型的表征中常存在对应高层次概念的线性方向。本文研究这些表征在(模拟)对话中的动态变化:它们如何在对话过程中沿这些方向演化。我们发现,线性表征在对话中可能发生剧烈变化;例如,初始时作为事实的信息在对话末尾可能被表示为非事实,反之亦然。这种变化具有内容依赖性:对话相关的信息表征易变,而通用信息通常保持稳定。即使在解耦事实性与表面回应模式的维度上,这种变化仍具鲁棒性,且在不同模型家族和模型层级间普遍存在。这些变化不依赖于同策略对话;甚至仅重播由完全不同的模型撰写的对话脚本也能产生类似效果。但若仅在上下文中引入科幻故事且明确标注为虚构,则适应性明显减弱。此外,沿表征方向操控在对话不同阶段会产生截然不同的影响。结果支持一种观点:表征可能随模型根据对话提示扮演特定角色而演变。这给可解释性和可控生成带来挑战——静态特征解读或假设特征范围始终对应特定真实值的探测器可能误导。然而,此类表征动态也为理解模型如何适应上下文开辟了新方向。
原文摘要 · Abstract (English)
Language model representations often contain linear directions that correspond to high-level concepts. Here, we study the dynamics of these representations: how representations evolve along these dimensions within the context of (simulated) conversations. We find that linear representations can change dramatically over a conversation; for example, information that is represented as factual at the beginning of a conversation can be represented as non-factual at the end and vice versa. These changes are content-dependent; while representations of conversation-relevant information may change, generic information is generally preserved. These changes are robust even for dimensions that disentangle factuality from more superficial response patterns, and occur across different model families and layers of the model. These representation changes do not require on-policy conversations; even replaying a conversation script written by an entirely different model can produce similar changes. However, adaptation is much weaker from simply having a sci-fi story in context that is framed more explicitly as such. We also show that steering along a representational direction can have dramatically different effects at different points in a conversation. These results are consistent with the idea that representations may evolve in response to the model playing a particular role that is cued by a conversation. Our findings may pose challenges for interpretability and steering -- in particular, they imply that it may be misleading to use static interpretations of features or directions, or probes that assume a particular range of features consistently corresponds to a particular ground-truth value. However, these types of representational dynamics also point to exciting new research directions for understanding how models adapt to context.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。