揭示大模型内在与提示触发的价值表达机制差异
Dual Mechanisms of Value Expression: Intrinsic vs. Prompted Values in Large Language Models
- 通过残差流中的价值向量和MLP神经元分析机制
- 两类机制部分共享组件,但各有独特作用
- 内在机制提升响应多样性,提示机制强化指令服从
大型语言模型表达价值观主要有两种方式:(1) 内在表达,反映训练中习得的固有价值;(2) 提示触发表达,由明确提示激发。鉴于其在对齐中的广泛应用,厘清二者底层机制至关重要,尤其在于它们是否高度重叠或依赖不同路径。本文从机制层面分析该问题,采用两种方法:(1) 价值向量,从残差流中提取表示价值机制的特征方向;(2) 价值神经元,贡献于价值向量的MLP神经元。结果表明,内在与提示价值机制部分共享关键组件,可跨语言泛化,并重构模型内部表示中的理论价值关联。然而,每种机制亦具独特成分:内在机制在更多样化的价值场景中激活,促进回应多样性;提示机制则增强指令遵从性,甚至在远距离任务如越狱攻击中生效。
原文摘要 · Abstract (English)
Large language models can express values in two main ways: (1) intrinsic expression, reflecting the model's inherent values learned during training, and (2) prompted expression, elicited by explicit prompts. Given their widespread use in value alignment, it is paramount to clearly understand their underlying mechanisms, particularly whether they mostly overlap (as one might expect) or rely on distinct mechanisms. We analyze this largely understudied problem at the mechanistic level using two approaches: (1) value vectors, feature directions representing value mechanisms extracted from the residual stream, and (2) value neurons, MLP neurons that contribute to value vectors. We demonstrate that intrinsic and prompted value mechanisms partly share common components crucial for inducing value expression, generalizing across languages and reconstructing theoretical inter-value correlations in the model's internal representations. Yet, each mechanism also possesses unique components that fulfill distinct roles. In particular, the intrinsic mechanism activates in more diverse value-related scenarios and promotes response diversity, whereas the prompted mechanism strengthens instruction compliance, taking effect even in distant tasks like jailbreaking.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。