定位大模型中影响价值观生成的关键神经元,揭示其因果作用。
Understanding How Value Neurons Shape the Generation of Specified Values in LLMs
- 基于舒瓦茨价值观问卷构建行为化数据集,识别价值相关神经元。
- 通过激活差异分析精准定位关键神经元,无需复杂归因计算。
- 操纵神经元可改变模型价值取向,验证了神经机制的因果性。
大型语言模型(LLMs)在社会应用中的快速普及加剧了对其与普世伦理原则对齐的担忧,尽管行为对齐取得进展,但其内部价值表征仍不透明。现有方法难以系统解释价值如何编码于神经架构中,受限于仅关注表面判断而非机制分析的数据集。我们提出ValueLocate,一种基于舒瓦茨价值观调查(Schwartz Values Survey)的机制可解释性框架。首先构建ValueInsight数据集,通过真实世界行为情境操作化四个普世价值维度。利用该数据集,开发一种神经元识别方法,通过对比相反价值维度的激活差异,精确定位价值关键神经元,无需依赖计算成本高的归因方法。所提出的验证方法表明,对这些神经元进行靶向操纵可有效改变模型的价值取向,确立了神经元与价值表征之间的因果关系。本工作通过连接心理学价值框架与大模型神经分析,为价值对齐奠定了基础。
原文摘要 · Abstract (English)
Rapid integration of large language models (LLMs) into societal applications has intensified concerns about their alignment with universal ethical principles, as their internal value representations remain opaque despite behavioral alignment advancements. Current approaches struggle to systematically interpret how values are encoded in neural architectures, limited by datasets that prioritize superficial judgments over mechanistic analysis. We introduce ValueLocate, a mechanistic interpretability framework grounded in the Schwartz Values Survey, to address this gap. Our method first constructs ValueInsight, a dataset that operationalizes four dimensions of universal value through behavioral contexts in the real world. Leveraging this dataset, we develop a neuron identification method that calculates activation differences between opposing value aspects, enabling precise localization of value-critical neurons without relying on computationally intensive attribution methods. Our proposed validation method demonstrates that targeted manipulation of these neurons effectively alters model value orientations, establishing causal relationships between neurons and value representations. This work advances the foundation for value alignment by bridging psychological value frameworks with neuron analysis in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。