arXiv:2504.04994cs.CLcs.AI2025-04被引 1

解析大模型中价值观的神经机制,揭示其如何影响决策。

Following the Whispers of Values: Unraveling Neural Mechanisms Behind Value-Oriented Behaviors in LLMs

  • 构建中文价值观基准C-voice,定位编码社会价值的神经元。
  • 通过关闭特定神经元,观察模型行为变化,验证价值影响机制。
  • 适用于研究大模型安全与价值观对齐,尤其关注中文语境。

尽管大语言模型表现优异,但其可能因编码的价值观而产生意外偏见和有害行为,亟需理解其背后的价值机制。现有研究多依赖外部响应评估,缺乏可解释性,且难以在真实场景中衡量社会价值观。本文提出新框架ValueExploration,旨在从神经层面探索大模型中与国家社会价值观相关的行为驱动机制。以中国社会价值观为例,首先构建C-voice——一个大规模双语基准,用于识别和评估大模型中的中文社会价值观。基于C-voice,通过激活差异定位负责编码这些价值观的神经元,并通过抑制这些神经元分析模型行为变化,揭示价值观影响决策的内在机制。在四个代表性大模型上的实验验证了该框架的有效性。相关基准与代码将公开。

原文摘要 · Abstract (English)

Despite the impressive performance of large language models (LLMs), they can present unintended biases and harmful behaviors driven by encoded values, emphasizing the urgent need to understand the value mechanisms behind them. However, current research primarily evaluates these values through external responses with a focus on AI safety, lacking interpretability and failing to assess social values in real-world contexts. In this paper, we propose a novel framework called ValueExploration, which aims to explore the behavior-driven mechanisms of National Social Values within LLMs at the neuron level. As a case study, we focus on Chinese Social Values and first construct C-voice, a large-scale bilingual benchmark for identifying and evaluating Chinese Social Values in LLMs. By leveraging C-voice, we then identify and locate the neurons responsible for encoding these values according to activation difference. Finally, by deactivating these neurons, we analyze shifts in model behavior, uncovering the internal mechanism by which values influence LLM decision-making. Extensive experiments on four representative LLMs validate the efficacy of our framework. The benchmark and code will be available.

大模型价值观对齐神经机制中文

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。