通过神经元级编辑,实现大模型价值观对齐的精准控制。
Controllable Value Alignment in Large Language Models through Neuron-Level Editing
- 定位稀疏的价值相关神经元,在推理时修改激活值以实现对齐。
- 相比现有方法,目标价值对齐更强,性能下降更小,泄露率显著降低。
- 适合需要可解释、可控价值观调整的研究与应用者。
随着大语言模型对人类行为和决策影响扩大,其与人类价值观对齐的重要性日益凸显。然而,现有基于调控的方法存在可控性不足的问题:引导某一价值时往往无意激活其他非目标价值。为此,我们提出价值泄露(value leakage)这一诊断概念,基于舒瓦茨价值观理论构建了归一化度量指标。针对此问题,我们提出NeVA——一种神经元级编辑框架,通过识别稀疏的价值相关神经元并在推理时进行激活值编辑,实现无需参数更新或重训练的精细控制。实验表明,NeVA在增强目标价值对齐的同时,通用能力退化更小,平均泄露显著降低,残余影响主要局限于语义相关的价值类别。整体上,NeVA提供了更可控、可解释的价值对齐机制。
原文摘要 · Abstract (English)
Aligning large language models (LLMs) with human values has become increasingly important as their influence on human behavior and decision-making expands. However, existing steering-based alignment methods suffer from limited controllability: steering a target value often unintentionally activates other, non-target values. To characterize this limitation, we introduce value leakage, a diagnostic notion that captures the unintended activation of non-target values during value steering, along with a normalized leakage metric grounded in Schwartz's value theory. In light of this analysis, we propose NeVA, a neuron-level editing framework for controllable value alignment in LLMs. NeVA identifies sparse, value-relevant neurons and performs inference-time activation editing, enabling fine-grained control without parameter updates or retraining. Experiments show that NeVA achieves stronger target value alignment while incurring smaller performance degradation on general capability. Moreover, NeVA significantly reduces the average leakage, with residual effects largely confined to semantically related value classes. Overall, NeVA offers a more controllable and interpretable mechanism for value alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。