揭示奖励模型中助人与无害目标的内在冲突机制
Understanding helpfulness and harmless tension in reward models
- 通过激活分析定位助人与无害相关神经元
- 混合目标模型性能低于单一目标,存在显著干扰
- 共享神经元主导行为表现,加剧对齐张力
奖励模型是基于人类反馈的强化学习(RLHF)的关键组件,用于使语言模型同时具备助人与无害的行为特性。然而,这些目标的内部机制及其相互冲突仍不清晰。本文研究了仅助人、仅无害及混合目标设置下训练的奖励模型中的对齐张力。结果发现,混合目标模型常表现不如单一目标模型,表明目标间存在干扰。通过基于激活的方法,我们识别出与各目标相关的神经元,并通过定向消融实验研究其功能。结果显示,这些神经元在支持对应目标的同时,往往对另一目标产生负面影响。此外,大量神经元在助人与无害之间共享,且这些共享神经元对模型行为具有不成比例的影响,是导致对齐张力的主要原因。本研究为对齐目标在奖励模型中的表征方式提供了机制性解释,揭示了多目标对齐持续挑战的原因,推动未来可解耦、可控制的对齐方法发展。
原文摘要 · Abstract (English)
Reward models are a key component of reinforcement learning from human feedback (RLHF), aligning language models toward both helpful and harmless behaviour. However, the internal mechanisms underlying these objectives and their conflicts remain poorly understood. We study alignment tension in reward models trained under helpfulness-only, harmlessness-only, and mixed-objective settings. We find that mixed-objective models often underperform single-objective models, indicating interference between objectives. Using activation-based methods, we identify neurons associated with each objective and study their functional roles via targeted ablations. We find that these neurons causally support their corresponding objectives while often negatively affecting the opposing one. We find that a substantial proportion of neurons are shared between helpfulness and harmlessness, and that these shared neurons exert a disproportionate influence on model behaviour, contributing to alignment tension. Additionally, our results provide insights and mechanistic interpretation into how alignment objectives are represented in reward models and why multi-objective alignment remains challenging, motivating future work on disentangled and controllable alignment methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。