提出无需微调的奖励头编辑方法,提升大模型奖励模型抗欺骗能力。
HARVE: Hacking-Aware Reward-Head Vector Editing for Robust Reward Models
- 通过残差流方向识别多维度欺骗子空间,直接修正奖励头向量。
- 在13种欺骗模式上验证,显著提升8个模型的抗骗鲁棒性。
- 适合关注大模型对齐安全的研究者与开发者使用。
奖励模型是大语言模型对齐的核心组件,但易受奖励欺骗攻击。为评估其鲁棒性,我们构建了涵盖高风险现实场景与通用设置的13种奖励欺骗模式的数据集RewardHackBench,发现8个主流奖励模型在特定子类别上存在严重失效。为此,我们提出HARVE——一种无需微调的标量奖励模型奖励头编辑方法。HARVE从与特定欺骗子类相关的残差流方向中识别多向欺骗子空间,并移除奖励头向量中与此子空间对齐的分量。该方法仅需少量对比性的优质-欺骗样本,无需梯度更新或微调,即可直接降低奖励头对欺骗特征的敏感性。在8个奖励模型上的全面实验表明,HARVE有效提升对抗欺骗的鲁棒性,优于微调基线,且保持模型整体能力。进一步分析表明,奖励欺骗更宜被建模为多维残差空间结构,而非孤立表面线索。
原文摘要 · Abstract (English)
Reward models are central to large language model (LLM) alignment, but they remain vulnerable to reward hacking. To evaluate reward-model robustness, we introduce RewardHackBench containing 13 reward-hacking patterns covering real life high-stakes domains and general settings, and we find severe failures on specific subcategories across eight reward models. To mitigate these failures, we propose HARVE, a training-free reward-head editing method for scalar reward models. Instead of fine-tuning the reward model, HARVE identifies a multi-directional hacking subspace from residual stream directions associated with selected hacking subcategories, and removes the component of the reward-head vector aligned with that subspace. This directly reduces the reward head's sensitivity to hacking-related features using only a small set of contrastive gold-hacked examples, without gradient updates or fine-tuning. Comprehensive experiments across eight reward models indicates that \model improves hacking robustness, outperforms fine-tuning baselines, and preserves reward-models' general capability. Further analyses suggest that reward hacking is better captured as a multidimensional residual-space structure than by isolated surface cues.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。