让奖励模型自己发现并修正错误,提升对对抗攻击的鲁棒性。
Teach a Reward Model to Correct Itself: Reward Guided Adversarial Failure Discovery for Robust Reward Modeling
- 用奖励模型引导生成错误样本,自动发现失效模式。
- 在HH和Beavertails数据集上显著提升鲁棒性,不损失奖励质量。
- 适合需要高可靠性奖励模型的对齐训练场景。
奖励建模(RM)通过捕捉人类偏好来对齐大语言模型,广泛应用于模型微调、响应过滤和排序任务。然而,由于人类偏好的内在复杂性及数据集覆盖有限,奖励模型在分布外或对抗扰动下常出现失效。现有方法通常依赖对偏好分布或失败特征的先验知识,限制了其在真实场景中的实用性。本文提出一种无需偏好分布先验的可计算方法,通过奖励引导的受控解码发现奖励模型失效模式。基于此,我们构建REFORM框架,利用奖励模型自身生成被错误评分的响应,作为对抗样本扩充训练数据,修复模型偏差。在Anthropic Helpful Harmless(HH)和PKU Beavertails两个常用偏好数据集上评估表明,REFORM在不牺牲奖励质量的前提下显著提升鲁棒性,且在直接评估和下游策略训练中保持性能,同时减少虚假相关性,改善对齐质量。
原文摘要 · Abstract (English)
Reward modeling (RM), which captures human preferences to align large language models (LLMs), is increasingly employed in tasks such as model finetuning, response filtering, and ranking. However, due to the inherent complexity of human preferences and the limited coverage of available datasets, reward models often fail under distributional shifts or adversarial perturbations. Existing approaches for identifying such failure modes typically rely on prior knowledge about preference distributions or failure attributes, limiting their practicality in real-world settings where such information is unavailable. In this work, we propose a tractable, preference-distribution agnostic method for discovering reward model failure modes via reward guided controlled decoding. Building on this, we introduce REFORM, a self-improving reward modeling framework that enhances robustness by using the reward model itself to guide the generation of falsely scored responses. These adversarial examples are then used to augment the training data and patch the reward model's misaligned behavior. We evaluate REFORM on two widely used preference datasets Anthropic Helpful Harmless (HH) and PKU Beavertails and demonstrate that it significantly improves robustness without sacrificing reward quality. Notably, REFORM preserves performance both in direct evaluation and in downstream policy training, and further improves alignment quality by removing spurious correlations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。