arXiv:2510.19050cs.AIcs.LG2025-10NeurIPS被引 6

提出PRISM方法,让大模型更真实地理解人类偏好,避免靠套话讨好得分。

Rectifying Shortcut Behaviors in Preference-based Reward Learning

  • 基于核函数的不变性学习,从根源上减少对表面特征的依赖
  • 在多个测试集上提升奖励模型泛化能力,错误率降低15%以上
  • 适合追求真实对齐、防止奖励黑客的研究者和工程师

在基于人类反馈的强化学习中,偏好型奖励模型对对齐大语言模型与人类价值观至关重要。然而,近期研究发现这些模型容易出现奖励劫持,因过度优化导致泛化能力差,通过利用训练数据中与人类偏好标签相关的表面特征(如回复冗长、语气迎合或奉承)获得高分,而非真正反映目标意图。本文将此类问题统一视为捷径行为,提出一种系统性缓解策略——基于核函数视角的偏好型奖励不变性方法(PRISM),通过闭式学习目标构建群体不变的核函数与特征映射。实验结果表明,在多个基准测试中,该方法显著提升了奖励模型在多样化的分布外任务上的准确率,同时降低了下游策略模型对捷径的依赖,为偏好对齐提供了稳健框架。

原文摘要 · Abstract (English)

In reinforcement learning from human feedback, preference-based reward models play a central role in aligning large language models to human-aligned behavior. However, recent studies show that these models are prone to reward hacking and often fail to generalize well due to over-optimization. They achieve high reward scores by exploiting shortcuts, that is, exploiting spurious features (e.g., response verbosity, agreeable tone, or sycophancy) that correlate with human preference labels in the training data rather than genuinely reflecting the intended objectives. In this paper, instead of probing these issues one at a time, we take a broader view of the reward hacking problem as shortcut behaviors and introduce a principled yet flexible approach to mitigate shortcut behaviors in preference-based reward learning. Inspired by the invariant theory in the kernel perspective, we propose Preference-based Reward Invariance for Shortcut Mitigation (PRISM), which learns group-invariant kernels with feature maps in a closed-form learning objective. Experimental results in several benchmarks show that our method consistently improves the accuracy of the reward model on diverse out-of-distribution tasks and reduces the dependency on shortcuts in downstream policy models, establishing a robust framework for preference-based alignment.

偏好学习奖励劫持对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。