arXiv:2506.03637cs.CLcs.AI2025-06被引 25

让奖励模型听懂自然语言指令,一键切换任务偏好。

RewardAnything: Generalizable Principle-Following Reward Models

  • 用自然语言定义奖励原则,无需重新训练
  • 在新原则下表现超越现有模型,无需微调
  • 适配强化学习对齐,提升大模型实用性

奖励模型(Reward Models, RM)是指导大语言模型优化的核心组件,但通常基于固定偏好数据集训练,仅能适应单一隐式偏好分布,难以应对多样现实需求——如某任务要求简洁回答,另一任务则需详细解释。当前方法依赖收集特定任务的偏好数据并重训,成本高且易引入偏差,限制实际应用。本文提出可泛化的、遵循原则的奖励模型,主张RM应理解并遵守动态提供的自然语言奖励原则,类似大模型的指令遵循能力。为此,我们构建了RABench基准,用于评估RM在多样化原则下的泛化能力。评估显示现有模型泛化能力薄弱。作为解决方案,我们提出RewardAnything,一种专门设计并训练以显式遵循自然语言原则的新型奖励模型。在传统基准上,仅通过指定明确原则即达到最先进性能;在RABench上,其可在不重训的情况下高效适应新原则。此外,RewardAnything可无缝集成至现有强化学习人类反馈(RLHF)流程,并通过案例研究展示了如何仅凭自然语言原则自动、高效对齐大语言模型。

原文摘要 · Abstract (English)

Reward Models, essential for guiding Large Language Model optimization, are typically trained on fixed preference datasets, resulting in rigid alignment to single, implicit preference distributions. This prevents adaptation to diverse real-world needs-from conciseness in one task to detailed explanations in another. The standard practice of collecting task-specific preference data and retraining reward models is resource-intensive, often producing biased rewards, and limits practical application. We introduce generalizable, principle-following reward models. We propose that RMs should understand and adhere to dynamically provided natural language specifications of reward principles, similar to instruction-following in LLMs. To measure this capability, we develop RABench, a comprehensive benchmark for RMs focusing on generalization across diverse principles. Evaluations on RABench reveal poor generalization of current RMs. As a solution, we present RewardAnything, a novel RM designed and trained to explicitly follow natural language principles. We achieve SotA performance with RewardAnything in traditional RM benchmark simply by specifying a well-defined principle, and results on RABench show we excel in adapting to novel principles without retraining. Furthermore, RewardAnything integrates seamlessly with existing RLHF methods and we show by a case study on how to automatically and efficiently align LLMs with only natural language principles.

奖励模型指令遵循RLHF泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。