arXiv:2502.21038cs.LGcs.AI2025-02ICLR被引 8

探索多种人类反馈类型在强化学习对齐中的应用潜力

Reward Learning from Multiple Feedback Types

  • 构建六种模拟反馈类型,支持多源反馈学习
  • 在十种环境中验证多类型反馈优于纯偏好基线
  • 为未来人机对齐提供新思路,适合对齐研究者参考

从偏好反馈中学习奖励已成为对齐智能体模型的重要工具。尽管二元比较形式的偏好反馈是获取大规模人工反馈的成熟方法,但人类在其他场景下的反馈更为多样。这些多样化反馈能更好支持标注者目标,且多种反馈源可能相互补充或携带类型特异性偏差。然而,多类型反馈的学习尚未得到充分研究。本文通过生成六种高质量模拟反馈类型,在十种强化学习环境中评估其表现。我们实现了针对所有六种反馈类型的奖励模型与下游强化学习训练,并实证表明多样化反馈可有效提升奖励建模性能。该工作首次提供了多类型反馈在强化学习人类反馈(RLHF)中潜力的有力证据。

原文摘要 · Abstract (English)

Learning rewards from preference feedback has become an important tool in the alignment of agentic models. Preference-based feedback, often implemented as a binary comparison between multiple completions, is an established method to acquire large-scale human feedback. However, human feedback in other contexts is often much more diverse. Such diverse feedback can better support the goals of a human annotator, and the simultaneous use of multiple sources might be mutually informative for the learning process or carry type-dependent biases for the reward learning process. Despite these potential benefits, learning from different feedback types has yet to be explored extensively. In this paper, we bridge this gap by enabling experimentation and evaluating multi-type feedback in a broad set of environments. We present a process to generate high-quality simulated feedback of six different types. Then, we implement reward models and downstream RL training for all six feedback types. Based on the simulated feedback, we investigate the use of types of feedback across ten RL environments and compare them to pure preference-based baselines. We show empirically that diverse types of feedback can be utilized and lead to strong reward modeling performance. This work is the first strong indicator of the potential of multi-type feedback for RLHF.

强化学习奖励学习人机对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。