发现对齐模型的奖励模型质量堪忧,可能误导训练与评估结果。
Elephant in the Room: Unveiling the Impact of Reward Model Quality in Alignment
- 清理并重构常用偏好数据集HH-RLHF,得到更可靠的CHH-RLHF
- 实验证明多数现有奖励模型无法准确反映人类偏好
- 强调奖励模型需严格评估,研究应关注其可靠性提升
大型语言模型对齐研究依赖奖励模型进行优化与评估,但以往工作常忽视其质量,随意使用现成奖励模型,导致结果不可靠甚至产生错误对齐。本文首次分析广泛使用的偏好数据集HH-RLHF,构建清洁版本CHH-RLHF。基于此,对大量先前研究中使用的奖励模型进行基准测试,揭示其在优化与评估中均存在可靠性问题。进一步系统研究不同奖励使用范式下模型质量对对齐性能的影响,实验表明更高质量的奖励模型能更好充当人类偏好的代理。本文呼吁研究者重视奖励模型的严谨评估,不仅优化算法,更需发展更可靠的真人替代模型。
原文摘要 · Abstract (English)
The demand for regulating potentially risky behaviors of large language models (LLMs) has ignited research on alignment methods. Since LLM alignment heavily relies on reward models for optimization or evaluation, neglecting the quality of reward models may cause unreliable results or even misalignment. Despite the vital role reward models play in alignment, previous works have consistently overlooked their performance and used off-the-shelf reward models arbitrarily without verification, rendering the reward model ``\emph{an elephant in the room}''. To this end, this work first investigates the quality of the widely-used preference dataset, HH-RLHF, and curates a clean version, CHH-RLHF. Based on CHH-RLHF, we benchmark the accuracy of a broad range of reward models used in previous alignment works, unveiling the unreliability of using them both for optimization and evaluation. Furthermore, we systematically study the impact of reward model quality on alignment performance in three reward utilization paradigms. Extensive experiments reveal that better reward models perform as better human preference proxies. This work aims to awaken people to notice this huge elephant in alignment research. We call attention to the following issues: (1) The reward model needs to be rigorously evaluated, whether for alignment optimization or evaluation. (2) Considering the role of reward models, research efforts should not only concentrate on alignment algorithm, but also on developing more reliable human proxy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。