arXiv:2503.22230cs.LG2025-03NeurIPS被引 32

通过优化数据构建提升人类反馈强化学习效果

Exploring Data Scaling Trends and Effects in Reinforcement Learning from Human Feedback

  • 融合推理验证器与生成奖励模型,缓解奖励滥用
  • 优先训练数学编码任务,显著提升整体性能
  • 提出预PPO提示选择法,保持响应多样性

人类反馈强化学习(RLHF)对对齐大语言模型与人类偏好至关重要。现有研究多关注算法改进,却忽视了提示数据构建的重要性。本文探究了RLHF性能扩展中的数据瓶颈,特别是奖励滥用和响应多样性下降问题。提出混合奖励系统,结合推理任务验证器(RTV)与生成式奖励模型(GenRM),有效缓解奖励滥用。设计新型提示选择方法Pre-PPO,维持响应多样性并提升学习效率。实验表明,优先在训练初期处理数学与编程任务可显著改善性能。不同方法对比显示:RTV对奖励滥用最具鲁棒性,次为含真实标签的GenRM,最后为基于SFT Best-of-N响应的GenRM。所提策略能快速捕捉任务细微差异,大幅提高整体RLHF表现。本工作强调数据构建的关键作用,并提供可落地的解决方案。

原文摘要 · Abstract (English)

Reinforcement Learning from Human Feedback (RLHF) is crucial for aligning large language models with human preferences. While recent research has focused on algorithmic improvements, the importance of prompt-data construction has been overlooked. This paper addresses this gap by exploring data-driven bottlenecks in RLHF performance scaling, particularly reward hacking and decreasing response diversity. We introduce a hybrid reward system combining reasoning task verifiers (RTV) and a generative reward model (GenRM) to mitigate reward hacking. We also propose a novel prompt-selection method, Pre-PPO, to maintain response diversity and enhance learning effectiveness. Additionally, we find that prioritizing mathematical and coding tasks early in RLHF training significantly improves performance. Experiments across two model sizes validate our methods' effectiveness and scalability. Results show that RTV is most resistant to reward hacking, followed by GenRM with ground truth, and then GenRM with SFT Best-of-N responses. Our strategies enable rapid capture of subtle task-specific distinctions, leading to substantial improvements in overall RLHF performance. This work highlights the importance of careful data construction and provides practical methods to overcome performance barriers in RLHF.

强化学习人类反馈数据构建语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。