arXiv:2505.20075cs.AI2025-05ACL被引 32

通过难度分级的训练数据,提升智能体反馈强化学习的泛化能力。

Curriculum-RLAIF: Curriculum Alignment with Reinforcement Learning from AI Feedback

  • 按数据难度构建分层训练课程,动态调整样本挑战性。
  • 相比基线方法,策略模型对齐性能显著提升,且推理成本不变。
  • 适合追求高效对齐的AI系统研发者使用。

通过人工智能反馈进行强化学习(RLAIF)训练的奖励模型常因泛化能力有限而影响策略模型的对齐效果。这一问题源于分布偏移、偏好标签噪声以及样本难度与模型能力不匹配等多重因素。本文提出一种以数据为中心的新框架——Curriculum-RLAIF,从数据难度统一视角出发,构建具有不同难度等级的偏好对,并据此生成特定训练课程。大量实验表明,采用Curriculum-RLAIF训练的奖励模型具备更强泛化能力,显著提升策略模型的对齐性能,且相较于多种非课程基线方法,无需额外推理开销。进一步分析与对比验证了该方法在简洁性、效率和有效性上的优势。

原文摘要 · Abstract (English)

Reward models trained through Reinforcement Learning from AI Feedback (RLAIF) methods frequently suffer from limited generalizability, which hinders the alignment performance of policy models. This challenge stems from various issues, including distribution shift, preference label noise, and mismatch of overly challenging samples with model capacity. In this paper, we aim to enhance the generalizability of reward models through a data-centric approach, driven by the insight that these issues are inherently intertwined from a uniform perspective of data difficulty. Accordingly, we propose a novel framework, Curriculum-RLAIF, which constructs preference pairs with varying difficulty levels and then produces a specific curriculum for reward model training. Comprehensive experimental results suggest that reward models trained with Curriculum-RLAIF achieve improved generalizability, boosting the alignment performance of policy models by a significant margin without incurring additional inference costs compared to various existing non-curriculum baselines. Further analysis and comparison with alternative strategies highlight the superiority of Curriculum-RLAIF in simplicity, efficiency, and effectiveness.

强化学习对齐训练奖励建模课程学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。