arXiv:2506.08965cs.LGcs.AI2025-06

用少量数据训练出高性能奖励模型,提升强化学习效率

GFRIEND: Generative Few-shot Reward Inference through EfficieNt DPO

  • 通过思维链采样挖掘多样高质量偏好关系
  • 在少样本下达到与大规模数据相当的性能
  • 适合资源有限但需高效训练奖励模型的场景

从人类反馈中进行强化学习(RLHF)时,以少量数据训练高性能奖励模型对提升效率和可扩展性至关重要。本文提出一种数据增强与扩展框架,使在小规模数据上训练的生成式奖励模型达到与大规模数据训练相当的性能。传统方法如直接偏好优化(DPO)受限于样本配对效率低和数据多样性不足。本工作引入偏好精炼机制,采用思维链(CoT)采样发现多样且高质量的偏好关系,并结合困惑度评分机制分配精细偏好等级,再利用多级直接偏好优化(M-DPO)捕捉样本间更细微的偏好差异。实验表明,该方法显著提升数据效率与模型性能,使少样本训练的奖励模型表现媲美大规模数据训练结果。研究凸显了数据高效策略在奖励模型优化中的潜力,为低资源RLHF应用提供可靠解决方案。

原文摘要 · Abstract (English)

The ability to train high-performing reward models with few-shot data is critical for enhancing the efficiency and scalability of Reinforcement Learning from Human Feedback (RLHF). We propose a data augmentation and expansion framework that enables generative reward models trained on small datasets to achieve comparable performance to those trained on large-scale datasets. Traditional methods to train a generative reward model, such as Direct Preference Optimization (DPO), are constrained by inefficiencies in sample pairing and limited data diversity. This work introduces preference refinement, which employs Chain-of-Thought (CoT) sampling to uncover diverse and high-quality preference relationships. It also incorporates a perplexity-based scoring mechanism to assign nuanced preference levels and utilizes Multi-level Direct Preference Optimization (M-DPO) to enable the model to capture finer-grained preference differences between samples. Experimental results demonstrate that the proposed method significantly enhances data efficiency and model performance, enabling reward models trained in a few-shot setting to achieve results on par with those trained on large-scale datasets. This study underscores the potential of data-efficient strategies in advancing reward model optimization, offering a robust solution for low-resource RLHF applications.

奖励模型少样本学习强化学习高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。