arXiv:2507.01352cs.CLcs.AI2025-07被引 182

用人类与AI协作构建4000万条高质量偏好数据,训练出更强的开源奖励模型。

Skywork-Reward-V2: Scaling Preference Data Curation via Human-AI Synergy

  • 设计人机协同两阶段流程,结合人工标注精度与大模型规模化优势。
  • 在2600万条精选数据上训练的8个版本模型,在7个基准测试中均达顶尖水平。
  • 适合关注模型对齐、安全性和公平性的研究者与开发者使用。

尽管奖励模型(RMs)在基于人类反馈的强化学习(RLHF)中至关重要,但当前最先进的开源奖励模型在多数评估基准上表现不佳,难以捕捉细微的人类偏好。我们推测其脆弱性主要源于偏好数据集的局限性:范围狭窄、合成标注或缺乏严格质量控制。为此,我们提出了包含4000万条偏好对的大规模数据集SynPref-40M。为实现规模化数据构建,设计了人机协同的两阶段流水线,利用人类标注质量与大语言模型(LLM)的可扩展性互补优势:人类提供验证标注,大模型在人类指导下自动筛选与清洗数据。基于该数据混合集,训练出涵盖0.6B至8B参数的八个奖励模型——Skywork-Reward-V2。实证表明,该系列模型在对齐人类偏好、客观正确性、安全性、抵抗风格偏见及Best-of-N扩展能力方面均具广泛适应性,在七个主流奖励模型基准上达到当前最优性能,超越生成式奖励模型,并展现出优异下游表现。消融实验确认,性能提升不仅来自数据量,更得益于高质量的协同标注过程。该工作标志着开源奖励模型的重大进展,证明人机协同能显著提升数据质量。

原文摘要 · Abstract (English)

Despite the critical role of reward models (RMs) in Reinforcement Learning from Human Feedback (RLHF), current state-of-the-art open RMs perform poorly on most existing evaluation benchmarks, failing to capture nuanced human preferences. We hypothesize that this brittleness stems primarily from limitations in preference datasets, which are often narrowly scoped, synthetically labeled, or lack rigorous quality control. To address these challenges, we present SynPref-40M, a large-scale preference dataset comprising 40 million preference pairs. To enable data curation at scale, we design a human-AI synergistic two-stage pipeline that leverages the complementary strengths of human annotation quality and AI scalability. In this pipeline, humans provide verified annotations, while LLMs perform automatic curation based on human guidance. Training on this preference mixture, we introduce Skywork-Reward-V2, a suite of eight reward models ranging from 0.6B to 8B parameters, trained on a carefully curated subset of 26 million preference pairs from SynPref-40M. We demonstrate that Skywork-Reward-V2 is versatile across a wide range of capabilities, including alignment with human preferences, objective correctness, safety, resistance to stylistic biases, and best-of-N scaling. These reward models achieve state-of-the-art performance across seven major reward model benchmarks, outperform generative reward models, and demonstrate strong downstream performance. Ablation studies confirm that effectiveness stems not only from data scale but also from high-quality curation. The Skywork-Reward-V2 series represents substantial progress in open reward models, demonstrating how human-AI curation synergy can unlock significantly higher data quality.

奖励模型人机协同数据构建RLHF

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。