arXiv:2410.18451cs.AIcs.CL2024-10被引 336

用精简数据提升大模型奖励建模效果,仅8万组数据即达顶尖性能。

Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

  • 通过数据筛选与优化,构建仅含8万组的高质量偏好数据集
  • 基于该数据集训练的模型在RewardBench榜单排名第一
  • 适合追求高效、低成本奖励模型训练的研究者和开发者

本文提出一系列以数据为中心的改进方法,用于增强大语言模型的奖励建模能力。我们设计了有效的数据选择与过滤策略,构建了高质量的开源偏好数据集——Skywork-Reward,仅包含80,000组偏好对,显著小于现有数据集。基于此数据集,我们开发了Skywork-Reward-Gemma-27B和Skywork-Reward-Llama-3.1-8B两个模型系列,其中前者目前位居RewardBench排行榜首位。我们的技术与数据集已直接提升多个主流模型在RewardBench上的表现,充分体现了其在真实偏好学习应用中的实用价值。

原文摘要 · Abstract (English)

In this report, we introduce a collection of methods to enhance reward modeling for LLMs, focusing specifically on data-centric techniques. We propose effective data selection and filtering strategies for curating high-quality open-source preference datasets, culminating in the Skywork-Reward data collection, which contains only 80K preference pairs -- significantly smaller than existing datasets. Using this curated dataset, we developed the Skywork-Reward model series -- Skywork-Reward-Gemma-27B and Skywork-Reward-Llama-3.1-8B -- with the former currently holding the top position on the RewardBench leaderboard. Notably, our techniques and datasets have directly enhanced the performance of many top-ranked models on RewardBench, highlighting the practical impact of our contributions in real-world preference learning applications.

奖励建模数据优化LLM高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。