arXiv:2505.11475cs.CLcs.AI2025-05NeurIPS被引 64

开源4万+高质量人类标注偏好数据,支持多任务多语言模型对齐。

HelpSteer3-Preference: Open Human-Annotated Preference Data across Diverse Tasks and Languages

  • 构建跨领域多语言的高质人工标注偏好数据集
  • 训练的奖励模型在RM-Bench和JudgeBench上分别达82.4%和73.7%
  • 支持生成式奖励模型与强化学习对齐,适用于多场景应用

偏好数据集对训练通用领域、指令跟随型语言模型至关重要,尤其依赖基于人类反馈的强化学习(RLHF)。随着每次数据发布,人们对数据质量与多样性的期待不断提升,亟需持续提升开放偏好数据的水平。为此,我们推出HelpSteer3-Preference:一个采用宽松许可协议(CC-BY-4.0)的高质量、人工标注偏好数据集,包含超过40,000个样本,覆盖大语言模型在真实世界中的多样化应用场景,包括科学、技术、工程、数学(STEM)、编程及多语言任务。利用该数据集,我们训练的奖励模型在RM-Bench上达到82.4%,在JudgeBench上达到73.7%,相比此前最佳结果提升约10个百分点。我们还展示了该数据集可用于训练生成式奖励模型,并实现策略模型通过我们的奖励模型进行RLHF对齐。数据集(CC-BY-4.0):https://huggingface.co/datasets/nvidia/HelpSteer3#preference;模型(NVIDIA开放模型):https://huggingface.co/collections/nvidia/reward-models-68377c5955575f71fcc7a2a3

原文摘要 · Abstract (English)

Preference datasets are essential for training general-domain, instruction-following language models with Reinforcement Learning from Human Feedback (RLHF). Each subsequent data release raises expectations for future data collection, meaning there is a constant need to advance the quality and diversity of openly available preference data. To address this need, we introduce HelpSteer3-Preference, a permissively licensed (CC-BY-4.0), high-quality, human-annotated preference dataset comprising of over 40,000 samples. These samples span diverse real-world applications of large language models (LLMs), including tasks relating to STEM, coding and multilingual scenarios. Using HelpSteer3-Preference, we train Reward Models (RMs) that achieve top performance on RM-Bench (82.4%) and JudgeBench (73.7%). This represents a substantial improvement (~10% absolute) over the previously best-reported results from existing RMs. We demonstrate HelpSteer3-Preference can also be applied to train Generative RMs and how policy models can be aligned with RLHF using our RMs. Dataset (CC-BY-4.0): https://huggingface.co/datasets/nvidia/HelpSteer3#preference Models (NVIDIA Open Model): https://huggingface.co/collections/nvidia/reward-models-68377c5955575f71fcc7a2a3

偏好数据多语言奖励模型对齐训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。