arXiv:2502.17173cs.CLcs.AI2025-02ACL被引 2

构建中文奖励模型评估与训练数据集,提升大模型对中文偏好对齐能力

Cheems: A Practical Guidance for Building and Evaluating Chinese Reward Models from Scratch

  • 提出人类标注的中文奖励模型评测基准CheemsBench
  • 构建大规模人机协作标注的中文偏好数据集CheemsPreference
  • 证明高质量人工监督对中文奖励模型至关重要

奖励模型(RMs)在对齐大语言模型(LLMs)与人类偏好方面至关重要。然而,现有研究多集中于英语,依赖大量合成数据,导致中文场景下的数据集和评测基准有限且不可靠。为此,我们提出CheemsBench——一个完全由人类标注的中文语境下奖励模型评测基准;以及CheemsPreference——一个通过人机协作构建的大规模、多样化的中文偏好数据集,用于支持中文奖励模型训练。我们在CheemsBench上系统评估了开源判别式与生成式奖励模型,发现其在捕捉中文偏好方面存在显著局限。基于CheemsPreference,我们构建的奖励模型在该基准上达到领先性能,证明了人工监督在奖励模型训练中的必要性。结果表明,单纯依赖大规模AI生成数据难以充分捕捉人类偏好,强调高质量人工标注的重要性。

原文摘要 · Abstract (English)

Reward models (RMs) are crucial for aligning large language models (LLMs) with human preferences. However, most RM research is centered on English and relies heavily on synthetic resources, which leads to limited and less reliable datasets and benchmarks for Chinese. To address this gap, we introduce CheemsBench, a fully human-annotated RM evaluation benchmark within Chinese contexts, and CheemsPreference, a large-scale and diverse preference dataset annotated through human-machine collaboration to support Chinese RM training. We systematically evaluate open-source discriminative and generative RMs on CheemsBench and observe significant limitations in their ability to capture human preferences in Chinese scenarios. Additionally, based on CheemsPreference, we construct an RM that achieves state-of-the-art performance on CheemsBench, demonstrating the necessity of human supervision in RM training. Our findings reveal that scaled AI-generated data struggles to fully capture human preferences, emphasizing the importance of high-quality human supervision in RM development.

奖励模型中文NLP人机协作数据标注

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。