arXiv:2409.09603cs.AIcs.CL2024-09被引 24

为强化学习对齐设计数据评估体系,提升偏好数据使用效率

Towards Data-Centric RLHF: Simple Metrics for Preference Dataset Comparison

  • 从规模、标签噪声、信息量三方面构建数据评估指标
  • 发现不同数据集在标注质量与信息密度上差异显著
  • 适合关注数据质量的RLHF研究者与数据收集团队

将语言模型对齐人类偏好需要能揭示这些偏好的数据。理想情况下,可针对每个下游应用精心收集和定制专属偏好数据。然而实践中,通常仅使用少数几个公开可用的偏好数据集来训练强化学习中的人类反馈(RLHF)的奖励模型。尽管新偏好数据集不断涌现,但目前尚无系统方法用于衡量和比较这些数据集。本文从规模、标签噪声和信息内容三个维度系统研究偏好数据集,提出相应的量化指标,并揭示了理解偏好数据集的新比较维度。本工作是迈向以数据为中心的对齐策略的第一步,为提升RLHF训练效率和迭代数据收集提供支持。

原文摘要 · Abstract (English)

The goal of aligning language models to human preferences requires data that reveal these preferences. Ideally, time and money can be spent carefully collecting and tailoring bespoke preference data to each downstream application. However, in practice, a select few publicly available preference datasets are often used to train reward models for reinforcement learning from human feedback (RLHF). While new preference datasets are being introduced with increasing frequency, there are currently no existing efforts to measure and compare these datasets. In this paper, we systematically study preference datasets through three perspectives: scale, label noise, and information content. We propose specific metrics for each of these perspectives and uncover different axes of comparison for a better understanding of preference datasets. Our work is a first step towards a data-centric approach to alignment by providing perspectives that aid in training efficiency and iterative data collection for RLHF.

RLHF数据评估偏好数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。