提出新指标筛选偏好数据,让大模型更准对齐人类价值观。
Larger or Smaller Reward Margins to Select Preferences for Alignment?
- 用‘对齐潜力’度量评估数据与目标奖励差距,指导优选。
- 实测在不同模型上均优于现有方法,提升对齐效果。
- 适用于自生成数据,规模越大、迭代越多越有效,适合研究者用。
偏好学习对大语言模型与人类价值观对齐至关重要,数据质量起关键作用。现有评估指标多基于显式或隐式奖励差距,但常对同一数据给出矛盾判断。为此,本文提出‘对齐潜力’指标,量化模型当前隐式奖励差距与目标显式奖励之间的距离,从而评估其对齐潜力。实验表明,基于该指标选择的数据训练,能持续提升对齐性能,优于多种基线模型和优化目标。该方法还可拓展至自对弈数据生成框架,在自生成内容中识别高质量数据。在此场景下,该方法在不同训练设置下超越当前最先进水平,并随数据集规模和训练迭代次数增加,对齐性能持续提升。
原文摘要 · Abstract (English)
Preference learning is critical for aligning large language models (LLMs) with human values, with the quality of preference datasets playing a crucial role in this process. While existing metrics primarily assess data quality based on either explicit or implicit reward margins, they often provide contradictory evaluations for the same data. To address this issue, we introduce the alignment potential metric, which quantifies the gap from the model's current implicit reward margin to the target explicit reward margin, thereby estimating the model's potential to align with the preference data. Empirical results demonstrate that training on data selected by this metric consistently enhances alignment performance, surpassing existing metrics across different base models and optimization objectives. Furthermore, our method extends to self-play data generation frameworks, where the metric is used to identify high-quality data within the self-generated content by LLMs. Under this data generation scenario, our method surpasses current state-of-the-art (SOTA) results across various training settings and demonstrates continuous improvements in alignment performance as dataset size and training iterations increase.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。