arXiv:2505.23114cs.CL2025-05ACL被引 2

用数据地图筛选高质量偏好数据,33%样本达全量效果

Alignment Data Map for Efficient Preference Data Selection and Diagnosis

  • 通过多方法评估数据对齐分数,结合响应质量与差异性筛选
  • 仅用33%高质低变样本,对齐效果媲美甚至超越全量数据
  • 可识别标签错误,适合需要降本增效的模型对齐团队

人类偏好数据对大语言模型与人类价值观对齐至关重要,但收集成本高、效率低,亟需高效数据选择方法以降低标注成本并保持对齐效果。为此,我们提出Alignment Data Map,一种用于识别和选择有效偏好数据的数据分析工具。首先,通过大语言模型作为裁判、显式奖励模型及基于参考的方法评估偏好数据的对齐分数。Alignment Data Map综合考虑响应质量与响应间变异性。实验发现,仅使用33%表现出高质量且低变异性样本进行训练,在MT-Bench、Evol-Instruct和AlpacaEval上达到与全量数据相当或更优的对齐性能。此外,该工具通过分析标注标签与对齐分数的相关性,可检测潜在标签误标问题,提升标注准确性。代码已开源:https://github.com/01choco/Alignment-Data-Map。

原文摘要 · Abstract (English)

Human preference data is essential for aligning large language models (LLMs) with human values, but collecting such data is often costly and inefficient-motivating the need for efficient data selection methods that reduce annotation costs while preserving alignment effectiveness. To address this issue, we propose Alignment Data Map, a data analysis tool for identifying and selecting effective preference data. We first evaluate alignment scores of the preference data by LLM-as-a-judge, explicit reward model, and reference-based approaches. The Alignment Data Map considers both response quality and inter-response variability based on the alignment scores. From our experimental findings, training on only 33% of samples that exhibit high-quality and low-variability, achieves comparable or superior alignment performance on MT-Bench, Evol-Instruct, and AlpacaEval, compared to training with the full dataset. In addition, Alignment Data Map detects potential label misannotations by analyzing correlations between annotated labels and alignment scores, improving annotation accuracy. The implementation is available at https://github.com/01choco/Alignment-Data-Map.

模型对齐数据筛选偏好学习高效训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。