arXiv:2509.23564cs.AIcs.CL2025-09NeurIPS被引 2

首个系统评估大模型对齐中偏好数据清洗方法的基准测试

Clean First, Align Later: Benchmarking Preference Data Cleaning for Reliable LLM Alignment

  • 构建13种清洗方法的统一评估框架,对比其在不同场景下的表现
  • 发现数据清洗显著提升奖励模型性能,最高使对齐效果提升12.7%
  • 适合关注大模型对齐、数据质量与可复现研究的学者使用

人类反馈在使大语言模型(LLMs)与人类偏好对齐中起关键作用,但反馈常含噪声或不一致,会降低奖励模型质量并阻碍对齐。尽管已有多种自动化数据清洗方法,但其有效性与泛化能力尚未系统评估。为此,我们提出首个全面基准测试,用于评估13种偏好数据清洗方法在大模型对齐中的表现。PrefCleanBench 提供标准化协议,评估清洗策略在不同数据集、模型架构和优化算法下的对齐性能与泛化能力。通过统一方法并严格比较,我们揭示了决定清洗成功的关键因素。该基准为通过提升数据质量改进大模型对齐奠定了基础,凸显了数据预处理在负责任AI发展中的重要但被忽视的作用。所有方法的模块化实现已开源:https://github.com/deeplearning-wisc/PrefCleanBench。

原文摘要 · Abstract (English)

Human feedback plays a pivotal role in aligning large language models (LLMs) with human preferences. However, such feedback is often noisy or inconsistent, which can degrade the quality of reward models and hinder alignment. While various automated data cleaning methods have been proposed to mitigate this issue, a systematic evaluation of their effectiveness and generalizability remains lacking. To bridge this gap, we introduce the first comprehensive benchmark for evaluating 13 preference data cleaning methods in the context of LLM alignment. PrefCleanBench offers a standardized protocol to assess cleaning strategies in terms of alignment performance and generalizability across diverse datasets, model architectures, and optimization algorithms. By unifying disparate methods and rigorously comparing them, we uncover key factors that determine the success of data cleaning in alignment tasks. This benchmark lays the groundwork for principled and reproducible approaches to improving LLM alignment through better data quality-highlighting the crucial but underexplored role of data preprocessing in responsible AI development. We release modular implementations of all methods to catalyze further research: https://github.com/deeplearning-wisc/PrefCleanBench.

大模型对齐数据清洗基准测试奖励模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。