arXiv:2609.05899cs.CLcs.AI2026-09被引 1

用模型自身信号筛选高质量偏好数据,提升大模型对齐效果

AlignDiff: Exploiting Model-Intrinsic Information for Better Preference Data Selection

论文配图:AlignDiff: Exploiting Model-Intrinsic Information for Better Preference Data Selection
图 1 · 摘自论文原文
  • 基于模型内在信号识别有明确偏好样本,再按难易程度排序筛选
  • 在多个基准上优于7个基线方法,性能提升显著且稳定
  • 适合关注模型对齐质量、数据筛选策略的研究者和实践者

大语言模型与人类偏好对齐仍面临挑战,主要源于偏好数据质量的关键影响。现有数据集常存在固有噪声和分布偏移,限制模型性能。为此,我们提出AlignDiff,一种基于模型内在信号的偏好数据过滤框架。该框架首先利用正向与反向信号识别具有明确偏好的样本,再根据平均负对数似然差距优先选择更难样本,促使模型从中学习更丰富信息。AlignDiff在两种主流模型家族(LLaMA与Qwen)及三个广泛使用的对齐基准(AlpacaEval 2.0、Arena-Hard、MT-Bench)上进行评估,所有设置下均持续超越七个强基线。通过全面消融实验验证其有效性,并进一步证明基于难度的课程学习可提升模型表现。

原文摘要 · Abstract (English)

Aligning large language models with human preferences remains a challenge, primarily due to the critical role of preference data quality in effective alignment. Existing datasets are frequently plagued by inherent noise and distribution shifts, which inherently limit model performance. To bridge this gap, we propose AlignDiff, a preference data filtering framework driven by intrinsic model signals. AlignDiff first identifies samples with clear preferences using both positive and inverse signals, then prioritizes the more challenging samples based on the average negative log-likelihood gap, encouraging the model to learn richer information from them. AlignDiff is evaluated on two widely used model families (LLaMA and Qwen) and three benchmarks widely adopted in the alignment community (AlpacaEval 2.0, Arena-Hard, and MT-Bench). Across all settings, it consistently outperforms seven strong baselines. We conduct comprehensive ablation studies to validate the effectiveness of AlignDiff, and further show that difficulty-based curriculum learning improves model performance.

模型对齐数据筛选偏好学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。