arXiv:2508.04149cs.CLcs.AI2025-08被引 7

按难易度选偏好数据,用10%样本实现更好对齐效果

Difficulty-Based Preference Data Selection by DPO Implicit Reward Gap

  • 基于DPO隐式奖励差距筛选难样本
  • 仅用10%数据即超越多个基线方法
  • 适合资源有限时高效对齐大模型

对齐大语言模型与人类偏好是人工智能研究的关键挑战。尽管强化学习人类反馈(RLHF)和直接偏好优化(DPO)广泛应用,但它们通常依赖大规模、高成本的偏好数据集。当前缺乏针对偏好数据的高质量数据选择方法。本文提出一种基于难度的数据选择策略,基于DPO隐式奖励机制。通过选择DPO隐式奖励差距较小的样本(表明更难案例),提升数据效率与模型对齐效果。该方法在多个数据集和对齐任务中持续优于五个强基线,在仅使用原数据10%的情况下实现更优性能。这一原理清晰、高效的选数方法为资源受限下的大模型对齐提供了可行方案。

原文摘要 · Abstract (English)

Aligning large language models (LLMs) with human preferences is a critical challenge in AI research. While methods like Reinforcement Learning from Human Feedback (RLHF) and Direct Preference Optimization (DPO) are widely used, they often rely on large, costly preference datasets. The current work lacks methods for high-quality data selection specifically for preference data. In this work, we introduce a novel difficulty-based data selection strategy for preference datasets, grounded in the DPO implicit reward mechanism. By selecting preference data examples with smaller DPO implicit reward gaps, which are indicative of more challenging cases, we improve data efficiency and model alignment. Our approach consistently outperforms five strong baselines across multiple datasets and alignment tasks, achieving superior performance with only 10\% of the original data. This principled, efficient selection method offers a promising solution for scaling LLM alignment with limited resources.

大模型对齐数据选择DPO

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。