arXiv:2607.09796cs.LG2026-07

在噪声偏好数据下,无需元数据也能提升大模型对齐效果。

Metadata-Free Meta-Reweighted Direct Preference Optimization under Noisy Preference Labels

  • 构建双层优化框架,自动学习偏好数据权重
  • 在不同噪声率下优于多个DPO基线方法
  • 无需元数据,适合真实场景中的模型对齐

直接偏好优化(DPO)因无需显式奖励建模和强化学习,已成为对齐大语言模型与人类偏好的重要方法。然而其性能高度依赖偏好数据质量,真实场景中的噪声数据会削弱对齐效果。为此,本文提出一种双层优化框架,并在理想条件下证明其可恢复干净数据下的DPO最优解。进一步推导出在标签翻转噪声下可学习权重函数的先验形式。考虑到高质量元数据难以获取,提出提示增强一致性方法,实现无元数据下的元学习。为降低大模型元学习中高阶梯度的计算成本,结合中心差分近似与LoRA微调,设计可扩展训练方案。在TL;DR摘要和Anthropic Helpful and Harmless对话数据集上的实验表明,该方法在不同噪声率下均优于多个DPO基线。

原文摘要 · Abstract (English)

Direct Preference Optimization (DPO) has become an important method for aligning large language models (LLMs) with human preferences because it removes the need for explicit reward modeling and reinforcement learning. However, its performance depends heavily on the quality of preference data, and noisy preference data in real-world settings can weaken alignment performance. To address this issue, we propose a bilevel optimization framework and prove, under some idealized conditions, that this framework can recover the DPO optimum under clean data. We further derive a prior form for the learnable weighting function under label-flipping noise. Considering that high-quality metadata may be difficult to obtain, we propose a prompt augmentation consistency method that enables meta-learning even when metadata is completely unavailable. To reduce the high cost of higher-order gradients in LLM meta-learning, we combine central-difference approximation with LoRA fine-tuning and develop a scalable training scheme. Experiments on TL;DR summarization and Anthropic Helpful and Harmless dialogue show that the proposed method improves alignment performance over multiple DPO baselines under different noise rates.

大模型对齐偏好优化噪声数据元学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。