arXiv:2608.30597cs.LGcs.CL2026-08

改进偏好学习中的噪声标签,让模型更稳定可靠。

PLC-DPO: Posterior Label Correction in Noisy and Ambiguous Preference Optimization

论文配图:PLC-DPO: Posterior Label Correction in Noisy and Ambiguous Preference Optimization
图 1 · 摘自论文原文
  • 根据策略与参考模型的差距动态修正偏好信号方向
  • 在57组实验中平均胜率比DPO高5个百分点
  • 适合处理有错误或模糊标注的数据场景

直接偏好优化(DPO)通过成对比较简化对齐过程,但假设所有观察到的偏好都是可靠的。真实数据常违反此假设,导致反转、弱或模糊标签,引发有害的策略更新。为此,我们提出后验标签修正的DPO(PLC-DPO),通过将每对训练信号路由为干净、翻转或平局三种情况,实现鲁棒的偏好优化。核心思想是利用校准后的策略-参考边界作为在线证据,采取适当的修正动作。这将噪声偏好学习重新定义为主动修正监督方向与强度,而非仅过滤可疑样本。在57个数据集-模型-评估组合中,PLC-DPO相较于DPO取得最高平均胜率(60.5对比次优方法的55.5)。注入噪声、平局压力测试、人类不一致分析及自确认诊断进一步表明,该路由机制保持稳定,并能有效区分翻转与弱方向性样本。

原文摘要 · Abstract (English)

Direct Preference Optimization (DPO) simplifies alignment through pairwise comparisons but assumes all observed preferences are reliable. Real data often violates this assumption, leading to reversed, weak, or ambiguous labels that cause harmful policy updates. To address this, we propose Posterior Label Correction DPO (PLC-DPO) to robustly optimize preferences by routing each pair's training signal as a clean, flip, or tie case. The key idea is to use the calibrated policy-reference margin as online evidence to take appropriate correction actions. This reframes noisy preference learning as actively correcting supervision direction and strength rather than merely filtering suspicious examples. Across 57 dataset-model-benchmark cells, PLC-DPO obtains the best mean win rate against DPO (60.5 vs. 55.5 for the next-best method). Injected-noise and tie stress tests, human disagreement analysis, and self-confirmation diagnostics further show that the routing remains stable and distinguishes flipped from weakly directional pairs.

偏好学习鲁棒优化标签修正

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。