轻量级方法提升语言模型对偏好噪声的鲁棒性
Lightweight Robust Direct Preference Optimization
- 基于偏好分布不确定性设计轻量级鲁棒优化框架
- 在标准对齐任务和真实公共卫生任务中均显著提升抗噪能力
- 适合需要稳定微调且数据质量不高的实际应用场景
直接偏好优化(DPO)因其稳定性和简单性成为大语言模型微调的主流方法,但对数据噪声敏感且易过拟合。已有研究采用分布鲁棒优化(DRO)缓解数据噪声与分布偏移问题,但常导致过度保守且计算开销高。本文提出DPO-PRO(DPO with Preference Robustness),一种基于DPO的轻量级鲁棒微调算法,通过简洁的DRO形式建模偏好分布的不确定性。与以往方法不同,DPO-PRO仅关注偏好不确定性,避免不必要的保守性,计算开销可忽略。我们进一步证明DPO-PRO等价于一种正则化DPO目标,在弱偏好信号下惩罚模型过自信。在标准对齐基准和真实世界公共卫生任务上的实验表明,相比现有DPO变体,DPO-PRO在面对噪声偏好信号时表现出更强的鲁棒性。
原文摘要 · Abstract (English)
Direct Preference Optimization (DPO) has become a popular method for fine-tuning large language models (LLMs) due to its stability and simplicity. However, it is also known to be sensitive to noise in the data and prone to overfitting. Recent works have proposed using distributionally robust optimization (DRO) to address potential noise and distributional shift in the data. However, these methods often suffer from excessive conservatism and high computational cost. We propose DPO-PRO (DPO with Preference Robustness), a robust fine-tuning algorithm based on DPO which accounts for uncertainty in the preference distribution through a lightweight DRO formulation. Unlike prior DRO-based variants, DPO-PRO focuses solely on uncertainty in preferences, avoiding unnecessary conservatism and incurring negligible computational overhead. We further show that DPO-PRO is equivalent to a regularized DPO objective that penalizes model overconfidence under weak preference signals. We evaluate DPO-PRO on standard alignment benchmarks and a real-world public health task. Experimental results show that our method consistently improves robustness to noisy preference signals compared to existing DPO variants.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。