arXiv:2502.01930cs.LGcs.AI2025-02NeurIPS被引 15

解决大模型对齐中的偏好分布偏移问题,提升实际应用的鲁棒性。

Robust LLM Alignment via Distributionally Robust Direct Preference Optimization

  • 基于分布鲁棒优化框架,提出两种新算法WDPO与KLDPO。
  • 在偏好分布偏移场景下,对齐效果显著优于传统方法。
  • 适合关注模型实际部署鲁棒性的研究人员和工程师。

大语言模型(LLM)对齐面临的主要挑战是偏好分布偏移问题。现有对齐算法依赖静态偏好数据集,假设其能准确反映真实用户偏好,但用户偏好在地理区域、人口统计、语言模式及文化趋势上存在显著差异。这种分布偏移导致许多实际应用中出现灾难性对齐失败。本文采用分布鲁棒优化的严谨框架,提出两种新型分布鲁棒直接偏好优化(DPO)算法:Wasserstein DPO(WDPO)与Kullback-Leibler DPO(KLDPO)。我们刻画了WDPO和KLDPO最优策略参数学习的样本复杂度,并通过设计合适的近似方法,开发出可扩展的梯度下降式学习算法以处理两者复杂的极小极大损失函数。在基准数据集和大语言模型上的实验证明,当存在偏好分布偏移时,WDPO与KLDPO显著提升了对齐性能。

原文摘要 · Abstract (English)

A major challenge in aligning large language models (LLMs) with human preferences is the issue of distribution shift. LLM alignment algorithms rely on static preference datasets, assuming that they accurately represent real-world user preferences. However, user preferences vary significantly across geographical regions, demographics, linguistic patterns, and evolving cultural trends. This preference distribution shift leads to catastrophic alignment failures in many real-world applications. We address this problem using the principled framework of distributionally robust optimization, and develop two novel distributionally robust direct preference optimization (DPO) algorithms, namely, Wasserstein DPO (WDPO) and Kullback-Leibler DPO (KLDPO). We characterize the sample complexity of learning the optimal policy parameters for WDPO and KLDPO. Moreover, we propose scalable gradient descent-style learning algorithms by developing suitable approximations for the challenging minimax loss functions of WDPO and KLDPO. Our empirical experiments using benchmark data sets and LLMs demonstrate the superior performance of WDPO and KLDPO in substantially improving the alignment when there is a preference distribution shift.

大模型对齐分布鲁棒偏好优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。