arXiv:2511.07691cs.CLcs.AI2025-11中稿 · IJCNLP-AACL 2025 F…被引 2

CAPO通过动态调整学习信号,让多语言模型更准确地对齐人类偏好。

CAPO: Confidence Aware Preference Optimization Learning for Multilingual Preferences

  • 根据偏好对的置信度动态调节损失,取代固定处理方式
  • 在多语言场景下奖励准确率提升至少16%,优选与次优响应差距更大
  • 适合需要高鲁棒性多语言对齐的应用场景

偏好优化是大语言模型对齐人类偏好的关键后训练技术,通常通过在排序响应对上微调实现。尽管直接偏好优化(DPO)在英语中表现良好,但在多语言环境下往往难以泛化。本文提出一种简单而有效的方法——置信度感知偏好优化(CAPO),将DPO对偏好对的固定处理替换为基于相对奖励的动态损失缩放机制。通过根据每个偏好对的置信度调节学习信号,CAPO增强了对噪声或低边际比较的鲁棒性,这类情况在多语言文本中普遍存在。实证结果显示,CAPO在奖励准确率上优于现有基线至少16%,并通过扩大不同语言中优选与次优响应间的差距,提升了对齐效果。

原文摘要 · Abstract (English)

Preference optimization is a critical post-training technique used to align large language models (LLMs) with human preferences, typically by fine-tuning on ranked response pairs. While methods like Direct Preference Optimization (DPO) have proven effective in English, they often fail to generalize robustly to multilingual settings. We propose a simple yet effective alternative, Confidence-Aware Preference Optimization (CAPO), which replaces DPO's fixed treatment of preference pairs with a dynamic loss scaling mechanism based on a relative reward. By modulating the learning signal according to the confidence in each preference pair, CAPO enhances robustness to noisy or low-margin comparisons, typically encountered in multilingual text. Empirically, CAPO outperforms existing preference optimization baselines by at least 16% in reward accuracy, and improves alignment by widening the gap between preferred and dispreferred responses across languages.

偏好优化多语言LLM对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。