arXiv:2604.01837cs.CL2026-04

用最优传输理论提升大模型对齐效果,更稳定且更懂人类偏好。

PLOT: Enhancing Preference Learning via Optimal Transport

  • 将偏好学习建模为最优传输问题,从词元层面优化输出分布。
  • 在7个子偏好任务上均提升对齐效果,同时保持语言流畅性。
  • 适合关注大模型对齐理论基础与高效训练的研究者。

大语言模型的偏好学习已取得显著进展,但现有方法仍受限于性能提升有限、计算成本高、超参数敏感以及对全局词元关系建模不足等问题。本文提出PLOT,通过最优传输导出的词元级损失函数,增强基于微调的对齐效果。将偏好学习形式化为最优传输问题,使模型输出与人类偏好对齐的同时保留原始语言模型分布,确保稳定性与鲁棒性。此外,PLOT利用词元嵌入捕捉语义关系,实现全局信息驱动的优化。在两类偏好(人类价值观与逻辑与问题求解)共七个子偏好上的实验表明,PLOT持续提升对齐性能,同时保持语言流畅性与连贯性。结果证实最优传输是偏好学习的合理范式,为大模型偏好学习提供了理论基础与新视角。

原文摘要 · Abstract (English)

Preference learning in Large Language Models (LLMs) has advanced significantly, yet existing methods remain limited by modest performance gains, high computational costs, hyperparameter sensitivity, and insufficient modeling of global token-level relationships. We introduce PLOT, which enhances Preference Learning in fine-tuning-based alignment through a token-level loss derived from Optimal Transport. By formulating preference learning as an Optimal Transport Problem, PLOT aligns model outputs with human preferences while preserving the original distribution of LLMs, ensuring stability and robustness. Furthermore, PLOT leverages token embeddings to capture semantic relationships, enabling globally informed optimization. Experiments across two preference categories - Human Values and Logic & Problem Solving - spanning seven subpreferences demonstrate that PLOT consistently improves alignment performance while maintaining fluency and coherence. These results substantiate optimal transport as a principled methodology for preference learning, establishing a theoretically grounded framework that provides new insights for preference learning of LLMs.

大模型对齐最优传输偏好学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。