arXiv:2505.10892cs.LG2025-05被引 16

让大模型同时满足多个冲突目标,提升对齐效果

Multi-Objective Preference Optimization: Improving Human Alignment of Generative Models

  • 基于配对偏好直接优化,用约束条件平衡多个目标
  • 在人类偏好数据上微调后,性能超越基线且更稳定
  • 适合需要兼顾帮助性与安全性的实际应用

通过强化学习与偏好优化方法(如DPO、IPO)对大语言模型进行后训练,显著提升了对齐效果,但这些方法假设仅存在单一目标。现实中,人类表达的目标往往多重且相互冲突,如有益性和无害性,无法自然标量化解。本文研究多目标偏好对齐问题,即策略需同时平衡多个目标。提出多目标偏好优化(MOPO),一种受约束的KL正则化框架,在最大化主要目标的同时,通过可调安全阈值确保次要目标不低于下限。MOPO直接作用于配对偏好,无需点级奖励,支持简单闭式迭代更新。实验表明,MOPO在合成基准上可恢复帕累托最优策略;在人类偏好数据上微调后,生成的百亿参数模型获得更高奖励,且帕累托优于基线,优化过程稳定可靠。

原文摘要 · Abstract (English)

Post-training LLMs with RLHF and preference optimization methods (e.g., DPO, IPO) has greatly improved alignment, yet these approaches assume a single objective. In reality, humans express multiple, often conflicting objectives, such as helpfulness and harmlessness, with no natural scalarization. We study the multi-objective preference alignment problem, where a policy must balance several objectives simultaneously. We propose Multi-Objective Preference Optimization (MOPO), a constrained KL-regularized framework that maximizes a primary objective while enforcing lower bounds on secondary objectives via tunable safety thresholds. MOPO operates directly on pairwise preferences without point-wise rewards, and admits simple closed-form iterative updates. Empirically, MOPO recovers Pareto-optimal policies on synthetic benchmarks and, when fine-tuned on human-preference data, yields multi-billion parameter models that achieve higher rewards and Pareto-dominate baselines, with stable and robust optimization dynamics.

大模型对齐多目标优化偏好学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。