提出几何感知的多目标优化方法,让大模型更公平地兼顾有用性、真实性和安全性。
MGDA-Decoupled: Geometry-Aware Multi-Objective Optimisation for DPO-based LLM Alignment
- 基于几何分析找到各目标的共享下降方向,显式考虑收敛动态差异。
- 在UltraFeedback数据集上,对齐效果超越其他方法,各项指标均最优。
- 无需强化学习或奖励模型,直接在轻量DPO框架内实现高效优化,适合追求公平性的研究者。
将大语言模型对齐至人类期望价值需平衡多个可能冲突的目标(如有用性、真实性、无害性),构成多目标优化挑战。现有对齐流程依赖固定加权方案,易系统性低估较难优化或少数目标,造成过程不公。为此,本文提出MGDA-Decoupled,一种基于几何的多目标优化算法,可在显式考虑各目标收敛动态的前提下,寻找共享下降方向。相比依赖强化学习(如GAPO)或显式奖励模型(如MODPO)的方法,本方法完全运行于轻量级直接偏好优化(DPO)范式中。在UltraFeedback数据集上的实验表明,几何感知方法——尤其是MGDA-Decoupled——在整体及各目标上的胜率均优于基准,表现最佳。
原文摘要 · Abstract (English)
Aligning large language models (LLMs) to desirable human values requires balancing multiple, potentially conflicting objectives such as helpfulness, truthfulness, and harmlessness, which presents a multi-objective optimisation challenge. Most alignment pipelines rely on a fixed scalarisation of these objectives, which can introduce procedural unfairness by systematically under-weighting harder-to-optimise or minority objectives. To promote more equitable trade-offs, we introduce MGDA-Decoupled, a geometry-based multi-objective optimisation algorithm that finds a shared descent direction while explicitly accounting for each objective's convergence dynamics. In contrast to prior methods that depend on reinforcement learning (e.g., GAPO) or explicit reward models (e.g., MODPO), our approach operates entirely within the lightweight Direct Preference Optimisation (DPO) paradigm. Experiments on the UltraFeedback dataset show that geometry-aware methods -- and MGDA-Decoupled in particular -- achieve the highest win rates against golden responses, both overall and per objective.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。