发现大模型多目标对齐中的干扰现象并提出缓解方法
Uncovering Cross-Objective Interference in Multi-Objective Alignment
- 通过协方差规律揭示目标间干扰机制
- 提出CTWA方法有效减少目标性能下降
- 适用于需平衡多个训练目标的场景
我们研究了大语言模型多目标对齐中的一种持续性失败模式:训练仅提升部分目标性能,导致其他目标退化。我们将此现象形式化为跨目标干扰,并首次系统性地考察了不同标量化的算法,发现干扰普遍存在且具有强模型依赖性。为解释该现象,我们推导出局部协方差定律,表明当某一目标奖励与标量得分正相关时,其性能会提升。我们将分析扩展至现代对齐中使用的截断代理目标,在温和条件下仍验证了协方差定律的有效性。基于此,我们提出即插即用的协方差目标权重自适应(CTWA)方法,通过维持目标奖励与训练信号间的正协方差来有效缓解跨目标干扰。最后,我们在Polyak–Łojasiewicz条件下进行了全局收敛分析,揭示非凸标量化优化实现全局收敛的条件,以及干扰如何依赖于具体模型几何性质。
原文摘要 · Abstract (English)
We study a persistent failure mode in multi-objective alignment for large language models (LLMs): training improves performance on only a subset of objectives while causing others to degrade. We formalize this phenomenon as cross-objective interference and conduct the first systematic study across scalarization algorithms, showing that interference is pervasive and exhibits strong model dependence. To explain this phenomenon, we derive a local covariance law showing that an objective improves when its reward exhibits positive covariance with the scalarized score. We extend this analysis to clipped surrogate objectives used in modern alignment, demonstrating that the covariance law remains valid under mild conditions despite clipping. Building on this analysis, we propose Covariance Targeted Weight Adaptation (CTWA), a plug-and-play method that maintains positive covariance between objective rewards and the training signal to effectively mitigate cross-objective interference. Finally, we complement these local improvement conditions with a global convergence analysis under the Polyak--Łojasiewicz condition, establishing when non-convex scalarized optimization achieves global convergence and how cross-objective interference depends on specific model geometric properties.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。