arXiv:2509.24610cs.LGcs.CL2025-09被引 6

通过正交分解解决多目标对齐中的参数冲突问题。

OrthAlign: Orthogonal Subspace Decomposition for Non-Interfering Multi-Objective Alignment

  • 将参数更新空间分解为互不干扰的正交子空间,实现多目标并行优化。
  • 在帮助性、无害性和真实性上分别提升34.61%至50.89%,整体奖励平均提升13.96%。
  • 理论保证稳定收敛,适合需要多维度精准对齐的LLM应用场景。

大型语言模型对齐在处理多重人类偏好时面临核心困境:某一维度的改进常以牺牲其他维度为代价,导致不可避免的权衡,如帮助性与无害性之间的冲突。现有方法主要依赖约束优化算法和数据选择策略缓解矛盾,但忽视了在参数层面直接解决冲突的本质问题。本文提出OrthAlign,开创性地采用正交子空间分解方法,在梯度层面根本性解决多目标偏好对齐中的冲突问题。该方法将参数更新空间划分为正交子空间,确保不同偏好的优化在数学上互不干扰。基于此,我们提供了理论证明:当参数增量满足正交子空间约束与谱范数边界时,更新呈现线性利普希茨增长而非指数级不稳定性,保障所有偏好维度的稳定收敛。大量实验表明:(1)经过多目标对齐后,单个偏好提升最高达34.61%至50.89%;(2)整体奖励平均提升13.96%。

原文摘要 · Abstract (English)

Large language model (LLM) alignment faces a critical dilemma when addressing multiple human preferences: improvements in one dimension frequently come at the expense of others, creating unavoidable trade-offs between competing objectives like helpfulness and harmlessness. While prior work mainly focuses on constraint-based optimization algorithms and data selection strategies to mitigate conflicts, these approaches overlook the fundamental issue of resolving conflicts directly at the parameter level. In this paper, we present OrthAlign, an innovative approach that pioneers a new paradigm by leveraging orthogonal subspace decomposition to fundamentally resolve gradient-level conflicts in multi-objective preference alignment. OrthAlign strategically decomposes parameter update spaces into orthogonal subspaces, ensuring that optimization toward different preferences occurs in mathematically non-interfering directions. Building upon this, we provide theoretical guarantees demonstrating that when parameter increments satisfy both orthogonal subspace constraints and spectral norm bounds, the resulting updates exhibit linear Lipschitz growth rather than exponential instability, ensuring stable convergence across all preference dimensions. Extensive experiments show that: I. OrthAlign achieves maximum single-preference improvements ranging from 34.61% to 50.89% after multiple-objective alignment across helpful, harmless, and truthful dimensions. II. With an average overall reward improvement of 13.96%.

多目标对齐正交分解参数优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。