arXiv:2512.00391cs.LGcs.AI2025-12

通过方向对齐提升模型融合效果,解决参数不一致问题

From Coefficients to Directions: Rethinking Model Merging with Directional Alignment

  • 提出方向对齐框架,统一优化参数与特征空间方向
  • 在多个基准上实现更优性能,尤其适用于不同训练模型
  • 适合关注模型融合效率与稳定性的研究者

模型融合已成为一种无需联合训练即可整合多个独立训练模型的有效方法。现有方法多基于参数分解或系数优化,虽减少训练成本并取得良好表现,但忽视了参数与特征空间中的方向信息作用。实践中,简单融合会破坏主导参数方向的一致性,导致结构失衡,影响性能。此外,系数优化隐含假设各模型特征方向兼容,但神经坍缩现象表明类别特征具有结构性方向模式,独立训练的模型间可能不一致,仅靠系数优化不足以保证融合质量。本文强调方向对齐的重要性,提出统一几何框架——方向对齐融合(Merging with Directional Alignment),在参数与特征空间中同步对齐方向结构。分析表明该方法增强结构一致性,跨基准、模型规模与任务配置的实验验证其有效性。

原文摘要 · Abstract (English)

Model merging has emerged as a practical paradigm for integrating multiple independently trained models into a single model without joint retraining. Previous studies have demonstrated the effectiveness of combining parameters through strategies such as parameter decomposition, coefficient optimization, and subspace learning, significantly reducing the need for expensive joint training and achieving strong empirical performance across diverse tasks. However, these approaches predominantly treat merging as a problem of parameter space decomposition or fusion coefficient optimization, while overlooking the critical role of directional information in both parameter and feature spaces. In practice, naïve merging introduces inconsistencies in dominant parameter directions and disrupts structural coherence across models, which can degrade performance. Moreover, coefficient-based optimization methods implicitly assume compatible feature-space directions across models. However, Neural Collapse indicates that class features follow structured directional patterns, which may differ across independently trained models, making coefficient optimization alone insufficient. In this work, we emphasize the importance of \emph{directional alignment} and introduce a unified geometric framework, \emph{Merging with Directional Alignment} (\method{}), which aligns directional structures consistently in both the parameter and feature spaces. Our analysis shows that directional alignment improves structural coherence, and extensive experiments across benchmarks, model scales, and task configurations further validate the effectiveness of our approach.

模型融合方向对齐神经坍缩参数优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。