arXiv:2503.01268cs.LGcs.AI2025-03

突破模型合并的限制,实现与集成相当甚至更优的性能

Multi-Level Collaboration in Model Merging

  • 提出理论关联模型合并与集成的性能关系
  • 5个CLIP-ViT-B/32模型合并达95.44%准确率,接近集成的95.46%
  • 适用于多模型、不同规模,无需同源预训练

参数级模型合并是多任务学习中的新兴范式,具有巨大潜力。以往研究揭示其与预测级模型集成(常被视为合并上限)之间的联系,但依赖于特定前提:仅限两模型、使用ViT架构、且均从同一预训练检查点微调。本文探讨若移除这些限制,合并与集成是否仍能保持性能一致性。我们首先建立合并与集成间的性能关联理论,发现即使在无上述限制条件下,合并仍可达到与集成几乎相同甚至更优的性能。为验证其可行性,提出名为Neural Ligand(NeuLig)的验证框架,其学习过程基于具备理论支持的专用损失函数。实验表明,NeuLig在模型规模和协作模型数量上均具强鲁棒性。例如,在5个CLIP-ViT-B/32模型场景下,参数级合并达到95.44%准确率,与预测级集成的95.46%几乎持平。

原文摘要 · Abstract (English)

Parameter-level model merging is an emerging paradigm in multi-task learning with significant promise. Previous research has explored its connections with prediction-level model ensembling-commonly viewed as the upper bound for merging-to reveal the potential of achieving performance consistency between the two. However, this observation relies on certain preconditions, such as being limited to two models, using ViT-based models, and all models are fine-tuned from the same pre-trained checkpoint. To further understand the intrinsic connections between model merging and model ensembling, this paper explores an interesting possibility: If these restrictions are removed, can performance consistency still be achieved between merging and ensembling? To answer this question, we first theoretically establish a performance correlation between merging and ensembling. We find that even when previous restrictions are not met, there is still a way for model merging to attain a near-identical and superior performance similar to that of ensembling. To verify whether our findings are practical, we introduce a validation framework termed Neural Ligand (NeuLig). The learning process of NeuLig is meticulously designed with a specialized loss function supported by theoretical foundations. Experimental results demonstrate the robust resilience of NeuLig in terms of both model scale and the number of collaborating models. For instance, for the case involving 5 CLIP-ViT-B/32 models, parameter-level merging achieves the same performance as prediction-level ensembling (merging: 95.44% vs. ensembling: 95.46%).

模型合并多任务学习集成学习视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。