arXiv:2601.06672cs.CL2026-01ACL被引 3

揭示模型合并成功率的决定因素,发现基础模型知识是关键

Will it Merge? On The Causes of Model Mergeability

  • 提出可量化的合并能力定义,用于评估模型是否易合并
  • 发现基础模型熟悉的任务更易实现高质量合并
  • 设计加权合并方法,更好保留基础模型的弱知识

模型合并作为一种无需重新训练即可将多个微调模型融合为单一多任务模型的有前景技术,其成功与否的决定因素仍不明确。本文提出一个具体且可测量的合并能力定义,探究影响合并性能的多种潜在原因。研究发现,基础模型的知识水平是主导因素:在基础模型已掌握的任务上微调的模型,比在基础模型难以处理的任务上微调的模型更易合并。基于该定义,我们提出一种简单的加权合并策略,能更有效地保留基础模型中的弱知识。

原文摘要 · Abstract (English)

Model merging has emerged as a promising technique for combining multiple fine-tuned models into a single multitask model without retraining. However, the factors that determine whether merging will succeed or fail remain poorly understood. In this work, we investigate why specific models are merged better than others. To do so, we propose a concrete, measurable definition of mergeability. We investigate several potential causes for high or low mergeability, highlighting the base model knowledge as a dominant factor: Models fine-tuned on instances that the base model knows better are more mergeable than models fine-tuned on instances that the base model struggles with. Based on our mergeability definition, we explore a simple weighted merging technique that better preserves weak knowledge in the base model.

模型合并微调知识保留

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。