让不同结构的模型互换参数,动态提升基础模型能力
Model Assembly Learning with Heterogeneous Layer Weight Merging
- 通过迭代融合异构模型参数,实现跨架构知识迁移
- 可在层宽不匹配情况下完成参数合并,突破传统限制
- 为多模型融合提供可操作的规则和实践指南
模型合并通过整合多个模型的参数,在无需额外数据或训练的情况下获得通用能力。以往方法通过排列不变性将参数对齐至同一损失盆地,实现线性模式连通性。本文提出模型组装学习(MAL),一种新型模型合并范式,通过在开放式的模型库中迭代整合多样模型的参数,持续增强基础模型的能力。与以往要求相同架构的方法不同,MAL支持异构架构间的参数合并,并可选择性地融合各层参数。具体而言,基础模型可从多个预训练模型的不同层中吸收参数。我们系统研究了异构参数合并的条件与基本设置,解决基础模型与目标模型之间层宽可能存在的所有不匹配问题。此外,我们归纳出关键规律并提供切实可行的实施指南。
原文摘要 · Abstract (English)
Model merging acquires general capabilities without extra data or training by combining multiple models' parameters. Previous approaches achieve linear mode connectivity by aligning parameters into the same loss basin using permutation invariance. In this paper, we introduce Model Assembly Learning (MAL), a novel paradigm for model merging that iteratively integrates parameters from diverse models in an open-ended model zoo to enhance the base model's capabilities. Unlike previous works that require identical architectures, MAL allows the merging of heterogeneous architectures and selective parameters across layers. Specifically, the base model can incorporate parameters from different layers of multiple pre-trained models. We systematically investigate the conditions and fundamental settings of heterogeneous parameter merging, addressing all possible mismatches in layer widths between the base and target models. Furthermore, we establish key laws and provide practical guidelines for effectively implementing MAL.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。