无需标注数据,融合异构多模态大模型能力
AdaMMS: Model Merging for Heterogeneous Multimodal Large Language Models with Unsupervised Coefficient Optimization
- 设计映射函数适配不同架构的多模态模型
- 通过线性插值缓解参数空间不对称问题
- 无监督调参策略提升跨模型融合效果
近期模型融合方法在整合多个大语言模型的能力方面表现出强大潜力。然而,现有方法主要针对同构模型(架构相同),在处理具有固有异构性的多模态大语言模型(MLLMs)时面临挑战,包括模型架构差异及参数空间不对称性。本文提出AdaMMS,一种专为异构MLLMs设计的新融合方法,包含三个步骤:映射、融合与搜索。首先,设计模型间映射函数以实现异构架构下的融合;其次,在权重上应用线性插值,主动适应参数空间的不对称性;最后,在超参数搜索阶段提出一种无监督的超参数选择方法。作为首个可在无标注数据条件下融合异构多模态大模型的方法,大量实验表明,AdaMMS在多种模型组合和视觉-语言基准测试中均优于以往融合方法。
原文摘要 · Abstract (English)
Recently, model merging methods have demonstrated powerful strengths in combining abilities on various tasks from multiple Large Language Models (LLMs). While previous model merging methods mainly focus on merging homogeneous models with identical architecture, they meet challenges when dealing with Multimodal Large Language Models (MLLMs) with inherent heterogeneous property, including differences in model architecture and the asymmetry in the parameter space. In this work, we propose AdaMMS, a novel model merging method tailored for heterogeneous MLLMs. Our method tackles the challenges in three steps: mapping, merging and searching. Specifically, we first design mapping function between models to apply model merging on MLLMs with different architecture. Then we apply linear interpolation on model weights to actively adapt the asymmetry in the heterogeneous MLLMs. Finally in the hyper-parameter searching step, we propose an unsupervised hyper-parameter selection method for model merging. As the first model merging method capable of merging heterogeneous MLLMs without labeled data, extensive experiments on various model combinations demonstrated that AdaMMS outperforms previous model merging methods on various vision-language benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。