用模型融合保持多模态模型的语言能力
Model Merging to Maintain Language-Only Performance in Developmentally Plausible Multimodal Models
- 通过加权线性插值融合多模态与纯语言模型参数
- 融合后语言任务准确率提升,多模态性能不变
- 适合关注多模态模型语言能力的开发者
当前最先进的视觉-语言模型参数量巨大,训练数据远超儿童语言习得量。本文针对BabyLM挑战的多模态赛道,采用发展上合理的低资源数据集构建纯语言与多模态模型,所提多模态模型优于以往BabyLM基线。已有研究发现,多模态模型在仅考察语法的语言任务中表现较差。为此,本文聚焦于维持多模态模型的语言能力,探索模型融合策略:通过加权线性插值将多模态模型与纯语言模型参数合并。实验结果验证了多模态模型在语言任务中的不足,且融合纯语言模型参数可部分缓解该问题,同时保持多模态性能。
原文摘要 · Abstract (English)
State-of-the-art vision-and-language models consist of many parameters and learn from enormous datasets, surpassing the amounts of linguistic data that children are exposed to as they acquire a language. This paper presents our approach to the multimodal track of the BabyLM challenge addressing this discrepancy. We develop language-only and multimodal models in low-resource settings using developmentally plausible datasets, with our multimodal models outperforming previous BabyLM baselines. One finding in the multimodal language model literature is that these models tend to underperform in \textit{language-only} tasks. Therefore, we focus on maintaining language-only abilities in multimodal models. To this end, we experiment with \textit{model merging}, where we fuse the parameters of multimodal models with those of language-only models using weighted linear interpolation. Our results corroborate the findings that multimodal models underperform in language-only benchmarks that focus on grammar, and model merging with text-only models can help alleviate this problem to some extent, while maintaining multimodal performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。