将专家大模型技能注入视觉语言模型,提升特定领域能力。
Investigating Cross-Modal Skill Injection: Scenarios, Methods, and Hyperparameters
- 通过模型融合将领域专家LLM的技能注入VLM,无需额外训练数据。
- 指令遵循和跨语言任务表现优异,数学推理能力仍不足。
- 发现TA和DARE方法最优,且超参数对性能影响显著。
视觉-语言模型(VLM)在通用多模态理解上表现出色,但在持续演进的领域特定技能获取上效率低下。传统增强方法如监督微调(SFT)需大量数据标注和计算资源。模型融合成为高效替代方案,可在不增加训练数据或显著计算开销的前提下,将大型语言模型(LLM)的领域专长转移至VLM。与同质LLM融合仅聚合已有能力不同,跨模态技能注入旨在通过整合领域专家LLM,催生新的跨模态能力。然而现有研究缺乏对跨模态技能注入适用场景、方法及超参数的系统分析。本研究从三个维度展开:场景、方法与超参数。结果显示,跨模态技能注入在指令遵循和跨语言任务中表现良好,但在数学推理任务中效果有限;经典方法如TA和DARE始终优于其他融合方法;同时我们对这些方法依赖的关键超参数进行了系统性定量分析。
原文摘要 · Abstract (English)
Vision-Language Models (VLMs) have demonstrated remarkable proficiency in general multi-modal understanding; yet they struggle to efficiently acquire continually evolving domain-specific skills. Conventional approaches to enhancing VLM capabilities, such as Supervised Fine-Tuning (SFT), require extensive dataset curation and substantial computational resources. Model merging has emerged as an efficient alternative that enables the transfer of domain-specific expertise from Large Language Models (LLMs) to VLMs without incurring additional training data requirements or significant computational overhead. Unlike conventional merging of homogeneous LLMs, which mainly aggregates existing capabilities, cross-modal skill injection aims to induce emergent cross-modal capabilities by integrating a domain-expert LLM into a VLM. However, existing research lacks a systematic analysis of the applicability and methodology of cross-modal skill injection. In this study, we investigate cross-modal skill injection across three main aspects: scenarios, methods, and hyperparameters. For scenarios, we find that cross-modal skill injection generally performs well in instruction-following and cross-lingual settings, yet struggles with mathematical reasoning. For methods, we find that classic approaches such as TA and DARE consistently achieve superior performance over alternative merging methods. We also provide a systematic and quantitative analysis of the hyperparameter tuning that these classic methods critically depend on.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。