arXiv:2505.05464cs.CL2025-05ICML被引 47

用模型合并让视觉模型学会推理,无需训练即可提升理解能力。

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

  • 跨模态合并视觉与语言模型参数,实现推理能力迁移。
  • 合并后所有层均参与推理,早期层仍主导感知功能。
  • 为多模态模型机制研究提供新工具,适合对模型可解释性感兴趣者。

视觉-语言模型(VLMs)将视觉感知与大型语言模型(LLMs)的通用能力(如推理)结合,但二者如何协同作用仍不明确。本文提出通过模型合并方式,将不同模态的模型参数连接起来,实现跨模态融合。不同于以往仅合并同类模型的方法,本工作探索将推理能力强的LLM与视觉模型合并,以在无训练情况下将推理能力迁移到视觉模型中。大量实验表明,该方法有效实现了能力迁移。进一步分析发现,感知能力主要存在于模型早期层,而推理则由中后期层主导;合并后,所有层均参与推理过程,但感知能力的分层分布基本不变。这些结果揭示了模型合并在多模态集成与机制解析方面的潜力。

原文摘要 · Abstract (English)

Vision-Language Models (VLMs) combine visual perception with the general capabilities, such as reasoning, of Large Language Models (LLMs). However, the mechanisms by which these two abilities can be combined and contribute remain poorly understood. In this work, we explore to compose perception and reasoning through model merging that connects parameters of different models. Unlike previous works that often focus on merging models of the same kind, we propose merging models across modalities, enabling the incorporation of the reasoning capabilities of LLMs into VLMs. Through extensive experiments, we demonstrate that model merging offers a successful pathway to transfer reasoning abilities from LLMs to VLMs in a training-free manner. Moreover, we utilize the merged models to understand the internal mechanism of perception and reasoning and how merging affects it. We find that perception capabilities are predominantly encoded in the early layers of the model, whereas reasoning is largely facilitated by the middle-to-late layers. After merging, we observe that all layers begin to contribute to reasoning, whereas the distribution of perception abilities across layers remains largely unchanged. These observations shed light on the potential of model merging as a tool for multimodal integration and interpretation.

模型合并视觉推理多模态可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。