发现小模型只学到了大模型的高阶推理能力,感知能力几乎没变。
Understanding Knowledge Transfer Mechanism in Heterogeneous MLLM Fusion: A Simple Linear Approach

- 用线性探针分析跨尺度知识迁移,找出关键注入方向
- 实验显示仅高阶推理显著提升,感知能力基本不变
- 适合关注多模态模型融合机制的研究者
无训练融合异构多模态大模型(MLLMs)为跨尺度能力迁移提供了直接路径,但性能提升并未揭示小模型实际继承了什么。现有研究多局限于有限任务集或综合指标;当评估扩展至更广泛的任务集合时,不同能力是否能跨尺度迁移仍不明确。为此,我们提出跨尺度定向参数注入(CDPI),一种简单的线性探针,用于分析异构融合中的跨尺度知识转移。局部理论分析表明,知识迁移的选择性由一阶上对共享注入方向的能力响应决定,而二阶曲率效应限制了有效转移范围。在四个Qwen3-VL模型组合与十二个多模态基准上的实验揭示出一致的选择性模式:增益集中于推理,尤其是高阶推理,而感知性能接近原目标模型水平。组件消融显示高阶推理提升主要来自语言模型部分,比例分析表明正向选择性迁移主要发生在小比例情形。这些发现将跨尺度异构MLLM融合重新理解为在窄且低干扰范围内以语言侧推理为主的有选择性迁移,而非广泛的能力继承。
原文摘要 · Abstract (English)
Training-free fusion of heterogeneous multimodal large language models (MLLMs) provides a direct route for cross-scale capability transfer, yet improvements in aggregate performance do not reveal what a smaller model actually inherits. Existing studies are largely designed and evaluated on limited task sets or aggregate metrics; as evaluation expands to broader task collections, whether different capabilities can transfer across scales remains poorly understood. To investigate this question, we introduce Cross-Scale Directional Parameter Injection (CDPI), a simple linear probe to analyze cross-scale knowledge transfer during heterogeneous fusion. A local theoretical analysis indicates that knowledge transfer selectivity is determined at first order by capability-dependent responses to a shared injection direction, while second-order curvature effects constrain the effective transfer regime. Across four Qwen3-VL model pairs and twelve multimodal benchmarks, our experiments reveal a consistent pattern of selectivity: gains concentrate on reasoning, particularly high-level reasoning, whereas perception performance remains close to that of the original target model. Component-wise ablations further show that high-level reasoning gains arise primarily from the language model, while ratio analysis finds that positive selective transfer occurs mainly in the small-ratio regime. These findings recast cross-scale heterogeneous MLLM fusion as selective language-side reasoning transfer within a narrow, low-interference regime, rather than broad capability inheritance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。