视觉输入让模型看懂却不会推理,提出路由干扰机制并改进效果
Seeing but Not Thinking: Routing Distraction in Multimodal Mixture-of-Experts
- 发现视觉输入在中间层路由偏差,导致推理专家激活不足
- 干预路由后多模型多任务提升,复杂视觉推理最高增3.17%
- 识别出具泛化能力的推理专家,可跨任务迁移
多模态混合专家(MoE)模型在视觉语言任务中表现优异,但存在一种奇怪现象:模型能准确感知图像内容,却在后续推理中失败,而相同问题以纯文本形式则能正确解答。通过系统分析,我们验证了跨模态语义共享的存在,排除了语义对齐失败的单一解释。进一步发现,视觉专家与领域专家在层级上呈现分离,图像输入在中间层引发显著的路由偏移,而该层正是领域专家集中所在。基于此,我们提出路由干扰假说:处理视觉输入时,路由机制未能充分激活相关推理专家。为此设计路由引导干预方法以增强领域专家激活。在三个多模态MoE模型、六个基准测试上的实验表明,性能持续提升,复杂视觉推理任务最高提升3.17%。分析还显示,领域专家的识别对应认知功能而非样本特异性解法,支持跨任务有效迁移。
原文摘要 · Abstract (English)
Multimodal Mixture-of-Experts (MoE) models have achieved remarkable performance on vision-language tasks. However, we identify a puzzling phenomenon termed Seeing but Not Thinking: models accurately perceive image content yet fail in subsequent reasoning, while correctly solving identical problems presented as pure text. Through systematic analysis, we first verify that cross-modal semantic sharing exists in MoE architectures, ruling out semantic alignment failure as the sole explanation. We then reveal that visual experts and domain experts exhibit layer-wise separation, with image inputs inducing significant routing divergence from text inputs in middle layers where domain experts concentrate. Based on these findings, we propose the Routing Distraction hypothesis: when processing visual inputs, the routing mechanism fails to adequately activate task-relevant reasoning experts. To validate this hypothesis, we design a routing-guided intervention method that enhances domain expert activation. Experiments on three multimodal MoE models across six benchmarks demonstrate consistent improvements, with gains of up to 3.17% on complex visual reasoning tasks. Our analysis further reveals that domain expert identification locates cognitive functions rather than sample-specific solutions, enabling effective transfer across tasks with different information structures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。