揭示视觉思维如何提升多模态推理,统一解释不同方法的原理
Visual Thoughts: A Unified Perspective of Understanding Multimodal Chain-of-Thought
- 提出视觉思维概念,解释多模态思维链提升性能的本质
- 发现视觉思维表达形式影响清晰度与简洁性,决定改进程度
- 揭示视觉思维作为图像与深层推理的中间桥梁,促进信息传递
大型视觉语言模型在多模态任务中表现优异,多模态思维链(MCoT)进一步提升了性能与可解释性。现有MCoT方法分为两类:(i) 文本式多模态思维链(T-MCoT),输入多模态内容,输出文本;(ii) 交错式多模态思维链(I-MCoT),生成图文交错输出。尽管进展显著,其机制仍不明确。本文首次揭示,MCoT通过引入视觉思维提升性能——无论格式如何,只要表达清晰简洁,就能将图像信息融入推理过程。我们定义了四种视觉思维表达形式,并系统分析其在清晰度与简洁性上的差异,发现不同形式导致不同程度的性能提升。此外,研究发现视觉思维在模型内部充当图像与深层Transformer层之间的中介,促进更高级别的视觉信息传输。该发现为未来MCoT研究提供新视角。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) have achieved significant success in multimodal tasks, with multimodal chain-of-thought (MCoT) further enhancing performance and interpretability. Recent MCoT methods fall into two categories: (i) Textual-MCoT (T-MCoT), which takes multimodal input and produces textual output; and (ii) Interleaved-MCoT (I-MCoT), which generates interleaved image-text outputs. Despite advances in both approaches, the mechanisms driving these improvements are not fully understood. To fill this gap, we first reveal that MCoT boosts LVLMs by incorporating visual thoughts, which convey image information to the reasoning process regardless of the MCoT format, depending only on clarity and conciseness of expression. Furthermore, to explore visual thoughts systematically, we define four distinct forms of visual thought expressions and analyze them comprehensively. Our findings demonstrate that these forms differ in clarity and conciseness, yielding varying levels of MCoT improvement. Additionally, we explore the internal nature of visual thoughts, finding that visual thoughts serve as intermediaries between the input image and reasoning to deeper transformer layers, enabling more advanced visual information transmission. We hope that the visual thoughts can inspire further breakthroughs for future MCoT research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。