arXiv:2606.22565cs.CLcs.AI2026-06ACL被引 2

探究多模态思维链的优劣,发现它对推理有效但易弱化视觉理解。

Look Light, Think Heavy: What Multimodal Chain-of-Thought Reasoning Can and Cannot Do

论文配图:Look Light, Think Heavy: What Multimodal Chain-of-Thought Reasoning Can and Cannot Do
图 1 · 摘自论文原文
  • 对比12项任务,发现思维链在数学科学推理中有效,但会损害视觉定位与计数能力。
  • 开源多模态推理模型整体提升有限,过度侧重数学而忽视综合能力。
  • 模型存在'轻视视觉、重于语言'现象,视觉反思随推理过程持续减弱。

思维链(CoT)已成为提升大语言模型推理能力的标准方法,通过引导分步思考来增强表现,但在多模态任务中的有效性尚不明确。本文系统研究核心问题:多模态思维链能做什么,又在何处失效?我们评估了12个跨感知与推理类别的多模态任务,涵盖14个非推理模型和8个推理模型。结果表明:(1) 思维链并非万能,需根据任务需求选择使用;在感知任务中,其可能导致性能下降,如视觉定位和物体计数能力减弱;而在数学、科学及多图像推理任务中表现良好;(2) 现有开源多模态推理模型相比原模型仅带来微弱整体提升,可能因过度强调数学推理而牺牲更广泛能力;(3) 视觉推理仍是关键瓶颈,模型呈现‘看轻、想重’特征——语言反思上升,而视觉反思持续下降。这表明当前多模态思维链虽能处理语言层面的推演,却难以维持深层视觉内省。

原文摘要 · Abstract (English)

Chain-of-Thought (CoT) has become a standard method for improving reasoning capabilities in large language models (LLMs) by eliciting step-by-step thinking, but its effectiveness in multimodal tasks remains unclear. In this paper, we aim to systematically investigate the key question: What can multimodal Chain-of-Thought reasoning do, and where and why does it fall short? To this end, we evaluate 12 multimodal tasks across perception and reasoning categories using both 14 non-reasoning models and 8 reasoning models. Our analysis reveals several important findings: (1) CoT is not a free lunch and should be used selectively depending on the specific requirements of each task. For perception tasks, CoT can lead to undesirable side effects, such as reduced performance in visual grounding and object counting. In contrast, it proves effective for reasoning tasks involving mathematical, scientific, and multi-image reasoning; (2) Compared to original models, existing open-source multimodal reasoning models often yield only marginal overall improvements, possibly due to an overemphasis on mathematical reasoning at the expense of broader capabilities; (3) Visual reasoning remains a key bottleneck for current multimodal CoT, as models exhibit a Look Light, Think Heavy pattern where verbal reflection rises and falls during reasoning, whereas visual reflection consistently diminishes. These findings suggest that while multimodal CoT handles verbal reflection relatively well, it lacks the ability to maintain deep visual introspection throughout the reasoning process.

多模态思维链推理视觉理解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。