arXiv:2604.16060cs.CVcs.AI2026-04ACL被引 6

CoT推理会削弱多模态模型的视觉空间推理能力

Chain-of-Thought Degrades Visual Spatial Reasoning Capabilities of Multimodal LLMs

  • 用链式思维提示让多模态模型在空间任务上表现更差
  • 17个模型在13个空间基准上均出现性能下降
  • 适合关注视觉推理可靠性的研究者阅读

基于链式思维(CoT)的多模态推理模型(MRMs)在数学和逻辑问题解决中取得了突破,但我们发现该范式在泛化空间智能方面存在严重缺陷。我们在13个空间基准上对17个模型进行了全面评估,发现CoT提示导致视觉空间推理性能持续下降。通过创新的No-Image++消融实验,我们证实多模态模型和文本提示的单模态模型存在严重捷径学习现象,在图像缺失时仍会从文本先验中幻觉出视觉细节。这些发现挑战了纯文本链式思维在空间任务中的有效性,强调需要以视觉为中心的推理范式。

原文摘要 · Abstract (English)

Multimodal Reasoning Models (MRMs) leveraging Chain-of-Thought (CoT) based thinking have revolutionized mathematical and logical problem-solving. However, we show that this paradigm struggles with generalized spatial intelligence. We perform a comprehensive evaluation of seventeen models across thirteen spatial benchmarks and identify a critical gap: CoT prompting consistently degrades performance in visual spatial reasoning. Furthermore, through a novel No-Image++ ablation, we demonstrate that MRMs and CoT prompted MLMs suffer from severe shortcut learning, and hallucinate visual details from textual priors even when the image is absent. These findings challenge the efficacy of text-only CoT for spatial tasks and underscore the need for vision-centric reasoning paradigms.

多模态链式思维空间推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。