arXiv:2501.07542cs.CLcs.CV2025-01ICML被引 220

让大模型用图像思考,提升空间推理能力

Imagine while Reasoning in Space: Multimodal Visualization-of-Thought

论文配图:Imagine while Reasoning in Space: Multimodal Visualization-of-Thought
图 1 · 摘自论文原文
  • 通过生成推理过程的图像可视化,实现多模态思维
  • 引入词元差异损失,显著提升图像一致性和真实度
  • 在复杂空间任务中表现优于传统文本推理,适合视觉思维场景

链式思维(CoT)提示在提升大语言模型和多模态大语言模型的复杂推理能力方面已证明极为有效,但在复杂的空间推理任务中仍表现不佳。然而,人类认知不仅依赖语言,还能以图像形式进行思考。受此启发,我们提出一种新的推理范式——多模态思维可视化(MVoT),通过生成模型推理轨迹的图像,使多模态大模型具备视觉化思维能力。为确保可视化质量,我们在自回归多模态大模型中引入词元差异损失,显著提升了图像的一致性与保真度。我们在多个动态空间推理任务上验证该方法,实验结果表明,MVoT在各项任务中均表现出竞争力,尤其在传统CoT失效的高难度场景中展现出稳健且可靠的性能提升。最终,MVoT为需要视觉思维辅助的复杂推理任务开辟了新可能。

原文摘要 · Abstract (English)

Chain-of-Thought (CoT) prompting has proven highly effective for enhancing complex reasoning in Large Language Models (LLMs) and Multimodal Large Language Models (MLLMs). Yet, it struggles in complex spatial reasoning tasks. Nonetheless, human cognition extends beyond language alone, enabling the remarkable capability to think in both words and images. Inspired by this mechanism, we propose a new reasoning paradigm, Multimodal Visualization-of-Thought (MVoT). It enables visual thinking in MLLMs by generating image visualizations of their reasoning traces. To ensure high-quality visualization, we introduce token discrepancy loss into autoregressive MLLMs. This innovation significantly improves both visual coherence and fidelity. We validate this approach through several dynamic spatial reasoning tasks. Experimental results reveal that MVoT demonstrates competitive performance across tasks. Moreover, it exhibits robust and reliable improvements in the most challenging scenarios where CoT fails. Ultimately, MVoT establishes new possibilities for complex reasoning tasks where visual thinking can effectively complement verbal reasoning.

视觉思维空间推理多模态模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。