arXiv:2507.20529cs.CVcs.AI2025-07被引 5

通过视觉与文本双重思考,提升模型空间推理能力。

Enhancing Spatial Reasoning through Visual and Textual Thinking

  • 双阶段思维:先生成目标位置特征,再基于视觉和对话逐步推理
  • 在多个空间理解任务上显著优于现有模型,无需额外标注信息
  • 改进数据集标注质量并重构输入格式,增强模型泛化能力

空间推理任务旨在理解二维与三维空间中的关系,是视觉问答和机器人技术的基础能力。尽管近年来视觉语言模型发展迅速,但在空间推理方面仍表现不佳。本文提出SpatialVTS方法,通过同步进行视觉与文本的双重思考来增强空间推理能力。在视觉思考阶段,模型自动生成关键目标的位置相关特征,不仅包含问题中提及的对象,还考虑潜在相关对象。在文本思考阶段,模型基于视觉线索和对话进行长期推理,逐步得出答案。为有效支持训练,我们对现有空间推理数据集进行了人工修正,剔除因自动标注导致的大量错误标签,重构数据输入格式以提升泛化能力,并构建具有逻辑细节的推理过程。无需引入额外信息(如掩码或深度图),本模型在多个空间理解任务上的整体平均性能显著优于其他模型。

原文摘要 · Abstract (English)

The spatial reasoning task aims to reason about the spatial relationships in 2D and 3D space, which is a fundamental capability for Visual Question Answering (VQA) and robotics. Although vision language models (VLMs) have developed rapidly in recent years, they are still struggling with the spatial reasoning task. In this paper, we introduce a method that can enhance Spatial reasoning through Visual and Textual thinking Simultaneously (SpatialVTS). In the spatial visual thinking phase, our model is trained to generate location-related specific tokens of essential targets automatically. Not only are the objects mentioned in the problem addressed, but also the potential objects related to the reasoning are considered. During the spatial textual thinking phase, Our model conducts long-term thinking based on visual cues and dialogues, gradually inferring the answers to spatial reasoning problems. To effectively support the model's training, we perform manual corrections to the existing spatial reasoning dataset, eliminating numerous incorrect labels resulting from automatic annotation, restructuring the data input format to enhance generalization ability, and developing thinking processes with logical reasoning details. Without introducing additional information (such as masks or depth), our model's overall average level in several spatial understanding tasks has significantly improved compared with other models.

空间推理视觉语言模型多模态推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。