让视觉语言模型通过画图来推理空间关系,提升几何理解能力。
Reinforcing Spatial Reasoning in Vision-Language Models with Interwoven Thinking and Visual Drawing

- 用画框和辅助线等基础绘图操作,让模型在视觉空间中直接推理。
- 在多个空间推理任务上平均提升18.4%,超越现有方法。
- 适合需要精准空间分析的场景,如导航、多视角理解。
随着大语言模型在文本推理上的进步,学术界对增强大视觉语言模型(LVLMs)的多模态推理能力日益关注。然而,现有方法多采用以文本为中心的直接推理方式,仅依赖多模态输入,缺乏对空间推理任务中精确几何理解与连续空间追踪的需求支持,这类任务人类可通过心理可视化与操作实现。为此,我们提出「画图推理」新范式,使LVLMs通过视觉空间中的基本绘图操作进行推理。通过赋予模型标注边界框、绘制辅助线等能力,使其能通过直接视觉操作表达和分析空间关系,同时避免以往工具集成方法中专用感知工具带来的性能瓶颈。为培养此能力,我们构建三阶段训练框架:基于合成数据的冷启动训练建立基础绘图能力,通过反思拒绝采样强化自我反思行为,再通过强化学习直接优化目标奖励。大量实验表明,所提出的VILASR模型在迷宫导航、静态空间推理、视频推理及多视角推理等多样化空间推理基准上持续领先,平均提升18.4%。
原文摘要 · Abstract (English)
As textual reasoning with large language models (LLMs) has advanced significantly, there has been growing interest in enhancing the multimodal reasoning capabilities of large vision-language models (LVLMs). However, existing methods primarily approach multimodal reasoning in a straightforward, text-centric manner, where both reasoning and answer derivation are conducted purely through text, with the only difference being the presence of multimodal input. As a result, these methods often encounter fundamental limitations in spatial reasoning tasks that demand precise geometric understanding and continuous spatial tracking-capabilities that humans achieve through mental visualization and manipulation. To address the limitations, we propose drawing to reason in space, a novel paradigm that enables LVLMs to reason through elementary drawing operations in the visual space. By equipping models with basic drawing operations, including annotating bounding boxes and drawing auxiliary lines, we empower them to express and analyze spatial relationships through direct visual manipulation, meanwhile avoiding the performance ceiling imposed by specialized perception tools in previous tool-integrated reasoning approaches. To cultivate this capability, we develop a three-stage training framework: cold-start training with synthetic data to establish basic drawing abilities, reflective rejection sampling to enhance self-reflection behaviors, and reinforcement learning to directly optimize for target rewards. Extensive experiments demonstrate that our model, named VILASR, consistently outperforms existing methods across diverse spatial reasoning benchmarks, involving maze navigation, static spatial reasoning, video-based reasoning, and multi-view-based reasoning tasks, with an average improvement of 18.4%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。